Systems, apparatuses, methods, and non-transitory computer-readable storage devices for enhancing and distributing synthetic dataset generating models
By generating synthetic datasets to train dataset generating models, the method addresses the limitations of federated learning, enhancing its flexibility, scalability, and privacy, and reducing communication costs, thus improving the efficiency and reliability of machine learning models across diverse environments.
Patent Information
- Application Number
- PCT/CN2024/070722
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-10
AI Technical Summary
Current machine learning models, particularly in federated learning, face challenges such as task-specific alignment, synchronization complexity, and increased communication requirements, which limit their flexibility and scalability.
A method involving the generation of synthetic datasets using authentic datasets to train dataset generating models, allowing for the combination of these datasets to create updated models capable of simulating multiple authentic datasets, facilitating decentralized and flexible training across a network infrastructure.
This approach enhances the universality and applicability of federated learning by reducing communication costs, improving knowledge transfer efficiency, and ensuring consistent performance despite data heterogeneity, while maintaining privacy and adaptability to varying resource constraints.
Smart Images

Figure CN2024070722_10072025_PF_FP_ABST
Abstract
Description
SYSTEMS, APPARATUSES, METHODS, AND NON-TRANSITORY COMPUTER-READABLE STORAGE DEVICES FOR ENHANCING AND DISTRIBUTING SYNTHETIC DATASET GENERATING MODELS
[0001] FIELD OF THE DISCLOSURE
[0002] The present disclosure relates generally to systems, apparatuses, methods, and non-transitory computer-readable storage devices for training artificial intelligence (AI) models, and more particularly to systems, methods, and non-transitory computer-readable storage devices for enhancing and distributing synthetic dataset generating models through network-based collaborative training.BACKGROUND
[0003] In the field of AI, the training of machine learning models is a fundamental aspect that drives the advancement and applicability of AI technologies in various domains. There are two main methods for using private user data to train these models: Centralized Learning and Federated Learning (FL) . Centralized learning involves aggregating user data on a central server where model training takes place. While this method is straightforward, it raises significant privacy concerns due to the need to collect raw, private user data. On the other hand, federated learning represents a collaborative approach where models are trained without sharing private data. In this method, the model parameters, rather than the raw data, are shared among the participants. This approach inherently respects user privacy better than centralized methods because it avoids the direct collection of personal data. However, both methodologies have their own challenges and limitations for the advancement of AI technologies.
[0004] The current state of machine learning, particularly in federated learning, presents unique technical challenges. Federated learning models tend to be task-specific, aligned with specific client objectives, and require a level of synchronization that relies on central servers. This need for synchronization, which is essential to the process, introduces complexity in managing the varying paces of the participating nodes. In addition, the architecture of these models typically requires uniformity across the network, impacting the flexibility of the system. The iterative nature of these methods, essential for refining the models, also leads to increased communication requirements between clients. Therefore, further improvement in the training of AI models is desired.SUMMARY
[0005] According to one aspect of this disclosure, there is provided a method, which comprises, at a present node of a network infrastructure: receiving a first dataset generating model from a first upstream node of the network infrastructure, the first dataset generating model being trained using a first authentic dataset at the first upstream node; generating a first synthetic dataset using the received first dataset generating model, the first synthetic dataset simulating the first authentic dataset; combining the first synthetic dataset with an expansion dataset into a combined dataset; and training a dataset generating model using the combined dataset to obtain an updated dataset generating model capable of generating an updated synthetic dataset simulating both the first authentic dataset and the expansion dataset.
[0006] In some embodiments, the dataset generating model may be the first dataset generating model or a generic generative AI model.
[0007] In some embodiments, the method may further comprise: outputting the updated dataset generating model to a first downstream node of the network infrastructure.
[0008] In some embodiments, the first authentic dataset may be collected at the first upstream node, and the expansion dataset may be a local authentic dataset collected at the present node.
[0009] In some embodiments, the method may further comprise: receiving a second dataset generating model from a second upstream node of the network infrastructure, the second dataset generating model being trained using a second authentic dataset at the second upstream node; and generating a second synthetic dataset using the received second dataset generating model, the second synthetic dataset simulating the second authentic dataset, wherein the expansion dataset is the second synthetic dataset.
[0010] In some embodiments, the method may further comprise: combining an additional training dataset into the combined dataset.
[0011] In some embodiments, the method may further comprise: outputting the updated dataset generating model back to the first upstream node and the second upstream node.
[0012] In some embodiments, the method may further comprise: generating the updated synthetic dataset with the updated dataset generating model; and training a task-specific generative AI model using the updated synthetic dataset.
[0013] According to another aspect of this disclosure, there is provided one or more circuits for performing actions, at a present node of a network infrastructure, comprising: receiving a first dataset generating model from a first upstream node of the network infrastructure, the first dataset generating model being trained using a first authentic dataset at the first upstream node; generating a first synthetic dataset using the received first dataset generating model, the first synthetic dataset simulating the first authentic dataset; combining the first synthetic dataset with an expansion dataset into a combined dataset; and training a dataset generating model using the combined dataset to obtain an updated dataset generating model capable of generating an updated synthetic dataset simulating both the first authentic dataset and the expansion dataset.
[0014] In some embodiments, the dataset generating model may be the first dataset generating model or a generic generative AI model.
[0015] In some embodiments, the first authentic dataset may be collected at the first upstream node, and the expansion dataset may be a local authentic dataset collected at the present node.
[0016] In some embodiments, the actions may further comprise: receiving a second dataset generating model from a second upstream node of the network infrastructure, the second dataset generating model being trained using a second authentic dataset at the second upstream node; and generating a second synthetic dataset using the received second dataset generating model, the second synthetic dataset simulating the second authentic dataset, wherein the expansion dataset is the second synthetic dataset.
[0017] In some embodiments, the action may further comprise: combining an additional training dataset into the combined dataset.
[0018] In some embodiments, the actions may further comprise: generating the updated synthetic dataset with the updated dataset generating model; and training a task-specific generative AI model using the updated synthetic dataset.
[0019] According to yet another aspect of this disclosure, there is provided one or more non-transitory computer-readable storage devices comprising computer-executable instructions, wherein the instructions, when executed, cause one or more circuits to perform actions comprising: receiving a first dataset generating model from a first upstream node of the network infrastructure, the first dataset generating model being trained using a first authentic dataset at the first upstream node; generating a first synthetic dataset using the received first dataset generating model, the first synthetic dataset simulating the first authentic dataset; combining the first synthetic dataset with an expansion dataset into a combined dataset; and training a dataset generating model using the combined dataset to obtain an updated dataset generating model capable of generating an updated synthetic dataset simulating both the first authentic dataset and the expansion dataset.
[0020] In some embodiments, the dataset generating model may be the first dataset generating model or a generic generative AI model.
[0021] In some embodiments, the first authentic dataset may be collected at the first upstream node, and the expansion dataset may be a local authentic dataset collected at the present node.
[0022] In some embodiments, the actions may further comprise: receiving a second dataset generating model from a second upstream node of the network infrastructure, the second dataset generating model being trained using a second authentic dataset at the second upstream node; and generating a second synthetic dataset using the received second dataset generating model, the second synthetic dataset simulating the second authentic dataset, wherein the expansion dataset is the second synthetic dataset.
[0023] In some embodiments, the actions may further comprise: combining an additional training dataset into the combined dataset.
[0024] In some embodiments, the actions may further comprise: generating the updated synthetic dataset with the updated dataset generating model; and training a task-specific generative AI model using the updated synthetic dataset.
[0025] The above-described method, one or more circuits, and one or more non-transitory computer-readable storage devices of this disclosure provide notable advantages. First, the embodiments described herein are task-agnostic, allowing any user, regardless of their specific task, to participate in FL. This universality dramatically expands the scope and applicability of FL, enabling diverse and widespread collaboration. Second, the embodiments described herein facilitate knowledge transfer among users in a one-shot manner. This efficiency in knowledge dissemination not only streamlines the learning process, but also significantly reduces the time and computational resources required. Third, the embodiments described herein are robust to data heterogeneity, ensuring consistent performance and reliability even in scenarios where data varies significantly across users. In addition, the embodiments described herein significantly reduce communication costs, addressing one of the key challenges of traditional FL methods. This reduction in communication overhead not only makes the process more efficient, but also more feasible in bandwidth-constrained environments. In addition, the embodiments described herein can be fully decentralized, providing a flexible and scalable solution that does not rely on central servers, thereby enhancing privacy and security. Further, the embodiments described herein support heterogeneous resource requirements, making them adaptable to different hardware and computational constraints. This adaptability ensures that the system can be deployed in a wide range of environments, from high-performance servers to edge devices, making it a versatile solution in the field of AI.BRIEF DESCRIPTION OF THE DRAWINGS
[0026] For a more complete understanding of the disclosure, reference is made to the following description and accompanying drawings, in which:
[0027] FIG. 1 is a simplified schematic diagram of an artificial intelligence (AI) system according to some embodiments of this disclosure;
[0028] FIG. 2 is a schematic diagram showing the hardware structure of the infrastructure layer of the AI system shown in FIG. 1, according to some embodiments of this disclosure;
[0029] FIG. 3 is a schematic diagram showing the hardware structure of a chip of the AI system shown in FIG. 1, according to some embodiments of this disclosure;
[0030] FIG. 4 is a schematic diagram of an AI model in the form of a deep neural network (DNN) used in the infrastructure layer shown in FIG. 2;
[0031] FIG. 5 is an example flowchart showing a method for training a dataset generating model within a network infrastructure according to some embodiments of this disclosure.
[0032] FIG. 6 is a diagram outlining a multi-step framework that facilitates various operations within a network according to some embodiments of this disclosure.
[0033] FIG. 7 is an example flowchart showing a method for training a dataset generating model involving a centralized node within a network infrastructure according to some embodiments of this disclosure.
[0034] FIG. 8 is a diagram of a network environment where dataset generating models are trained according to some embodiments of this disclosure.
[0035] FIG. 9 is a procedural flowchart corresponding to the network environment shown in FIG. 8, illustrating the operational sequence for training dataset generating models according to some embodiments of this disclosure.
[0036] FIG. 10 is a procedural flowchart showing a procedure for receiving dataset generating models from an application perspective according to some embodiments of this disclosure.
[0037] FIG. 11 is a procedural flowchart showing a downstream task training procedure for a task-specific model according to some embodiments of this disclosure.
[0038] FIG. 12 is a procedural flowchart showing a transmitting procedure for dataset generating models within a network environment according to some embodiments of this disclosure.DETAILED DESCRIPTION
[0039] A. ARTIFICIAL INTELLIGENCE SYSTEM
[0040] AI machines and systems usually comprise one or more AI models which may be trained using a large amount of relevant data for improving the precision of their perception, inference, and decision making.
[0041] Turning now to FIG. 1, an AI system according to some embodiments of this disclosure is shown and is generally identified using reference numeral 100. The AI system 100 comprises an infrastructure layer 102 for providing hardware basis of the AI system 100, a data processing layer 104 for processing relevant data and providing various functionalities 106 as needed and / or implemented, and an application layer 108 for providing intelligent products and industrial applications.
[0042] The infrastructure layer 102 comprises necessary input components 112 such as sensors and / or other input devices for collecting input data, computational components 114 such as one or more intelligent chips, circuitries, and / or integrated chips (ICs) , and / or the like for conducting necessary computations, and a suitable infrastructure platform 116 for AI tasks.
[0043] The one or more computational components 114 may be one or more central processing units (CPUs) , one or more neural processing units (NPUs; which are processing units having specialized circuits for AI-related computations and logics) , one or more graphic processing units (GPUs) , one or more application-specific integrated circuits (ASICs) , one or more field-programmable gate arrays (FPGAs) , and / or the like, and may comprise necessary circuits for hardware acceleration.
[0044] The platform 116 may be a distributed computation framework with networking support, and may comprise cloud storage and computation, an interconnection network, and the like.
[0045] In FIG. 1, the data collected by the input components 112 are conceptually represented by the data-source block 122 which may comprise any suitable data such as sensor data (for example, data collected by Internet-of-Things (IoT) devices) , service data, perception data (for example, forces, offsets, liquid levels, temperatures, humidities, and / or the like) , and / or the like, and may be in any suitable forms such as figures, images, voice clips, video clips, text, and / or the like.
[0046] The data processing layer 104 comprises one or more programs and / or program modules 124 in the form of software, firmware, and / or hardware circuits for processing the data of the data-source block 122 for various purposes such as data training, machine learning, deep learning, searching, inference, decision making, and / or the like.
[0047] In machine learning and deep learning, symbolic and formalized intelligent information modeling, extraction, preprocessing, training, and the like may be performed on the data-source block 122.
[0048] Inference refers to a process of simulating an intelligent inference manner of a human being in a computer or an intelligent system, to perform machine thinking and resolve a problem by using formalized information based on an inference control policy. Typical functions are searching and matching.
[0049] Decision making refers to a process of making a decision after inference is performed on intelligent information. Generally, functions such as classification, sorting, and inferencing (or prediction) are provided.
[0050] With the programs and / or program modules 124, the data processing layer 104 generally provides various functionalities 106 such as translation, text analysis, computer-vision processing, voice recognition, image recognition, and / or the like.
[0051] With the functionalities 106, the AI system 100 may provide various intelligent products and industrial applications 108 in various fields, which may be packages of overall AI solutions for productizing intelligent information decisions and implementing applications. Examples of the application fields of the intelligent products and industrial applications may be intelligent manufacturing, intelligent transportation, intelligent home, intelligent healthcare, intelligent security, automated driving, safe city, intelligent terminal, and the like.
[0052] FIG. 2 is a schematic diagram showing the hardware structure of the infrastructure layer 102, according to some embodiments of this disclosure. As shown, the infrastructure layer 102 comprises a data collection device 140 for collecting training data 142 for training an AI model 148 (such as a machine-learning (ML) model, a neural network (NN) model (for example, a convolutional neural network (CNN) model) , or the like) and storing the collected training data 142 into a training database 144. Herein, the training data 142 comprises a plurality of identified, annotated, or otherwise classified data samples that may be used for training (denoted “training samples” hereinafter) and their corresponding desired results (denoted “labels” hereinafter; that is, the target or desired predictions that the AI model 148 should make from the data samples) . Herein the training samples may be any suitable data samples to be used for training the AI model 148, such as one or more annotated images, one or more annotated text samples, one or more annotated audio clips, one or more annotated video clips, one or more annotated numerical data samples, and / or the like. The desired results are ideal results expected to be obtained by processing the training samples by using the trained or optimized AI model 148’ . One or more training devices 146 (such as one or more server computers forming the so-called “computer cloud” or simply the “cloud” , and / or one or more client computing devices (also called “edge devices” ) similar to or same as the execution devices 150) train the AI model 148 using the training data 142 retrieved from the training database 144 to train the AI model 148 for use by the computation module 174 (described in more detail later) .
[0053] As those skilled in the art will appreciate, in actual applications, the training data 142 maintained in the training database 144 may not necessarily be all collected by the data collection device 140, and may be received from other devices. Moreover, the training devices 146 may not necessarily perform training completely based on the training data 142 maintained in the training database 144 to obtain the trained AI model 148’ , and may obtain training data 142 from a cloud or another place to perform model training.
[0054] The trained AI model 148’ obtained by the training devices 146 through training may be applied to various systems or devices such as an execution device 150 which may be an edge device such as a mobile phone terminal, a tablet computer, a notebook computer, an augmented reality (AR) device, a virtual reality (VR) device, a vehicle-mounted terminal, a server, or the like. The execution device 150 comprises an I / O interface 152 for receiving input data 154 from an external device 156 (such as input data provided by a user 158) and / or outputting results 160 to the external device 156. The external device 156 may also provide training data 142 to the training database 144. The execution device 150 may also use its I / O interface 152 for receiving input data 154 directly from the user 158.
[0055] The execution device 150 also comprises a processing module 172 for performing preprocessing based on the input data 154 received by the I / O interface 152. For example, in cases where the input data 154 comprises one or more images, the processing module 172 may perform image preprocessing such as image filtering, image enhancement, image smoothing, image restoration, and / or the like.
[0056] The processed data 142 is then sent to a computation module 174 which uses the trained AI model 148’ to analyze the data received from the processing module 172 for prediction. As described above, the prediction results 160 may be output to the external device 156 via the I / O interface 152. Moreover, data 154 received by the execution device 150 and the prediction results 160 generated by the execution device 150 may be stored in a data storage system 176.
[0057] In the following, the AI model to be trained and the corresponding trained AI model are identified using the same reference numeral 148 for ease of description.
[0058] FIG. 3 is a schematic diagram showing the hardware structure of a computational component 114 according to some embodiments of this disclosure. The computational component 114 may be any processor suitable for large-scale exclusive OR operation processing, for example, a convolutional NPU, a tensor processing unit (TPU) , a GPU, or the like. The computational component 114 may be a part of the execution device 150 coupled to a host CPU 202 for use as the computational module 160 under the control of the host CPU 202. Alternatively, the computational component 114 may be in the training devices 146 to complete training work thereof and output the trained AI model 148.
[0059] As shown in FIG. 3, the computational component 114 is coupled to an external memory 204 via a bus interface unit (BIU) 212 for obtaining instructions and data (such as the input data 154 and weight data) therefrom. The instructions are transferred to an instruction fetch buffer 214. The input data 154 is transferred to an input memory 216 and a unified memory 218 via a storage-unit access controller (or a direct memory access controller, DMAC) 220, and the weight data is transferred to a weight memory 222 via the DMAC 220. In these embodiments, the instruction fetch buffer 214, the input memory 216, the unified memory 218, and the weight memory 222 are on-chip memories, and the input data 154 and the weight data may be organized in matrix forms (denoted “input matrix” and “weight matrix” , respectively) .
[0060] A controller 226 obtains the instructions from the instruction fetch buffer 214 and accordingly controls an operation circuit 228 to perform multiplications and additions using the input matrix from the input memory 216 and the weight matrix from the weight memory 222.
[0061] In some implementations, the operation circuit 228 comprises a plurality of processing engines (PEs; not shown) . In some implementations, the operation circuit 228 is a two-dimensional systolic array. The operation circuit 228 may alternatively be a one-dimensional systolic array or another electronic circuit that may perform mathematical operations such as multiplication and addition. In some implementations, the operation circuit 228 is a general-purpose matrix processor.
[0062] For example, the operation circuit 228 may obtain an input matrix A (for example, a matrix representing an input image) from the input memory 216 and a weight matrix B (for example, a convolution kernel) from the weight memory 222, buffer the weight matrix B on each PE of the operation circuit 228, and then perform a matrix operation on the input matrix A and the weight matrix B. The partial or final computation result obtained by the operation circuit 228 is stored into an accumulator 230.
[0063] If required, the output of the operation circuit 228 stored in the accumulator 230 may be further processed by a vector calculation unit 232 such as vector multiplication, vector addition, an exponential operation, a logarithmic operation, size comparison, and / or the like. The vector calculation unit 232 may comprise a plurality of operation processing engines, and is mainly used for calculation at a non-convolutional layer or a fully connected layer (FC) of the CNN, and may specifically perform calculation in pooling, normalization, and the like. For example, the vector calculation unit 232 may apply a non-linear function to the output of the operation circuit 228, for example a vector of an accumulated value, to generate an active value. In some implementations, the vector calculation unit 232 generates a normalized value, a combined value, or both a normalized value and a combined value.
[0064] In some implementations, the vector calculation unit 232 stores a processed vector into the unified memory 218. In some implementations, the vector processed by the vector calculation unit 232 may be stored into the input memory 216 and then used as an active input of the operation circuit 228, for example, for use at a subsequent layer in the CNN.
[0065] The data output from the operation circuit 228 and / or the vector calculation unit 232 may be transferred to the external memory 204.
[0066] FIG. 4 is a schematic diagram of the AI model 148 in the form of a deep neural network (DNN; which is a sophisticated form of an artificial neural network (ANN) . As shown, the DNN 148 comprises an input layer 302, a plurality of cascaded hidden layers 304, and an output layer 306. The trained AI model 148 may have a set of parameters optimized through the AI-model training.
[0067] The input layer 302 comprises a plurality of input nodes 312 for receiving input data and outputting the received data to the computation nodes 314 of the subsequent hidden layer 304. Each hidden layer 304 comprises a plurality of computation nodes 314. Each computation node 304 weights and combines the outputs of the input or computation nodes of the previous layer (that is, the input nodes 312 of the input layer 302 or the computation nodes 314 of the previous hidden layer 304, and each arrow representing a data transfer with a weight) . The output layer 306 also comprises one or more output node 316, each of which combines the outputs of the computation nodes 314 of the last hidden layer 304 for generating the outputs 356.
[0068] As those skilled in the art will appreciate, the AI model such as the DNN 148 shown in FIG. 4 generally requires training for optimization. For example, a training device 146 (see FIG. 2) may provide training data 142 (which comprises a plurality of training samples with corresponding desired results) to the input nodes 312 to run through the AI model 148 and generate outputs from the output nodes 316. By comparing the outputs obtained from the output nodes 316 with the desired results in the training data 142, a lost function may be established and the parameters of the AI model 148, such as the weights thereof, may be optimized by minimizing the lost function.
[0069] B. TRAINING OF DATASET GENERATING MODELS FOR GENERATING SYNTHETIC DATASETS
[0070] FIG. 5 illustrates an example flowchart for training a dataset generation model within a network infrastructure conducted in the absence of authentic data, according to an embodiment of the present disclosure. The network infrastructure illustrated includes three nodes: a first node 410, a second node 420, and a third node 430.
[0071] Within the scope of this disclosure, each node is defined as a computational entity that may represent an individual user device or a collection of devices collaborating to execute the operations as outlined. These nodes are endowed with networking capabilities, which enable them to establish and maintain communication with other nodes in the network. Such communication may occur over various channels, including but not limited to wireless radio frequency technologies, wired Ethernet connections, or through the use of emerging communication protocols that may offer enhanced bandwidth, reduced latency, or increased security.
[0072] In the context of this example, the first, second, and third nodes 410, 420, 430 are mobile devices with data exchange functionalities. These devices are not only equipped to transmit and receive data packets amongst themselves utilizing direct peer-to-peer connections or indirect routing mechanisms but are also capable of interfacing with the broader network infrastructure. This includes the ability to interact with fixed network stations, such as base stations or access points, which facilitate the nodes’ integration into the network. The nodes’ mobile nature, coupled with their robust communication capabilities, enables a versatile and resilient network topology that is adaptable to various user scenarios and network conditions.
[0073] Furthermore, these nodes are imbued with data processing and handling capabilities, allowing for a range of advanced operations including data packet encapsulation, prioritization of data traffic, data storage, and execution of security protocols. The nodes can negotiate data transmission parameters based on network policies, user preferences, or adaptive algorithms that seek to optimize the network's performance and resource utilization.
[0074] In the illustrated embodiment, the first node 410 acquires a first authentic dataset 412, for example, obtained from a user via a first mobile device 411. This first mobile device 411 may itself constitute the first node 410, or may operate independently. The first authentic dataset 412 comprises real user data, often of a private and sensitive nature. The types of data may range from browsing history, purchase records, personal health records and financial information to private communications and personalized user interactions. Such data is typically secured and not intended for dissemination beyond the boundaries of the first node.
[0075] The acquired first authentic dataset 412 in this embodiment is used to train a generic AI model or a pre-trained AI model to derive a first dataset generating model 413. This training process leverages the intrinsic data patterns and characteristics to enable the first dataset generating model 413, once trained, to generate new, synthetic datasets that simulate the statistical characteristics of the original data without compromising privacy. These generated datasets are constructed to mirror the statistical profile of the original dataset, and this facsimile is achieved without overstepping the boundaries of user privacy. Thus, the resulting output of the first dataset generating model 413 is a statistical facsimile of the first authentic dataset, but any sensitive or private user data attributes are substituted with artificially generated non-sensitive equivalents. In this way, it is ensured that the privacy of the data subjects is maintained, while still providing valuable data for further analysis and model training efforts. While the current example uses a generic AI model as the basis for training, the scope of this embodiment is not limited to such. Alternatives include the use of a pre-trained model, which may be obtained from another node within the network or from an external third party, providing a foundation enriched with a broader or different knowledge base.
[0076] As a result of the training, the first dataset generating model 413 acquires the capability to produce a synthetic dataset that closely mimics the statistical profile of the original, first authentic dataset 412, as previously described. This synthetic dataset can then be employed in the training of a first task-specific model 414. For example, such a model may take the form of a language generative AI model, designed to synthesize human-like textual content, or an image generative AI model, adept at crafting visual content endowed with particular attributes. The first task-specific model 414 is designed to cater to distinct use cases, offering specialized functionalities that may include language translation, content curation, or the provision of personalized user experiences. By leveraging synthetic data for training, the first task-specific model 414 may be trained without the use of any genuine user data, thereby significantly enhancing user privacy protections.
[0077] Training of either the first dataset generating model 413 or the first task-specific model 414 may be performed at the first node 410 using available computational resources, which may include processing power, memory, and data storage capacities. This training process is orchestrated to ensure efficiency and effectiveness, taking advantage of the node's capabilities to refine the models to a high degree of accuracy while protecting the user's private data. It is worth noting that the training methods used can be adapted based on the availability of resources, allowing for flexibility in computational allocation and energy consumption, thus ensuring an optimized training process tailored to the specific capacities of the first node 410.
[0078] After the training of the first dataset generating model 413 at the first node 410, the model is then transmitted to the second node 420. The second node 420, which may be a second user device 421 or an aggregation of devices, receives this trained model. With the first dataset generating model 413 now in place, the second node 420 uses it to generate a first synthetic dataset 412', which closely mimics the statistical profile of the original, first authentic dataset 412, inheriting its training from the first node 410. On the other hand, the second node 420 owns or receives a second authentic dataset 422 as an expansion dataset for expanding the volume of the dataset for training the dataset generating model in the second node 420. Like the first authentic dataset 412, the second authentic dataset 422 typically contains sensitive or private data, relating to a second user associated with the second user device 421.
[0079] A combination operation is then performed, in which the first synthetic dataset 412'is combined or fused with the second authentic dataset 422 at the second node 420. This combination may take various forms, including a simple amalgamation or a weighted amalgamation, among other sophisticated data integration techniques. The result of this combination is a combined dataset that not only preserves the integrity of each dataset, but also enriches the informational breadth. This combined dataset is then used to refine the previously received first dataset generating model 413 or to initiate the training of a new generic AI model. Upon completion of this process, the model is evolved into a second dataset generating model 423. This evolved model is now capable of generating a synthetic dataset that reflects the nuances of both the first authentic dataset 412 and the second authentic dataset 422, due to the cumulative training across the first and second nodes 410 and 420.
[0080] It should be appreciated that, in another embodiment, the second node 420 may receive another dataset generating model in place of the second authentic dataset 422, so that this distinct dataset generating model is capable of generating another synthetic dataset as the expansion dataset to be combined with the first synthetic dataset 412'.
[0081] As a result of the operations performed by the second node 420, the newly trained second dataset generating model 423 has the ability to produce an expanded synthetic dataset. This dataset is larger in volume and richer in diversity, effectively simulating the data from both the first and second authentic datasets. The expanded synthetic dataset may thus be used to train a second task-specific model 424 (like the first task-specific model 414) . The benefits of a continuous training of the dataset generating model include improving the representativeness of the data while preserving the privacy of the underlying authentic information. It also facilitates the scalability of data generation processes and the reliability of subsequent AI-driven analyses or applications.
[0082] After the training of the second dataset generating model 423 at the second node 420, the model is then transmitted to the third node 430. The third node 430, which may be a third user device 431 or an aggregation of devices, receives this trained model. With the second dataset generating model 423 now in place, the third node 430 uses it to generate an altered first synthetic dataset 412” and a second synthetic dataset 422', which are a statistical twin of the first authentic dataset 412 and the second authentic dataset 422, respectively, inheriting its training from the first node 410 and the second node 420. On the other hand, the third node 430 owns or receives a third authentic dataset 432, like the first authentic dataset 412 or the second authentic dataset 422, typically containing sensitive or private data, relating to a second user associated with the third user device 431.
[0083] A combination operation is then performed, in which the altered first synthetic dataset 412” is combined or fused with the third authentic dataset 432 at the third node 430. This combination may take various forms, including a simple amalgamation or a weighted amalgamation, among other sophisticated data integration techniques. The result of this combination is a combined dataset that not only preserves the integrity of each dataset, but also enriches the informational breadth. This combined dataset is then used to refine the previously received second dataset generating model 423 or to initiate the training of a new generic AI model. Upon completion of this process, the model is evolved into a third dataset generating model 433. This evolved model is now capable of generating a synthetic dataset that reflects the nuances of all of the first authentic dataset 412, the second authentic dataset 422, and the third authentic dataset 432, due to the cumulative training across the first, second, and third nodes 410, 420, and 430.
[0084] It should be appreciated that, in another embodiment, the third node 430 may receive another dataset generating model in place of the third authentic dataset 432, so that this distinct dataset generating model is capable of generating another synthetic dataset to be combined with the altered first synthetic dataset 412” and the second synthetic dataset 422'.
[0085] As a result of the operations performed by the third node 430, the newly trained third dataset generating model 433 has the ability to produce a further expanded synthetic dataset. This dataset is larger in volume and richer in diversity, effectively simulating the data from all of the first, second, and third authentic datasets. The further expanded synthetic dataset may thus be used to train a third task-specific model 434 (like the first task-specific model 414) . The benefits of a continuous training of the dataset generating model include improving the representativeness of the data while preserving the privacy of the underlying authentic information. It also facilitates the scalability of data generation processes and the reliability of subsequent AI-driven analyses or applications.
[0086] It should be appreciated that, in certain implementations, the altered first synthetic dataset 412” synthesized at the third node 430 may be identical to the first synthetic dataset 412'synthesized at the second node 420. Nevertheless, differences in the altered first synthetic dataset 412” may arise as a result of differences in available resources between different nodes within the network infrastructure. Such variations may manifest themselves in the dimensional attributes of the dataset, including but not limited to the size, complexity, or variety of the data. Other differences may include the granularity of the data, the level of noise within the dataset, or other data quality metrics such as completeness and consistency. These variations reflect the resource-driven adaptability of the system, allowing each node to tailor the generation of synthetic datasets according to its operational capabilities and constraints.
[0087] In addition, the variation in resource allocation-which can include computational power, memory capacity, storage space, and processing speed-may result in the generation of synthetic datasets with different levels of fidelity or detail. For example, a node with greater processing power may produce a synthetic dataset with higher resolution or finer features, while a node with limited capabilities may produce a dataset that prioritizes broader patterns and generalizations over intricate details.
[0088] In some embodiments, the generation of synthetic datasets according to the resources available at a particular node may be determined by a predefined algorithm. This algorithm, which may be intrinsic to the functionality of the nodes in the network, can be sent to, received from, and subsequently stored by any node. It serves as a blueprint, outlining the steps for adjusting data characteristics such as volume, complexity, and variety to match the computational and storage capabilities of the particular node.
[0089] The algorithm may incorporate various decision-making heuristics, machine learning models, or deterministic rules that systematically reduce or enhance the attributes of the dataset. For example, the algorithm may employ data compression techniques to accommodate nodes with limited storage or increase data granularity when excess processing power is available. Further, the algorithm may adjust the fidelity of the synthetic dataset, trading off data detail against processing efficiency, as dictated by the node's current operating parameters.
[0090] Additionally or alternatively, the algorithm may be designed to be dynamic, allowing it to respond to changes in the node's resources in real time. It may prioritize more essential data features or downscale less essential elements, ensuring that the most valuable insights are retained even when resources are limited. The adaptability of the algorithm also means that it may receive updates or modifications that can be deployed network-wide to improve performance or introduce new functionality.
[0091] In all cases, the purpose of the predefined algorithm is to ensure that the synthetic datasets remain a robust and reliable foundation for model training and AI tasks, regardless of the varying computational landscapes they encounter. This level of responsiveness and flexibility underscores the innovative nature of the system, allowing it to harness the collective power of the network to optimize data generation and processing efforts.
[0092] While the embodiment depicted in FIG. 5 illustrates a sequential flow of operations, it should be understood that the present disclosure is not limited to such a linear progression. For example, the architecture may include a plurality of nodes configured in a looped network topology, thereby allowing the data set generating model to be trained in a cyclical sequence. In this arrangement, training of the model may be performed in a tandem order, with the training looping back to the first node after reaching the last node in the sequence. This cyclical training methodology promotes continuous improvement of the model as it iteratively acquires and integrates knowledge from the datasets of all participating nodes, progressively increasing its simulation capabilities and accuracy with each pass through the network loop.
[0093] Alternatively, nodes may be connected in a non-sequential configuration, following a predetermined order that is not strictly linear. Within this framework, certain nodes may be designated for repeated training, possibly because of their superior computational capabilities or strategic importance within the network. For example, a node with significant computational resources or one that is central to a cluster of nodes may be selected to perform additional training iterations. Such a node may act as a hub, aggregating insights and updates from other nodes before distributing an improved model back to the network. This selective iteration ensures that certain nodes play a more prominent role in model development, contributing to a tailored and resource-efficient training methodology.
[0094] In addition, the network configuration may be dynamically adjusted, with nodes added, removed, or reassigned training responsibilities based on real-time assessments of network conditions and resource availability. This dynamic adaptation allows for a flexible and scalable system that adapts to the changing needs of the network, ensuring optimal resource utilization and consistent model improvement across the infrastructure.
[0095] Referring now to FIG. 6, the framework described in the present disclosure operates on a number of operations executable by any given node "i" within the network. This framework enables participating nodes to use AI models, such as a generative model, to facilitate the transfer of knowledge across the network according to the various embodiments described herein.
[0096] First, the operation of transmission involves the ability of node "i" (530) to either send or receive a dataset generating model, herein referred to as a "generator" , from or to any other node (such as 510, 520, or 540) within the network. This operation also includes the routing of a trained generator to any other node (such as 510, 520, or 540) , ensuring the continuous flow of knowledge throughout the network.
[0097] Second, in the data generation operation, node "i" is equipped to generate synthetic dataset or data samples. This is achieved by using the received generators or by using a locally stored generator. The synthetic dataset thus generated serves as a proxy for real data, facilitating the augmentation of the node's data resources without compromising privacy.
[0098] The third operation, training generator, allows node "i" to refine or create a new local generator. This generator is created by integrating the synthetic data with the node's local data, giving the generator a richer understanding of different data patterns. Upon completion of this training, the improved local generator may be transmitted to one or more nodes, denoted by {j: i ≠ j} , thereby propagating the acquired knowledge throughout the network.
[0099] Subsequently, a save operation allows node "i" to store the newly created local generator. This storage serves not only as a repository for the node's immediate use, but also as a reservoir of knowledge that may be tapped or shared with other nodes in the future. Then, the combined local and synthetic data is used to train a downstream model.
[0100] FIG. 7 illustrates an example flowchart for training a dataset generation model with a centralized node in a network infrastructure, according to another embodiment of the present disclosure. The network infrastructure illustrated includes four nodes: a first node 610, a second node 620, a third node 630, and a centralized node 640. Each of the first, second, and third nodes may represent either a single user device or a collection of devices that collectively perform the described operations, all of which are equipped with networking capabilities that allow them to communicate with other nodes via wireless connections, Ethernet, or other communication technologies.
[0101] In the illustrated embodiment, the first node 610 acquires a first authentic dataset 612, either from another source or stored locally at the first node 610. The first authentic dataset 612 comprises real user data, often of a private and sensitive nature. For example, the types of data may range from browsing history, purchase records, personal health records and financial information to private communications and personalized user interactions. Such data is typically secured and not intended for dissemination beyond the boundaries of the first node.
[0102] The acquired first authentic dataset 612 in this embodiment is used to train a generic AI model or a pre-trained AI model to derive a first dataset generating model 613. This training process leverages the intrinsic data patterns and characteristics to enable the first dataset generating model 613, once trained, to generate new, synthetic datasets that simulate the statistical characteristics of the original data without compromising privacy. These generated datasets are constructed to mirror the statistical profile of the original dataset, and this facsimile is achieved without overstepping the boundaries of user privacy. Thus, the resulting output of the first dataset generating model 613 is a statistical facsimile of the first authentic dataset 612, but any sensitive or private user data attributes are substituted with artificially generated non-sensitive equivalents. In this way, it is ensured that the privacy of the data subjects is maintained, while still providing valuable data for further analysis and model training efforts. While the current example uses a generic AI model as the basis for training, the scope of this embodiment is not limited to such. Alternatives include the use of a pre-trained model, which may be obtained from another node within the network or from an external third party, providing a foundation enriched with a broader or different knowledge base. It should be understood that the first node 610 may store the first authentic dataset 612 that is collected from multiple users or user devices, thus serving as a primarily centralized hub.
[0103] Similarly, the second node 620 acquires a second authentic dataset 622, either from another source or stored locally at the second node 620; the third node 630 acquires a third authentic dataset 632, either from another source or stored locally at the third node 630. The acquired second and third authentic datasets 622, 632 are then used to train a generic AI model or a pre-trained AI model to derive second and third dataset generating models 623, 633, respectively.
[0104] Training for the dataset generating models-first model 613, second model 623, and third model 633-is executed at the corresponding first node 610, second node 620, and third node 630, respectively. This training utilizes the computational resources available at each node, which encompass processing power, memory, and data storage capacities. Nonetheless, it should be noted that the resources at one or more of these nodes-610, 620, 630-may be constrained. In such cases, additional training with more substantial resources and a more comprehensive dataset could be advantageous for enhancing the model's performance.
[0105] Upon completion of the training processes, the first dataset generating model 613, the second dataset generating model 623, and the third dataset generating model 633 are transmitted from their respective originating nodes-first node 610, second node 620, and third node 630-to the centralized node 640. Equipped with potentially more robust computational resources, the centralized node 640 acquires these trained models. Once in possession of the first, second, and third dataset generating models 613, 623, 633, the centralized node 640 proceeds to generate corresponding synthetic datasets: 612', 622', and 632'. These datasets are crafted to statistically emulate the authentic datasets from each originating node-first authentic dataset 612, second authentic dataset 622, and third authentic dataset 632-thus allowing for the creation of precise synthetic parallels without direct use of the original authentic data. In this example, either of the synthetic datasets 612', 622', 632'may be regarded as the expansion dataset for another one of the synthetic datasets 612', 622', 632'.
[0106] At the centralized node 640, an auxiliary training dataset 642 can be incorporated into the model refinement process. This auxiliary training dataset 642, which may be either stored locally on the centralized node 640 or sourced externally, introduces additional data elements that may be authentic or synthetic. The purpose of integrating this dataset is to leverage the superior computational capabilities of the centralized node 640. This capability allows for a considerable expansion of the dataset’s volume, or an adjustment thereof, to meet specific training requirements and enhance the model's learning potential. It should be appreciated that the inclusion of the auxiliary training dataset 642 is optional, providing flexibility to adapt the training process to the availability of data and computational resources.
[0107] A combination operation is then performed, in which the first synthetic dataset 612', the second synthetic dataset 622', and the third synthetic dataset 632' are combined with or without the auxiliary training dataset 642. This combination may take various forms, including a simple amalgamation or a weighted amalgamation, among other sophisticated data integration techniques. The result of this combination is a combined dataset 652 that not only preserves the integrity of each dataset, but also enriches the informational breadth. This combined dataset 652 is then used to refine any of the previously received first, second, and third dataset generating models 613, 623, 633 or to initiate the training of a new generic AI model. Upon completion of this process, the model is evolved into an enhanced dataset generating model 653. This enhanced model is now capable of generating a synthetic dataset that reflects the nuances of all of the first, second, and third authentic dataset 613, 623, 633, and optionally the auxiliary training dataset 642.
[0108] Then, the enhanced dataset generating model 653 may be transmitted back to each of the first, second, and third nodes 610, 620, 630, and be used to generate a first enhanced synthetic dataset 652', a second enhanced synthetic dataset 652”, and a third enhanced synthetic dataset 652”'. It should be appreciated that, in certain implementations, the first enhanced synthetic dataset 652', the second enhanced synthetic dataset 652”, and the third enhanced synthetic dataset 652”' synthesized at the three nodes may be identical to each other. Nevertheless, differences may arise as a result of differences in available resources between different nodes within the network infrastructure. Such variations may manifest themselves in the dimensional attributes of the dataset, including but not limited to the size, complexity, or variety of the data. Other differences may include the granularity of the data, the level of noise within the dataset, or other data quality metrics such as completeness and consistency. These variations reflect the resource-driven adaptability of the system, allowing each node to tailor the generation of synthetic datasets according to its operational capabilities and constraints.
[0109] In addition, the variation in resource allocation-which can include computational power, memory capacity, storage space, and processing speed-may result in the generation of synthetic datasets with different levels of fidelity or detail. For example, a node with greater processing power may produce a synthetic dataset with higher resolution or finer features, while a node with limited capabilities may produce a dataset that prioritizes broader patterns and generalizations over intricate details.
[0110] In some embodiments, the generation of synthetic datasets according to the resources available at a particular node (including the centralized node 640) may be determined by a predefined algorithm. This algorithm, which may be intrinsic to the functionality of the nodes in the network, can be sent to, received from, and subsequently stored by any node. It serves as a blueprint, outlining the steps for adjusting data characteristics such as volume, complexity, and variety to match the computational and storage capabilities of the particular node.
[0111] The algorithm may incorporate various decision-making heuristics, machine learning models, or deterministic rules that systematically reduce or enhance the attributes of the dataset. For example, the algorithm may employ data compression techniques to accommodate nodes with limited storage or increase data granularity when excess processing power is available. Further, the algorithm may adjust the fidelity of the synthetic dataset, trading off data detail against processing efficiency, as dictated by the node's current operating parameters.
[0112] Additionally or alternatively, the algorithm may be designed to be dynamic, allowing it to respond to changes in the node's resources in real time. It may prioritize more essential data features or downscale less essential elements, ensuring that the most valuable insights are retained even when resources are limited. The adaptability of the algorithm also means that it may receive updates or modifications that can be deployed network-wide to improve performance or introduce new functionality.
[0113] In all cases, the purpose of the predefined algorithm is to ensure that the synthetic datasets remain a robust and reliable foundation for model training and AI tasks, regardless of the varying computational landscapes they encounter. This level of responsiveness and flexibility underscores the innovative nature of the system, allowing it to harness the collective power of the network to optimize data generation and processing efforts.
[0114] Moreover, the first enhanced synthetic dataset 652', the second enhanced synthetic dataset 652”, and the third enhanced synthetic dataset 652”' may be respectively combined with the first authentic dataset 612, the second authentic dataset 622, and the third authentic dataset 632. These resultant composite datasets provide a rich foundation for further training of the first, second, and third dataset generating models 613, 623, 633. Through this further training phase, refined versions of the models-namely, the first updated dataset generating model 613', the second updated dataset generating model 623', and the third updated dataset generating model 633'-are produced, each enhanced by the depth and diversity of the combined datasets.
[0115] It should be understood that the embodiment depicted with respect to FIG. 7, which illustrates a configuration of three nodes in conjunction with a single centralized node, is exemplary and not restrictive. The disclosed framework is scalable and can accommodate any number of nodes. For instance, the network may consist solely of a singular node that operates in concert with the centralized node, utilizing an auxiliary training dataset to enrich the combined dataset, thereby broadening the scope of the data (as delineated in the embodiment corresponding to FIG. 5) . Alternatively, the network could be configured with two nodes, or it may extend beyond three nodes, each variation capable of interfacing with the centralized node.
[0116] The use of the centralized node, as detailed in the embodiment corresponding to FIG. 7, serves a dual purpose. First, the centralized node is typically equipped with more substantial resources as compared to individual nodes, which facilitates a more refined and potentially superior training outcome for the dataset-generating models. Second, certain network environments may preclude direct communication between nodes; in such scenarios, the dataset generating models will benefit from training within a centralized framework, as exemplified by this embodiment.
[0117] FIG. 8 illustrates a 5G or 6G network environment in which the training of dataset generating models may be adopted. In this scenario, the network 700 includes several elements that collectively enable the handling and processing of data. A control plane is primarily governed by a network controller 710, which is a central entity that orchestrates the flow and management of network operations. The data plane includes one or more Data Processing Functions (DPF) 720 that facilitate data handling and execution of network tasks.
[0118] Within this environment, two types of Processing Service Functions (PSFs) are identified: Type-1 PSF 730 and Type-2 PSF 740. The Type-1 PSF 730 may represent a node described in respect to other embodiments and be a PSF equipped with an AI model. The Type-1 PSF 730, as the first, second, or third node 610, 620, 630 as discussed above in respect of FIG. 7, is capable of tasks such as generating synthetic datasets from authentic datasets using dataset generating models, and may represent a wireless terminal such as a User Equipment (UE) . The goal for the Type-1 PSF 730 mainly includes training AI models using private datasets collectively and mitigating the negative impacts of data heterogeneity within these private datasets (such as by using generative AI models) . Meanwhile, the Type-2 PSF 740 may represent a server, such as the centralized node 640 as described in respect of FIG. 7, that does not inherently possess an AI model, but assists in the AI model training process. For example, the Type-2 PSF 740 may receive AI models from the Type-1 PSF (s) 730, update the models as an enhanced model, and send the enhanced model back to the Type-1 PSF (s) 730.
[0119] A Processing Service Controller (PSC) 750 is also depicted, which is tasked with managing, controlling, and influencing the behavior of both Type-1 and Type-2 PSFs 730, 740. This controller ensures that PSFs 730, 740 operate within the parameters and goals of the network, and effectively trains and deploys AI models as needed.
[0120] The terminologies used in the context of FIG. 8 may be extended to both 5G and 6G network architectures. In the 6G X-centric architecture, the network controller 710 may play the role of a Multi-access Connectivity Function (MCF) responsible for managing the multiple connectivity options inherent in a 6G environment. In this context, the PSC 750 may perform functions similar to those of the NET4AI module, which is expected to be a central hub for network intelligence and AI-driven functions. The Type-1 PSF 730 may operate within the 6G framework as a User Equipment (UE) , Application Server (AS) or Network Function (NF) , performing tasks such as data processing and AI model interactions directly at the edge of the network. Conversely, the Type-2 PSF 740 can act as a PSF within the NET4AI module, supporting the training and operation of AI models without necessarily hosting the AI capabilities themselves.
[0121] In the 5G context, the network controller 710 may play the role of a Session Management Function (SMF) , which orchestrates the establishment and maintenance of network sessions. The PSC 750 may correspond to an Application Function (AF) , which plays a role in application-specific traffic routing and policy enforcement. In addition, the Type-1 PSF 730 may typically represent the UE, which are end-user devices such as smartphones, tablets, or IoT devices that connect to the network. The Type-2 PSF 740 may correspond to an AS that manages the application logic and services.
[0122] The interplay between these elements facilitates the operation of the network. The network controller (MCF or SMF) is the overarching entity that provides seamless connectivity and coordination between the various PSFs and the broader network. The PSC or AF acts as an intermediary that facilitates communication, control, and management between the core of the network and the processing service functions that perform or support the specific AI-driven tasks. Through this structured interaction, the network may efficiently meet the requirements of both current 5G and future 6G ecosystems.
[0123] It should be appreciated that while FIG. 8 shows a 5G or 6G network environment, the principles and structures described herein may be adapted for use in other network environments, such as LTE, with appropriate modifications. This versatility underscores the applicability of the disclosure to various evolving network standards, ensuring its relevance and utility in both current and future telecommunications landscapes.
[0124] FIG. 9 illustrates a procedural flowchart for a 5G or 6G network environment depicted by FIG. 8, detailing the usage of information elements and the interaction between the Type-1 and Type-2 PSFs 730, 740 and the PSC 750. This flowchart exemplifies the dynamic process by which dataset generating models are trained, validated, and utilized across a network including various PSFs and the central controller 710, following the methodologies discussed in the preceding text. It should be understood that the procedures depicted herein represent just one embodiment, and the scope of this disclosure is not confined to these specific steps alone.
[0125] In Step 801, the PSC 750 orchestrates the selection of Type-1 and Type-2 PSFs 731, 732, 740 by employing a set of criteria that ensure optimal alignment with network demands. These criteria may encompass location-based proximity, which facilitates efficient data transmission and interaction by choosing nodes close to the data source. Furthermore, task relevance may be considered, wherein PSFs are selected for their specific data processing capabilities that align with the required tasks. Additionally, data similarity may be considered, with clustering based on the likeness of generative models' data features, ensuring that PSFs with similar data or feature sets can work in tandem. These selection methods initiate a strategic routing procedure that not only enhances the network's efficiency but also its overall data processing and AI model training efficacy.
[0126] Proceeding to Step 802 within the embodiment, the PSC 750 delineates configurations for the selected Type-2 PSF 740. This step involves defining a set of parameters that dictate how the Type-2 PSF 740 will interact with and process data. These parameters establish a framework for subsequent operations and include identifying compatible Type-1 PSFs for potential communication with the Type-2 PSF 740, related to Step 805 of this embodiment.
[0127] The aggregation configuration, related to Step 806 of this embodiment, encompasses a variety of factors. It may specify the task for which the data is being prepared, the characteristics of the combined dataset generating models such as type, size, and encoding, and the number of data samples required for generation. It may also decide if the Type-2 PSF 740 utilizes additional data, which may be either independent of (only using synthetic data) or dependent upon existing datasets. This could entail leveraging data-free techniques or relying on data-dependent methods. The configuration may also take into account the integration of previously trained dataset generating models (for example, if the Type-2 PSF uses previously trained dataset generating models to create additional data and thus update newly received models or train new models) , the allocation of time and resources for training sessions, and the protocols for the training procedure itself.
[0128] Furthermore, the evaluation configuration, related to Step 806 of this embodiment, includes establishing metrics to verify the quality of the received dataset generating models (for example, by distances between authentic data and synthetic data, fidelity, distribution shift, etc. ) as well as the resultant combined models (for example, performance on a downstream task that does not require real data, etc. ) . It may define target performance benchmarks that the models should meet to be deemed successful (for example, there are predefined values for the metrics, and if they are not met, the updated model may be discarded. ) .
[0129] Error handling protocols, related to Steps 805 and 807 of this embodiment, are put in place to address a number of potential issues. These protocols prepare the system to identify and rectify transmission and training complications, as well as to troubleshoot any concerns arising from the received models.
[0130] In Step 803 of the embodiment, the PSC 750 undertakes the configuration of a first Type-1 PSF 731 and a second Type-1 PSF 732. This step involves the establishment of operational parameters that guide the training, evaluation, and communication processes of the dataset generating models. For identification purposes related to Step 805 of this embodiment, the PSC 750 may assign identifiers (IDs) to each dataset generating model, facilitating delineation and tracking throughout the network.
[0131] The training configuration, related to Step 804 of this embodiment, may prescribe the type and size of the dataset generating models. It may specify the volume of data samples required for effective model training and the criteria for selecting data that will be used to train the dataset generating model. Moreover, it may consider the strategic use of models that have undergone prior training, optimizing the allocation of time dedicated to the training endeavors (for example, if the Type-1 PSFs use previously trained dataset generating models to create additional data and thus update newly received models or train new models) .
[0132] For the evaluation configuration, also related to Step 804 of this embodiment, the PSC 750 may establish a suite of metrics to assess the integrity and performance of the dataset generating models (for example, by distances between authentic data and synthetic data, fidelity, distribution shift, etc. ) . These metrics may be designed to verify the models against target performance and guide the selection of data used for validation purposes (for example, there are predefined values for the metrics, and if they are not met, the updated model may be discarded. ) .
[0133] Transmission configuration, related to Step 805 of this embodiment, may involve defining the ID or address of the corresponding Type-2 PSF, which will receive the trained models. It may also include setting up the data plane configuration, ensuring that the data flow between the PSFs adheres to the network's architectural standards and facilitates efficient data transmission.
[0134] Step 804 shows the actual training modules within the Type-1 and Type-2 PSFs 731, 732 where the dataset generating models are trained. The training process includes determining the type and size of the model, the number of samples to be used, and the specific data selection strategy for training the model. Previously trained models may be utilized to enhance the training process or to generate additional synthetic dataset (s) .
[0135] In Step 805, information is sent from the Type-1 PSFs 731, 732 to the Type-2 PSF 740 through the data plane 720. The information may contain identification as well as detailed configurations of the models, such as weights, architecture, and distribution. It may also include sampling instructions, related to Step 806 of this embodiment, for generating new data samples (for example, how to generate new samples and what type of data to generate from the dataset generating models in the Type-1 PSFs) , which may range from random sampling to using specific text prompts, tokens, images, embeddings, etc. The sampling instructions may also include the number of samples used to train the dataset generating models and the number of samples to generate.
[0136] In Step 806, the Type-2 PSF 740 undertakes the task of updating the dataset generating models to create an aggregate model. This operation is consistent with the processes associated with the centralized node 640, as depicted in FIG. 7, which details the integration of individual models. The resultant aggregate model as updated encapsulates the collective data features, enhancing the data synthesis.
[0137] Subsequent to this, Step 807 entails the issuance of a confirmation notification to the PSC 750. This notification serves a dual purpose: it confirms the successful update of the aggregate model, and it reports any detected errors. These errors are markers for areas of improvement, identified to be addressed in future iterations of the model training cycle, thus contributing to a process of continuous optimization.
[0138] In Step 808, the PSC 750 is presented with the opportunity to evaluate the performance and suitability of the Type-2 PSF 740. Depending on this assessment, the PSC 750 may decide to retain the current Type-2 PSF or elect to select or re-select an alternative PSF. This decision-making process is integral to the adaptive nature of the network, ensuring that the most capable PSF is tasked with the model aggregation at any given point, in alignment with the network's evolving requirements and objectives.
[0139] In Step 809, the PSC 750 orchestrates the selection of Type-1 PSFs 731, 732 as receivers. In other words, it generates a request for routing the generated model (s) to the appropriate nodes.
[0140] In Step 810, the PSC 750 directs the transmission configuration for the Type-1 PSFs, which includes the identifications or receiving addresses of the receiving Type-1 PSFs as well as the traffic configuration. Any errors encountered during this process are noted for future iterations and improvements.
[0141] In Step 811, the Type-2 PSF 740 sends the updated information to the Type-1 PSFs 731, 732 through the data plane 720. The information may contain identification as well as detailed configurations of the aggregate model, such as weights, architecture, and distribution. It may also include sampling instructions for generating new data samples (for example, how to generate new samples and what type of data to generate from the updated dataset generating model in the Type-2 PSF) , which may range from random sampling to using specific text prompts, tokens, images, embeddings, etc. The sampling instructions may also include the number of samples used to train the aggregate dataset generating model.
[0142] Lastly, Step 812 and Step 813 deal with the status and utilization of received models. The PSC 750 confirms the reception and evaluates whether the updated models meet the target performance metrics. Such metrics may include various forms of errors, fidelity checks, and compliance with expected outcomes. If the updated models fail to meet established performance thresholds, indicating suboptimal functionality, they may be either discarded or targeted for further refinement through additional training cycles.
[0143] Step 813 shifts the focus to downstream task modules that are integral to the practical application of the updated models. In this step, the PSC 750 facilitates the deployment of the models in real-world scenarios such as data analysis, predictive maintenance, or other specialized tasks for which the models were originally designed. This deployment relies on the ability of the models to effectively translate their learned patterns and insights into tangible benefits within the operational context. The effectiveness of these models and the synthetic datasets they generate in downstream tasks serves as the ultimate litmus test of the training and aggregation process, confirming the models' readiness to deliver on the network's AI-driven objectives.
[0144] It should be appreciated that while FIG. 9 is set in a 5G or 6G network context, the principles and procedures outlined are adaptable to other network environments with suitable modifications. This adaptability allows for the presented framework to be applied across a variety of evolving telecommunications standards, ensuring its relevance in the advancing landscape of network technologies.
[0145] FIG. 10 illustrates a procedural flowchart showing a receiving procedure from an application perspective, particularly detailing the training of a dataset generating model within a network. At the onset of this process, as depicted in Step 1001, a communication module 910 receives broadcasted new models, denoted as {Mknew} , from other nodes within the network. This initial reception of models is followed by Step 1002, model buffering, where the models are temporarily stored in preparation for further action.
[0146] The flow then proceeds to a trigger event in Step 1003 at a generating module 920, indicating that once a predetermined threshold number of new models, K (where K≥1) , is reached, the process advances to subsequent stages. This threshold ensures that a sufficient variety of models are collected when needed.
[0147] In case that a local dataset generating model is required and available, the generating module 920 may request the local model Mgen from a storage module 940 in Step 1004. The storage module 940 serves as a repository for the generated models and synthetic data, and ensures that relevant data is securely stored and readily accessible for further processing. Then, in Step 1005, the local model Mgen may be sent back to the generating module 920, so that it may be combined with the buffered model (s) .
[0148] In Step 1006, the generating module 920 utilizes the buffered models, which may or may not be combined with the local model Mgen, to create a synthetic dataset. This dataset is synthesized in such a manner that it closely mimics the characteristics of real-world datasets, thereby serving as a stand-in for actual data in the training process, as discussed above in respect of other embodiments described herein. The synthetic data creation is for maintaining the privacy of the original data sources while still providing valuable input for model training.
[0149] In Step 1007, the synthetic dataset is sent from the generating module 920 to a generator training module 930 for training the dataset generating model therein. If an additional local dataset is needed, the generator training module 930 may request such a dataset from the storage module 940 in Step 1008. Then, in Step 1009, the storage module 940 may send the additional dataset back to the generator training module 930 as an expansion dataset.
[0150] Step 1010 depicts the training phase, where the generated synthetic data, along with the additional local dataset obtained from the storage module 940, is employed to train the dataset generating model. This phase is capable of refining the model's ability to produce accurate and useful synthetic datasets. Then, in Step 1011, the updated model is sent to the storage module 940 to be stored therein.
[0151] Although not involved in the process depicted by FIG. 10, a downstream task module 950 is shown to be able to utilize the synthetic dataset generated by the trained dataset generating model for various applications.
[0152] FIG. 11 illustrates a procedural flowchart from an application perspective, specifically outlining a downstream task training procedure for a task-specific generative AI model Mtask. This flowchart delineates a sequence of operations that facilitate the effective training of a generative model within a networked system.
[0153] The process begins with the downstream task module 950, which requests for the local dataset from the storage module 940 in Step 1101, receives the local dataset in Step 1102, request for the local dataset generating model from the storage module Mgen 940 in Step 1103, and receives the local dataset generating model Mgen in Step 1104. The local dataset obtained in Step 1102 may be used to train the model obtained in Step 1104. Then, the downstream task module 950 sends the local dataset generating model Mgen to the generating module 920, which utilizes the model to generate a synthetic dataset in Step 1006, and then send the synthetic dataset back to the downstream task module 950 in Step 1007. In Step 1108, the downstream task module 950 utilizes the synthetic dataset obtained in Step 1107 to train the task-specific generative AI model Mtask. Lastly, the downstream task module 950 sends the trained task-specific generative AI model Mtask to the storage module 940 to be stored therein.
[0154] Throughout the steps depicted in FIG. 11, the flowchart presents a structured approach to training task-specific generative AI models within a networked environment. Each step is methodically placed to ensure a seamless progression from model reception to the application of the trained model in practical, real-world tasks.
[0155] FIG. 12 illustrates a procedural flowchart from an application perspective, specifically focusing on a transmitting procedure for dataset generating models within a network environment. This flowchart maps out the sequence of operations involved in handling requests for and retrieving a dataset generating model.
[0156] In Step 1201, this process begins with the communication module 910, which is responsible for interfacing with other nodes in the network. It is here that a request from another node to share a model is received, indicating a demand for the transmission of a dataset generating model.
[0157] Following the reception of the request, in Step 1202, the process proceeds to query the storage module 940 within the network for the stored model Mgen. Subsequently, the model Mgen is transmitted back to the communication module 910 in Step 1203, completing the request cycle by replying the model Mgen to the requesting node in Step 1204.
[0158] The embodiments have been described above with reference to flow, sequence, and block diagrams of methods, apparatuses, systems, and computer program products. In this regard, the depicted flow, sequence, and block diagrams illustrate the architecture, functionality, and operation of implementations of various embodiments. For instance, each block of the flow and block diagrams and operation in the sequence diagrams may represent a module, segment, or part of code, which comprises one or more executable instructions for implementing the specified action (s) . In some alternative embodiments, the action (s) noted in that block or operation may occur out of the order noted in those figures. For example, two blocks or operations shown in succession may, in some embodiments, be executed substantially concurrently, or the blocks or operations may sometimes be executed in the reverse order, depending upon the functionality involved. Some specific examples of the foregoing have been noted above but those noted examples are not necessarily the only examples. Each block of the flow and block diagrams and operation of the sequence diagrams, and combinations of those blocks and operations, may be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0159] The terminology used herein is only for the purpose of describing particular embodiments and is not intended to be limiting. Accordingly, as used herein, the singular forms “a” , “an” , and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and “comprising” , when used in this specification, specify the presence of one or more stated features, integers, steps, operations, elements, and components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and groups. Directional terms such as “top” , “bottom” , “upwards” , “downwards” , “vertically” , and “laterally” are used in the following description for the purpose of providing relative reference only, and are not intended to suggest any limitations on how any article is to be positioned during use, or to be mounted in an assembly or relative to an environment. Additionally, the term “connect” and variants of it such as “connected” , “connects” , and “connecting” as used in this description are intended to include indirect and direct connections unless otherwise indicated. For example, if a first device is connected to a second device, that coupling may be through a direct connection or through an indirect connection via other devices and connections. Similarly, if the first device is communicatively connected to the second device, communication may be through a direct connection or through an indirect connection via other devices and connections.
[0160] Phrases such as “at least one of A, B, and C” , “at least one of A, B, or C” , “one or more of A, B, and C” , and “A, B, and / or C” are intended to include both a single item from the enumerated list of items (i.e., only A, only B, or only C) and multiple items from the list (i.e., A and B, B and C, A and C, and A, B, and C) . Accordingly, the phrases “at least one of” , “one or more of” , and similar phrases when used in conjunction with a list are not meant to require that each item of the list be present, although each item of the list may be present.
[0161] It is contemplated that any part of any aspect or embodiment discussed in this specification can be implemented or combined with any part of any other aspect or embodiment discussed in this specification, so long as such those parts are not mutually exclusive with each other.
[0162] While every effort has been made to provide a detailed and accurate description of the disclosure herein, it should be noted that the scope of the disclosure is not limited to the exact configurations and embodiments described. The description provided is intended to illustrate the principles of the disclosure and not to limit the disclosure to the specific embodiments illustrated. It is intended that the scope of the disclosure be defined by the appended claims, their equivalents, and their potential applications in other fields.
Claims
1.A method, comprising, at a present node of a network infrastructure:receiving a first dataset generating model from a first upstream node of the network infrastructure, the first dataset generating model being trained using a first authentic dataset at the first upstream node;generating a first synthetic dataset using the received first dataset generating model, the first synthetic dataset simulating the first authentic dataset;combining the first synthetic dataset with an expansion dataset into a combined dataset; andtraining a dataset generating model using the combined dataset to obtain an updated dataset generating model capable of generating an updated synthetic dataset simulating both the first authentic dataset and the expansion dataset.2.The method of claim 1, wherein the dataset generating model is the first dataset generating model or a generic generative artificial intelligence (AI) model.3.The method of claim 1 or 2, further comprising:outputting the updated dataset generating model to a first downstream node of the network infrastructure.4.The method of any one of claims 1 to 3, wherein the first authentic dataset is collected at the first upstream node, and wherein the expansion dataset is a local authentic dataset collected at the present node.5.The method of claim 1 or 2, further comprising:receiving a second dataset generating model from a second upstream node of the network infrastructure, the second dataset generating model being trained using a second authentic dataset at the second upstream node; andgenerating a second synthetic dataset using the received second dataset generating model, the second synthetic dataset simulating the second authentic dataset,wherein the expansion dataset is the second synthetic dataset.6.The method of claim 5, further comprising:combining an additional training dataset into the combined dataset.7.The method of claim 5 or 6, further comprising:8.The method of any one of claims 1 to 7, further comprising:generating the updated synthetic dataset with the updated dataset generating model; andtraining a task-specific generative AI model using the updated synthetic dataset.9.One or more circuits for performing actions, at a present node of a network infrastructure, comprising:receiving a first dataset generating model from a first upstream node of the network infrastructure, the first dataset generating model being trained using a first authentic dataset at the first upstream node;generating a first synthetic dataset using the received first dataset generating model, the first synthetic dataset simulating the first authentic dataset;combining the first synthetic dataset with an expansion dataset into a combined dataset; andtraining a dataset generating model using the combined dataset to obtain an updated dataset generating model capable of generating an updated synthetic dataset simulating both the first authentic dataset and the expansion dataset.10.The one or more circuits of claim 9, wherein the dataset generating model is the first dataset generating model or a generic generative AI model.11.The one or more circuits of claim 9 or 10, wherein the first authentic dataset is collected at the first upstream node, and wherein the expansion dataset is a local authentic dataset collected at the present node.12.The one or more circuits of claim 9 or 10, wherein the actions further comprises:receiving a second dataset generating model from a second upstream node of the network infrastructure, the second dataset generating model being trained using a second authentic dataset at the second upstream node; andgenerating a second synthetic dataset using the received second dataset generating model, the second synthetic dataset simulating the second authentic dataset,wherein the expansion dataset is the second synthetic dataset.13.The one or more circuits of claim 12, wherein the action further comprises:combining an additional training dataset into the combined dataset.14.The one of more circuits of any one of claims 9 to 13, wherein the actions further comprises:generating the updated synthetic dataset with the updated dataset generating model; andtraining a task-specific generative AI model using the updated synthetic dataset.15.One or more non-transitory computer-readable storage devices comprising computer-executable instructions, wherein the instructions, when executed, cause one or more circuits to perform actions comprising:receiving a first dataset generating model from a first upstream node of the network infrastructure, the first dataset generating model being trained using a first authentic dataset at the first upstream node;generating a first synthetic dataset using the received first dataset generating model, the first synthetic dataset simulating the first authentic dataset;combining the first synthetic dataset with an expansion dataset into a combined dataset; andtraining a dataset generating model using the combined dataset to obtain an updated dataset generating model capable of generating an updated synthetic dataset simulating both the first authentic dataset and the expansion dataset.16.The one or more circuits of claim 15, wherein the dataset generating model is the first dataset generating model or a generic generative AI model.17.The one or more circuits of claim 15 or 16, wherein the first authentic dataset is collected at the first upstream node, and wherein the expansion dataset is a local authentic dataset collected at the present node.18.The one or more circuits of claim 15 or 16, wherein the actions further comprises:receiving a second dataset generating model from a second upstream node of the network infrastructure, the second dataset generating model being trained using a second authentic dataset at the second upstream node; andgenerating a second synthetic dataset using the received second dataset generating model, the second synthetic dataset simulating the second authentic dataset,wherein the expansion dataset is the second synthetic dataset.19.The one or more circuits of claim 18, wherein the actions further comprises:combining an additional training dataset into the combined dataset.20.The one of more circuits of any one of claims 15 to 19, wherein the actions further comprises:generating the updated synthetic dataset with the updated dataset generating model; andtraining a task-specific generative AI model using the updated synthetic dataset.
Citation Information
Patent Citations
Online teaching recommendation system with data enhancement
CN112784154A
Generating high-dimensional efficient synthetic data
CN114787826A
Cellular network fault diagnosis method based on knowledge and data fusion
CN115119242A
Generating synthetic data using reject inference processes for modifying lead scoring models
US20200027157A1