REFERENCE GRADIENT SKETCHING IN DISTRIBUTED ARTIFICIAL NEURAL NETWORK TRAINING: CERTIFIED, SELECTIVE COMMUNICATION, AND COPY-WRITE ATOMIC PARAMETER UPDATE METHOD AND SYSTEM

TR202613744A2Pending Publication Date: 2026-09-21ONUR GÜNDOĞDU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TR202613744
Authority / Receiving Office
TR · TR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-08-13
Publication Date
2026-09-21

Smart Images

  • Figure 00000017_0000
    Figure 00000017_0000
  • Figure 00000018_0000
    Figure 00000018_0000
  • Figure 00000019_0000
    Figure 00000019_0000
Patent Text Reader

Abstract

The invention relates to the full delta communication and block-based certification of candidate parameter updates in distributed neural network training prior to writing to the active model. Candidate optimizer deltas are generated in copy-write shadow pages; sketched in a common-seeded linear projection with at least one utility and one protection reference gradient. Block impact certificates are generated by calculating explicit inner product error bounds from the projection size, vector norms, and shared failure probability. During the preparation phase, certificates / sketches are transmitted instead of full deltas, and a common block commitment mask is created from the utility lower confidence limit and protection upper confidence limit. Only the full deltas of the accepted blocks are aggregated. After verifying that the aggregated real delta sketch matches the prepared sketch within identity tolerance, the relevant parameter and optimizer pages are atomically activated with a single generation pass; the active generation is preserved in case of failure.This prevents network transmission and HBM writing of rejected blocks, while ensuring distributed model generation consistency. Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

1 TARIFF CERTIFIED IN REFERENCE GRADIENT SKETCHING IN DISTRIBUTED ARTIFICIAL NEURAL NETWORK TRAINING. SELECTIVE COMMUNICATION AND COPY-WRITE ATOMIC PARAMETER UPDATE METHOD AND SYSTEM 5 TECHNICAL FIELD The invention enables the training of artificial neural networks on single or multiple accelerators; in particular. In distributed learning, candidate parameter updates created are added to the active model parameters. Technical impact via reference gradients at the parameter block level before implementation. certification in terms of the full deltas of only the certified blocks copy-write memory pages of the communication and parameter and optimizer status. Method, system and temporary for enabling generation-identified atomic processing via It relates to non-computer-readable storage media. 15 The invention is particularly suitable for education using GPUs, TPUs, NPUs, FPGAs, or matrix accelerators. Network communication throughput, high-bandwidth memory write-on, and model generation in clusters. It aims to reduce inconsistency problems together. PREVIOUS TECHNIQUE Large artificial neural networks can be data-parallel, model-parallel, expert-parallel, or pipelined. In parallel training, each worker node produces local gradients; all or a significant portion of the gradients or parameter deltas derived therefrom the section with all-reduce, reduce-scatter / all-gather or parameter server operations is aggregated. As the model size increases, network transfer to high-bandwidth memory 25 writing, updating optimizer status, and synchronizing versions between workers. These constitute the significant technical costs of the training phase. Known communication reduction methods analyze gradients based on magnitude, variance, ranking, delay, or It can be diluted using compression ratio criteria or in the form of small-sized sketches. It can transmit. However, if a gradient or delta component is large, the word 30 The issue is improving the Delta's utility objective while maintaining the old mission, safety, stability, or other aspects. That alone does not prove that it will not compromise the conservation objective. Selective parameter updating and gradient conflict methods use certain parameters. It can freeze or compare the directions of different task gradients. However The election result is mostly a block mask for the collective communication that will be carried out later, active 35 Parameter page writing mask and model generation that all workers will pass through on the same machine. It is not combined into a verifiable transaction record. 2 Speculative or shadow update schemes before activating candidate results It can test this. However, once the exact candidate deltas have been generated or reported, The subsequent verification involves network transmission and HBM writing for blocks to be rejected. It does not prevent; furthermore, the prepared candidate and the subsequently aggregated actual delta are the same. It does not require block-based sketch-based authentication to prove that it is. 5 Therefore, the primary focus of candidate updating is on the benefits and protections. limiting its effect to low-dimensional co-projections before full delta communication, word the subject transforms boundaries into a mask of shared communication and memory writing, accepted alone aggregating the exact deltas of the blocks, the aggregated true delta with the prepared island re-sketching and verifying its identity and matching the parameter and optimizer pages to the same 10 A need for a connected transaction protocol that enables atomic-level operation under a generation token. It is located. A BRIEF DESCRIPTION OF THE INVENTION The technical problem solved by the invention is that it does not provide benefit or protection in distributed education. 15 Parameter blocks exceeding their targets, full network communication, and high-bandwidth memory. Parsing before writing, preparation information regarding selected blocks This was later confirmed by the actual delta data reported in the news, and all workers were monitored using the parameter. The optimizer state is transferred to the same new generation without creating a mixed intermediate generation. One aim of the invention is to achieve full delta communication of parameter blocks to be rejected in distributed education. To prevent. Another objective of the invention is to provide parameter and optimizer status HBM for rejected blocks. The goal is to prevent its writing. Another objective of the invention is to establish at least one benefit and at least one conservation reference with the candidate delta. The goal is to calculate the effect between the gradients on a block basis, with explicit error and confidence limits. 25 Another purpose of the invention is to calculate a total failure probability budget using block-reference pairs. The goal is to determine the simultaneous level of trust in the shared commitment mask by distributing it among them. Another purpose of the invention is to aggregate the preparation sketch with the actual delta sketch. by comparing the wrong block, old generation, different quantization mode, or corrupted packet commitment. The goal is to prevent. 30 Another objective of the invention is to improve the accepted parameters and the optimizer status associated with them. Create pages in a copy-write format under the same token and with a single generation transition. It is to activate. Another purpose of the invention is to allow cancellation and optional return without creating an exact second model copy. The aim is to provide the opportunity to receive. 35 3 One aspect of the invention is that the active model parameters are independently addressable parameters. They are divided into blocks and active parameter pages under an active generation ID. Candidate optimizer deltas are obtained from local gradients calculated at worker nodes. These deltas are being created; copy-write shadow deltas are created without being written to the active pages. It is prepared on pages 5. For each block, a candidate delta, at least one benefit reference gradient, and at least one conservation reference gradient are required. The gradient is applied to a linear projection matrix generated from the same shared seed. The inner products in the projection space have a primary effect on the reference losses of the candidate delta. It estimates the degree of impact. Projection size, vector norms, separated failure. probability, quantization, reference staleness and, if present, second-order curvature limit are used for every 10 A clear total error limit is calculated for impact estimation. Worker nodes may use block impact certificates instead of full candidate deltas during the preparation phase, or The candidate submits the delta sketches to the process coordinator. The coordinator sets a lower confidence limit for the benefit as a threshold. by identifying the blocks that remain below the relevant tolerance on and for each protection upper safety limit It creates a common block commitment mask; the total failure probability budget covers all 15 in the certificate. The block-reference decisions are allocated in such a way that they are valid together. Candidate deltas with full or defined sensitivity of blocks accepted only in the common mask. It is aggregated through collective communication. The aggregated real delta is reprojected and from the batch sketch in the preparation phase, quantization and numerical summation error limits Matching is verified within the derived identity tolerance. 20 New parameter and optimizer status shadow pages are created for matching blocks. Active The root pointer or equivalent generation identifier of the page table is the only one tied to the commitment token. with a compare-and-swap, read-copy-update, accelerator barrier, or distributed generation process It is passed on to the next generation. If the mating or preparation condition is not met, the active generation is preserved and Shadow pages are invalidated. 25 BRIEF DESCRIPTION OF THE FIGURE Figure 1 shows the general architecture of the certified block-commitment training system. Figure 2 shows the active parameter pages, copy-write shadow pages, and atomic page table. The transition is shown. 30 Figure 3 shows the end-to-end steps of the method. Figure 4 shows the joint projection of the candidate delta and reference gradients, inner product estimation, and The process of generating a block domain certificate with an open error margin is demonstrated. Figure 5 shows worker readiness messages, shared trust budget, block commitment mask, and commitment. The token flow is shown. 35 4 Figure 6 shows sketch-to-authentication after aggregation with selective full delta communication. It is shown. Figure 7 shows the successful atomic transition, cancellation, and optional rollback scenarios. Figure 8 shows the integration of the optional dataset pre-evaluator with the training framework. It is shown. 5 DETAILED DESCRIPTION OF THE INVENTION The following descriptions are provided to facilitate understanding of the invention and are protected. It does not limit its scope to examples. Parameter block; a layer, tensor, tensor slice, attention 10 independent memory modules mapped to a title, expert network, channel group, or logical / physical memory region It is a subset of addressable parameters. The page must be an operating system page. contiguous or logical memory that is not present but can be versioned with a pointer or index table It is the region. Delta, the gradient itself, or SGD, momentum, Adam, AdamW, Adafactor, and This includes parameter variation generated from the gradient by a similar optimizer. 1. System Architecture 15 According to Figure 1, the system (100) has one or more worker processing nodes (110), active parameters repository (120), shadow delta repository (130), optimizer status repository (140), reference gradient bank (150), projection and sketch engine (160), certificate generator (170), transaction coordinator (180), selective communication engine (190), atomic commitment engine (200), optional It includes a validation and retrieval unit (210) and an optional dataset pre-evaluator (220). 20 Worker processing node (110), CPU, GPU, TPU, NPU, FPGA, dedicated matrix accelerator or It is a combination of these. In a multi-node application, workers use Ethernet, InfiniBand, NVLink, PCIe can communicate over CXL or equivalent bus. The collective communication engine is NCCL. The MPI can be a gRPC-based parameter server or an equivalent communication layer. It contains the active parameter store (120), the active page table (121), and the active generation ID (122). 25 The active generation ID indicates which parameter version workers initiated a training process from. The determining monotonic counter is the epoch-step composition, UUID, or completeness hash. (Ongoing) The active parameter pages will remain closed to writing until the process is complete. Shadow delta repository (130) holds candidate delta sheet (131-b) for each parameter block b. Physical 30 only for blocks containing a non-zero delta or exceeding a pre-selection criterion. Shadow pages can be allocated. Shadow pages can be HBM, RAM, CXL memory, NVMe buffer, or It can be stored in a compressed memory pool. 2. Reference gradient bank Reference gradient bank (150), at least one utility reference gradient (151) and at least one Includes the protection reference gradient (152). Benefit reference target mission, target area or 35 loss of validation subset; protection reference legacy task performance, critical class, security, calibration, privacy, physical stability, numerical finiteness, or otherwise It can be calculated from the regression metric. Reference gradients do not need to be updated at every training step. Updates are necessary for specific training periods. In the step interval, during model generation change, the loss slope or distribution shift threshold is 5. This can be done when the limit is exceeded. Each reference record has a source dataset and loss function ID. model generation in which it was produced, timestamp, upper limit of the norm, gradient staleness limit if any, and It has a validity window. 3. Candidate optimizer delta and first-order effect. The local gradient g(i,t,b) is calculated for worker i and parameter block b in training step t. 10 The optimizer generates the candidate delta Δ(i,t,b) without writing it to the active parameter page. Worker weights The aggregate candidate delta for a_i is in the following form: Δ(t,b) = Σ_i a_i · Δ(i,t,b) If the gradient of the reference loss number k around the active parameters is p(k,t,b), then the candidate The first-order loss change of delta is the inner product E(k,t,b)=p(k,t,b)^TΔ(t,b). Loss reduction 15 For the intended benefit objective, the positive benefit is defined as B(t,b)=-E(benefit,t,b); for the objective j to be protected. The regression can be defined as D(j,t,b)=E(protection_j,t,b). If the reference loss gradient is L(k,t,b)-Lipschitz in the block region, then the absolute Taylor residual is... The maximum value is (L(k,t,b) / 2)||Δ(t,b)||². If this limit is not met, the certificate is first. The degree assumption flag carries or a more cautious fixed upper limit of 20 for protection purposes. It is used. 4. Application of concrete projection, inner product, and error limiting. In a preferred implementation configuration, the size of block b is n_b, the projection size is m_b, and The common seed is determined as ξ(t,b). The projection and sketching engine (160) uses the seed in question. It generates the matrix R_b∈R^(m_b×n_b). In the Gaussian Johnson-Lindenstrauss application, every 25 The matrix element is independent of the form Z(r,l) / √m_b, where Z(r,l) is a standard normal random matrix. It is variable. In the alternative Rademacher application, Z(r,l) has an equal probability of +1 or -1. It is valuable. The candidate delta and reference gradient sketches are calculated as follows: sΔ(b) = R_b Δ(b) ; s_k(b) = R_b p_k(b) 30 The inner product estimate is of the form Ê(k,b)=sΔ(b)^T s_k(b). In Gaussian application, this estimate is... The expected value is equal to the actual inner product. The variance is as follows: Var[Ê(k,b)] = (||Δ(b)||² ||p_k(b)||² + (Δ(b)^T p_k(b))²) / m_b 6 Using the Cauchy-Schwarz inequality, the variance in question is calculated as 2||Δ(b)||²||p_k(b)||² / m_b It is bounded from above. For each block-reference pair, the allocated failure probability ρ(k,b) is calculated. The conservative projection error margin is calculated as follows: ε_proj(k,b) = ||Δ(b)|| · ||p_k(b)|| · √(2 / (m_b ρ(k,b))) The desired maximum projection inner product error is η(k,b), the candidate delta norm upper bound is D_b and 5 When the reference gradient norm upper limit G(k,b) is predetermined, the projection size is at least It is set to the following integer: m_b ≥ ceil( 2 D_b² G(k,b)² / (ρ(k,b) η(k,b)²) ) If there is more than one reference objective for the same block, m_b is the sub-subject calculated for all objectives. It is the largest of the limits. If the upper limit of the norm is exceeded, a larger limit of 10 is imposed for the relevant block. The projection is recreated, the full inner product is calculated, or the block certificate is excluded. The total error margin for the certificate does not necessarily have to consist solely of projection error. In the application format, the total error is generated from the following components: ε_top(k,b) = ε_proj(k,b) + ε_nic(k,b) + ε_bayat(k,b) + (L(k,b) / 2)||Δ(b)||² Here, ε_nic is the norm 15 of the quantized delta to be communicated later, relative to the exact island. The error is obtained by multiplying the p_k norm; ε_bayat is the current reference and the reference in the bank. The difference in norms between them is determined by multiplying the upper limit of the uniform norm by the delta norm. The quantization step delta norm error upper limit for coordinates q_b and n_b is √n_b q_b / 2. It can be obtained as follows: ε_nic(k,b)≤||p_k(b)||√n_b q_b / 2. For all Q block-reference decisions to be valid simultaneously in the same process, the coordinator must be 20. (180) distributes the total failure probability budget α to the values ​​ρ(k,b) and Σ_(k,b)ρ(k,b)≤α It applies the condition. In an equal distribution, ρ(k,b)=α / Q can be chosen. In this case, the block effect The common confidence level of the certificates is recorded as at least 1-α. In the CountSketch application, the R_b matrix contains d groups of independent rows, with a hash for each group. function h_r:{1,…,n_b}→{1,…,w} and sign function σ_r:{1,…,n_b}→{-1,+1} with 25 It is defined; each coordinate adds the value σ_r(l)x_l to the h_r(l) bucket. The certificate, w, d, hash and signal seeds, row-based inner product estimates, and selected probabilistic or calibration It carries a base error limit. However, the Gaussian matrix in its main implementation form and the above The explicit error limit is a directly applicable one, regardless of the use of CountSketch. This is an example. 30 5. Benefit and protection safety limits Estimated benefit for utility objective B_hat(b)=-Ê(fayda,b), estimated decline for conservation objective j D_hat(j,b)=Ê(protection_j,b) is calculated. The certificate producer (170) has the following lower and upper It establishes trust boundaries: LCB_utility(b) = B_hat(b) - ε_tot(utility,b) 35 7 UCB_protection(j,b) = D_line(j,b) + ε_top(protection_j,b) If the second-order limit is to be used on the utility side, the lower limit of the benefit, the possible curvature. It is created by deducting the residual. On the protection side, the residual curvature is always at the upper limit of the regression. Added. Numerical NaN / Inf, norm limit exceedance, or block with invalid reference window. The confidence limit is not considered valid for this. 5 6. Block effect certificate Certificate producer (170), machine-readable block impact certificate (171) for each block It creates a certificate with at least the transaction ID (TID), base active generation (v), block ID (b), and worker ID (v). or worker set, projection type and size m_b, projection seed ξ, candidate delta sketch and norm, utility and conservation inner product estimates, ε_project and ε_top values, LCB / UCB 10 values, ρ(k,b), total α, budget ID, quantization mode, reference versions, optimizer status version, numerical finiteness flag, shadow page integrity summary, and validation. It carries its window. The certificate is not a proof of absolute truth in a mathematical sense; it refers to the projection mentioned. candidate updating under distribution, norm upper limits, error components and probability budget 15 It is a machine-verifiable data structure that carries its effect and identity. Certificate or pledge. The HMAC token is certified as integrity through digital signature or trusted execution environment attestation value. It can be taken under protection. 7. Preparation phase and joint block commitment mask Each worker uses local candidate deltas 20 via the base generation v specified in the process identifier (181). and produces certificates. The preparation message (182) produces block certificates instead of full delta tensors. It carries local delta sketches, norms, and required version metadata. Linear projection. Therefore, local sketches are collected to create a comprehensive candidate delta sketch. The coordinator generates M(b)=1 for blocks that satisfy the following conditions together: • The LCB_benefit(b) value must be equal to or greater than the τ_benefit(b) threshold defined for the block; 25 • For each protection objective j, the UCB_protection(j,b) value must be equal to or equal to the τ_protection(j,b) tolerance. being small; • active generation, projection seed, block ID, optimizer version, and certificate validity matching information; • The candidate's delta norm, numerical finiteness, step limit, and atomicity group conditions must be met (30). ensuring; • communication bytes, HBM write bytes, latency, or maximum acceptable block count budget not exceeding. M(b)=1 ⇔ LCB_benefit(b)≥ ∧τ_benefit(b) UCB_protection(j,b)≤τ_protection(j,b), for all j 8 If a confidence interval intersects one of the thresholds, the projection size for the coordinator block is adjusted. can be enhanced, re-sketched with an independent second seed, but fully internal in the block in question. It can calculate the multiplication or reject the block. Timeout or unresolved ambiguity. It is not considered safe. 8. Commitment token and selective full delta communication 5 When the preparation conditions are met, the coordinator will provide the process ID, base generation v, target generation v+1, Common block mask M, projection seeds, total α confidence budget ID, preparation Commitment token (184) containing a summary of certificates, expiry and optional signature It publishes. Workers will not initiate full Delta communication until the token has been verified. The selective communication engine (190) selects only blocks with M(b)=1 with full or defined precision. Candidate deltas can be reduced using all-reduce, reduce-scatter / all-gather, parameter server push / pull, or It aggregates with equivalent operation. The exact deltas of blocks where M(b)=0 are not included in the collective communication; The local delta is discarded, either transferred to its buffer or carried over to the next step. Accepted blocks can be packaged according to size, topology, or availability time. Different Blocks can be transmitted in different quantization modes; the upper limit for quantization error is 15 for each mode. It is linked to certification and identity tolerance. 9. Sketch-identity verification after aggregation. After the actual deltas of the selected blocks are aggregated, the projection engine uses the same R_b It recalculates the sketch sΔ′(b)=R_bΔ′(b) using the matrix. The sketch preparation in question This is compared with sΔ(b), which is formed by aggregating local sketches in the stage: 20 ||sΔ′(b)-sΔ(b)||₂ ≤ δ_id(b) Identity tolerance δ_identity(b), quantization, collective summation order, floating-point rounding and is calculated from the upper limits of the allowed compression residual at the time of certificate creation. For the uniform quantization error e_q, δ_identity(b) is numerically summed with at least ||R_b||₂||e_q||₂. It includes the total of the error limit. If the measured difference exceeds this limit, it indicates a wrong block, wrong generation, or defective 25. The packet assumes a missing worker or a different quantization mode, and the relevant atomic group is committed. It cannot be done. 10. Copy-write parameter and optimizer pages For verified and accepted blocks, the active parameter page is not modified in place. Using the active page content and aggregated delta, a new parameter shadow page 30 is created. They are created the same way. In Adam or similar optimizers, the first and second moment pages are also the same. The block mask is updated using a copy-write method. The parameters of the rejected blocks are... And the optimizer pages will continue to be shared without creating a physical copy. Atomic commitment engine (200), all accepted parameters and their associated optimizers When the status pages are ready, the root pointer of the active page table, generation 35 9 It passes the identifier or version table to v+1 at a single commit point. Other training The nuclei either see all of the v generation or all of the v+1 generation; a mixed intermediate state. It does not observe. The application uses compare-and-swap, read-copy-update, lock, and accelerator kernel. This can be implemented with a barrier tolling system or a distributed generation token. 11. Cancellation, verification and undo 5 If an error occurs during the preparation, authentication, or page creation phase, Shadow Pages (202) are marked invalid and active generation v is preserved. Optional after commitment. verification unit (210), a probe dataset, numerical finiteness test or hardware integrity The test can check for generation v+1. If the check fails, the active page table can check for the previous one. The generation can be reverted to the root or an inverse delta can be applied. 10 The commitment token and transaction identifier determine which parameter and optimizer status blocks It determines that it will be undone. Previous pages continue to be shared throughout the storage window. Therefore, a complete second model copy is not required for a refund. 12. Single-node, distributed, and parallelism modes In a single-accelerator application, the coordinator and the worker can be present during the same working time. 15 The preparation phase involves evaluating the small in-device sketch pad. This is achieved through technical means, specifically by writing HBM parameters and optimizer scripts for rejected blocks. This is to prevent the collective rejection of blocks in multi-accelerator applications. Communication is prevented. The model uses a layer or tensor slice with a block ID held by a specific worker in a parallel application of 20. They are synchronized. In a parallel application, the same block of data is present in all workers, and a common mask is used. In the expert parallel model, each expert is a separate atomic group; in the pipeline parallel model, each The next stage could be a local certificate set. 13. Hardware and training framework integration. System PyTorch DistributedDataParallel, Fully Sharded Data Parallel, TensorFlow 25 Distribution Strategy can be integrated into JAX pmap / pjit or custom training runtime. Backpropagation hooks and the optimizer.step wrapper insert the candidate delta into the shadow buffer. It redirects. The collective call selector creates a tensor list from only M(b)=1 blocks. Page table, tensor storage pointers, CUDA graph parameter bindings, or custom runtime This can be implemented via the descriptor table. 30 Network queue, HBM bandwidth, DMA latency, collective time, and error counters in communication. and can be included in the writing budget. These counters replace the reference gradient impact certificate. It doesn't pass; it only provides additional technical input for budget and packaging decisions. 14. Optional dataset pre-evaluator The optional dataset pre-evaluator (220) is a dataset before the main backpropagation. It can predict whether or not it will be processed. The pre-evaluator output predicts the benefit threshold or communication. It can adjust its budget; however, the block impact certificate, selective communication, and atomic commitment chain It is not a substitute. 5 15. Example numerical study In an illustrative application, a forty-eight-block system trained on eight GPUs. The Transformer model is used. For each block, a Gaussian projection of dimension m_b=4096 is used. There are four reference objectives. The total confidence budget is α=0.01, and the following are considered in the evaluation. They are distributed into block-reference pairs. Workers are of the same generation v=12540 and the same projection seed 10 It generates certificates through this process. The coordinator is a partner who accepted thirty-one blocks as a result of confidence limits and budget conditions. It can create a mask. However, the exact deltas of these blocks are all-reduced. In a block, again... If the calculated actual delta sketch exceeds the identity tolerance, the block in question is atomic. It is cancelled along with its group. The parameter and optimizer shadow pages of the remaining blocks are 15. The root page table is created and set to v=12541 in a single pass. The numbers here are only for the calculation. This describes the flow and is not a claim based on measured performance or convergence results. 16. Experimental verification protocol Application validation uses intensive DDP, top-k dilution, sketch-only SGD, and selective methods. Parameter updating and speculative training baselines can be used. The technical data to be measured is 20. Values ​​include: full delta network bytes per iteration, provision sketch bytes, and bytes written to HBM. parameter and optimizer byte, collective duration, atomic transition time, accepted block The indicators are protection breach rate, false acceptance rate, target loss, and total wall clocks. Projection size m_b, total confidence budget α, block size, number of references, and The effect of the quantization step on decision quality is evaluated through separate experiments. (Disrupted 25) package, missing worker, old generation, different seed, timeout, and quantization mode mismatch error Authentication and fail-closed behavior are tested through injections. This specification... Unverified percentage savings do not constitute a guarantee of accuracy or security. Industrial applicability The invention encompasses 30 large-scale models of language, image, speech, suggestion, robotics, and industrial control. In data centers where they are trained; in bandwidth-limited edge or federated systems; in multiple It is applicable in GPU / TPU clusters and memory-limited systems with a single accelerator. The system, training time optimizer and software module for collective communication layer or a hardware-assisted controller within an accelerator, SmartNIC, or collective processor. It can be placed as follows: 35 11 The scope of protection of the invention is defined in the claims and the application given herein. It is not limited to these forms. A technical expert analyzes the linked process chain in the requests. protecting different optimizers, linear projection, collective communication, or page management. He can use techniques. REFERENCE NUMBERS GIVEN IN THE FIGURE 5 100: Certified block-commitment training system 110: Worker process node 120: Active parameter store 121: Active page table 10 122: Active generation identity 130: Shadow Delta Depot 131: Shadow delta / parameter page 140: Optimizer status repository 150: Reference gradient bank 15 151: Utility-reference gradient 152: Protection reference gradient 160: Projection and sketching engine 161: Projection seed 162: Candidate delta sketch 20 Fig. 163: Reference gradient sketch 164: Error limit calculator 165: Projection size selector 170: Certificate producer 171: Block effect certificate 25 180: Operations Coordinator 181: Transaction identifier 182: Preparation message 183: Block commitment mask 184: Commitment token 30 185: Joint trust budget 190: Selective communication engine 191: Collective consolidation process 200: Atomic Contracting Engine 201: Page marker / generation transition 35 202: Cancel and shadow page removal 210: Verification and Retrieval Unit 12 211: Post-commitment probe verification 212: Return to the previous generation 220: Optional dataset pre-evaluator 221: Preliminary assessment decision S101: Separation of active parameters into blocks and active generation 5 S102: Generating the local gradient and candidate optimizer delta on the shadow page. S103: Obtaining benefit and protection reference gradients. S104: Creation of candidate delta and reference gradient sketches S105: Creation of block effect certificates S106: Transmitting preparation messages and creating a group sketch 10 S107: Generation of the common block commitment mask and commitment token S108: Consolidation of the full deltas of only the accepted blocks S109: Verification of the aggregated delta sketch against the prepared sketch. S110: Creation of accepted parameter and optimizer shadow pages S111: Atomic migration of the page table to the next generation 15 S112: Cancellation or optional recall in case of failure.

Claims

13 REQUESTS 1. Training an artificial neural network using multiple processing nodes. It is a method performed by a computer and its characteristic is; (a) the artificial neural network is active 5 parameters, independently addressable parameter blocks under an active generation ID (b) stored in the active parameter pages; calculated in each of the transaction nodes from local gradients, belonging to parameter blocks without being applied to active parameter pages. (c) generating candidate optimizer deltas in copy-write shadow delta pages; each The candidate optimizer delta of the parameter block, at least one benefit reference gradient, and the maximum a small conservation reference gradient, produced from a shared seed and with a projection size of 10 sketching using a common linear projection matrix normalized with; (d) each Generating a first-order effect estimate from the sketch's intraproduct for block and reference purposes. projection size, candidate delta norm, reference gradient norm, and the block in question. Calculating a projection error limit from the failure probability allocated to the reference pair. and the impact estimate, margin of error, active generation identity, and candidate delta integrity 15 (e) creation of a block impact certificate containing information; a full candidate in the preparation phase Processing block impact certificates or candidate delta sketches instead of optimizer deltas from the nodes to a process coordinator and candidate delta sketches (f) aggregation of the total failure probability budget by the coordinator Distribution of block-reference decisions in the certificate, lower confidence threshold for benefit effect is 20 the defined benefit threshold is met and the upper confidence limit for each protective effect is the relevant decline. common block commitment mask showing parameter blocks that provide tolerance identification and active generation ID, transaction ID, the mask in question and the projection seed (g) issuance of the commitment token linked to the commitment token; block commitment only according to the commitment token. Collective 25 of the complete candidate optimizer deltas belonging to the accepted parameter blocks in the mask aggregated communication and the full deltas of rejected blocks to that communication. (h) not to be included; the same projection from aggregated actual candidate optimizer deltas With the sketch that was recreated with the matrix, and the sketch that was aggregated during the preparation phase, that it matches within the identity tolerance derived from quantization and numerical summation error limits (i) verification; and (i) the parameter blocks accepted if verification is successful 30 copy-write shadow pages relating to these and their corresponding optimizer status blocks their markers become atomically active with a single generation transition tied to the commitment token. if verification fails, the shadow identity will remain unchanged without changing the active generation ID. It includes steps to invalidate the pages.

2. The method according to claim 1 is that the parameter block is a neural network layer, tensor, tensor 35 at least from the segment, attention header, expert network, channel group or logical memory region It is being paired with someone. 14 3. The method is according to claim 1, and the linear projection matrix has dimensions of m_b×n_b. and its elements are independent of the shared seed standard normal variables with 1 / √m_b by scaling or scaling independent Rademacher signals by 1 / √m_b is the creation of.

4. The method according to claim 3 is the inner product estimation for block b and reference k. The projection error limit is defined as Ê(k,b)=(R_bΔ_b)^T(R_bp_k,b). It is calculated as ε_proj(k,b)=||Δ_b||₂||p_k,b||₂√(2 / (m_bρ(k,b))).

5. The method is according to claim 4, and the desired maximum projection error is η(k,b), candidate delta norm The projection dimension for the upper bound D_b and the upper bound G(k,b) of the reference gradient norm The condition m_b≥ceil(2D_b²G(k,b)² / (ρ(k,b)η(k,b)²)) is selected. 10 6. The method is according to claim 1, where the total failure probability budget α and Q are block-references. For the decision, ρ(k,b)=α / Q or weighted ρ(k,b) values ​​whose sum does not exceed α are used. The certificate must record the common confidence level of the confidence intervals as at least 1-α.

7. The method is according to Claim 1, where the total error limit is added to the projection error limit. Quantization error limit, reference gradient staleness limit, and reference loss gradient 15 It must contain at least one of the second-order Taylor residual limits determined from the Lipschitz constant.

8. The method is according to Claim 1, and the block impact certificate also includes the block ID, projection type, and size, projection seed, candidate delta norm, reference gradient version, optimizer Status version, numerical finiteness flag, timestamp, validity window, and shadow. The page must include at least one of the integrity summaries. 20 9. The method is according to Claim 1, and the active generation identities in worker messages are projected. When seeds, block IDs, or optimizer versions don't match, or the certificate... The commitment token will not be generated when its validity period expires.

10. The method is according to Claim 1, and the total communication budget of the block commitment mask is high. bandwidth memory write budget, latency budget, or maximum acceptable block 25 The number is limited to at least one of them.

11. According to claim 1, the method is relevant when the confidence interval intersects the benefit or protection threshold. Increasing the projection size for the block, re-sketch with an independent second seed. one of the following: creation, calculation of the exact inner product, or rejection of the block. It is the implementation. 30 12. The method according to Claim 1 is the complete candidate optimizer deltas of the accepted blocks. from all-reduce, reduce-scatter / all-gather or parameter server aggregation operations communication with someone and the full deltas of the rejected blocks in question regarding the aggregation not being subjected to the process.

13. The method, according to Claim 1, is to determine the minimum candidate delta of a rejected parameter block. a portion of it is now held in the buffer and will be used with the candidate delta in the next training step. It is the combination.

14. The method according to claim 1 is the operator norm of the projection matrix of the identity tolerance. multiplied by the quantization delta norm error upper bound and the sliding scale error resulting from collective summation. It includes the upper limit of the point error.

15. The method is according to Claim 1, and the parameter page is for the accepted parameter block. The first moment and second moment optimizer status pages share the same commitment token. It involves creating and activating it together using copy-write methods.

16. The method is according to claim 1, where the active page table of the atomic generation transition is root 10. pointer compare-and-swap operation, read-copy-up operation, accelerator kernel barrier or this is accomplished by exchanging it via a distributed generation token.

17. The method is according to Claim 1, and after commitment, a probe dataset or hardware. Validation is performed on the integrity test, and the active page is activated when the validation fails. This is a return to the previous generation root of the table. 15 18. The method is in accordance with Claim 1, and the benefit and protection reference gradients are calculated in specific steps. in the range, during model generation change, when the loss slope threshold is exceeded or distribution shift It is the renewal that occurs when it is perceived.

19. A system that performs distributed training of an artificial neural network; its characteristic is the active model. The parameters are divided into blocks in active parameter pages and under the active generation ID 20 Holding active parameter repository; block-based candidate optimizer without writing to active parameters. A copy-write shadow delta repository that holds deltas; at least one benefit and at least one protection. Reference gradient bank containing reference gradients; candidate deltas and references. gradients from shared seed to common linear projection matrix The inner product is 25 based on the converter, projection size, and block-reference failure probability. Projection, sketching, and error limit engine that calculates error limits; block-based utility estimate. Block effect including effect, protection effect, fault limits, active generation ID, and integrity information. The certificate producer that generates the certificates; instead of using full deltas during the preparation phase, the relevant deltas are used. Joint block commitment under total failure probability budget by obtaining certificates or sketches The transaction coordinator that produces the mask and generation-identity commitment token; only the mask is accepted 30. A selective communication engine that aggregates the exact deltas of the generated blocks through collective processing; and aggregated true delta sketch with prepared sketch quantization and numerical error Accepted parameter and optimizer status after verification within tolerance. pages that are activated with a single generation pass, the active generation when verification fails It includes a protective atomic commitment engine. 35 16 20. The system is defined in accordance with claim 19, and the worker processing nodes are GPU, TPU, NPU, FPGA or It must include at least one matrix accelerator and the selective communication engine must be InfiniBand. Communication can be via Ethernet, NVLink, PCIe, or CXL.

21. The system is defined according to claim 19 and includes an active parameter repository, a shadow delta repository, and an optimizer. The status store is stored in high-bandwidth memory via a logical page table. It is versioning.

22. According to Claim 19, the system uses a commitment token message verification code, digital signature, or A reliable execution environment means preserving its integrity through the attestation value.

23. According to claim 19, the system is the process coordinator and selector in a single-process node application. The communication engine and the atomic commitment engine run in the same accelerator runtime 10 high band for parameter and optimizer status for found and rejected blocks This is preventing large-scale memory writes.

24. When executed by one or more processors, any of prompts 1-18 According to the method, the non-volatile computer contains instructions that execute the procedure. A readable storage medium. 15 25. According to claim 24, it is a storage medium where instructions are retrieved within a deep learning framework. spread hook, optimizer step wrapper, projection size, and confidence budget as a calculator, collective communication call selector, and active page table transition module It is the implementation.