Data security training method and system for industrial large model under federated learning architecture

CN122824618APending Publication Date: 2026-09-25CRRC IND INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610618499.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]传统集中式工业大模型训练架构强制要求各个分散的生产节点将包含全量细节的原始生产数据通过公共网络链路直接传输汇集至云端服务器,这种高风险的数据物理迁移模式导致核心生产工艺参数与敏感故障特征完全暴露于外部传输环境及第三方存储介质中,极易引发严重的工业机密泄露与数据隐私安全危机,同时缺乏价值评估的无差别海量数据上传持续挤占宝贵的网络带宽资源,造成通信链路在高频次迭代过程中发生严重拥塞与延迟,且云端算力资源难以聚焦于高价值样本特征,致使整体训练效率低下并伴随巨大的资源浪费

Benefits of technology

[0014]与现有技术相比,本发明的优点和积极效果在于:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824618A_ABST
    Figure CN122824618A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data security, in particular to an industrial large model data security training method and system under a federated learning architecture, comprising the following steps: calculating the marginal performance contribution degree of node gradient by using a standard validation set, matching differential sparsity and quantization communication instructions according to the contribution degree, guiding the node to perform gradient truncation and encoding compression, calculating global aggregation weights based on contribution degree normalization numerical value, performing weighted accumulation and model parameter adjustment on compressed encrypted data packets, and constructing an industrial defect recognition global model.In the present application, by constructing a differential communication and aggregation mechanism based on marginal performance contribution degree feedback, the actual gain of each node parameter update on model performance is accurately quantified by using a standard validation set, the quantization accuracy and sparsity ratio of each node communication are dynamically regulated according to the contribution degree, the bandwidth resource is intelligently tilted to high value features, the communication bottleneck is broken, and the capture accuracy of the industrial large model on complex fault features is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data security technology, and in particular to a data security training method and system for large industrial models under a federated learning architecture. Background Technology

[0002] Data security technology involves protecting digital information to prevent unauthorized access, use, disclosure, destruction, modification, or damage, and ensuring the confidentiality, integrity, and availability of data. This field encompasses encryption technology, access control, data anonymization, and security auditing, aiming to build a defense system through technical means and management strategies to protect information assets from threats in the network environment. Traditional industrial large-scale model data security training methods refer to the process of building models using a centralized architecture. In this process, each distributed industrial production node directly transmits its locally collected raw production data to a central server deployed in the cloud via network links. The central server receives and stores the raw data containing production details and then uses its computing resources to perform unified iterative training and parameter updates on the large model.

[0003] Traditional centralized industrial large-scale model training architectures force each distributed production node to directly transmit raw production data containing full details to the cloud server via public network links. This high-risk physical data migration mode exposes core production process parameters and sensitive fault characteristics to the external transmission environment and third-party storage media, which can easily lead to serious industrial secret leaks and data privacy and security crises. At the same time, the indiscriminate uploading of massive amounts of data without value assessment continuously consumes valuable network bandwidth resources, causing severe congestion and delays in communication links during high-frequency iterations. Furthermore, cloud computing resources cannot be focused on high-value sample features, resulting in low overall training efficiency and huge resource waste. Summary of the Invention

[0004] To address the technical problems existing in the prior art, embodiments of the present invention provide a method for secure training of industrial large-scale models under a federated learning architecture, comprising the following steps: S1: Obtain a standard verification dataset of defect features of multiple types of industrial products, construct a temporary model copy, perform a recognition accuracy test on the temporary model copy, and obtain the marginal performance contribution. S2: Call the marginal performance contribution and perform a numerical comparison operation with the preset contribution judgment threshold. For nodes whose marginal performance contribution is greater than the contribution judgment threshold, match a 16-bit floating-point number width with a low sparsity truncation threshold that retains multiple features. For nodes whose marginal performance contribution is less than the contribution judgment threshold, match a 4-bit integer width with a high sparsity truncation threshold that retains fewer features. Construct differentiated communication interaction instructions. S3: Obtain the original gradient value matrix generated by local training, parse the sparsity configuration item and bit width configuration item in the differentiated communication interaction instruction, remove non-key elements in the original gradient value matrix whose absolute value is less than the sparsity configuration item, and construct a compressed encrypted transmission data packet. S4: Based on the compressed and encrypted transmission data packet, calculate the proportion of the marginal performance contribution of a single node to the sum of the marginal performance contributions of all nodes, and obtain the global aggregation weight factor.

[0005] As a further aspect of the present invention, the marginal performance contribution includes the accuracy gain value and the defect identification increment index; the differentiated communication interaction instructions include the quantization accuracy parameter and the gradient retention ratio; the compressed encrypted transmission data packet includes the sparse position index, the binary quantization payload, and the integrity check code; and the global aggregation weight factor includes the contribution normalization coefficient and the node contribution ratio coefficient.

[0006] As a further aspect of the present invention, the specific steps of S1 are as follows: S101: Obtain a standard verification dataset of defect features of multiple types of industrial products, read the network topology configuration parameters of the real-time global model, initialize an independent model running container in the server memory according to the network topology configuration parameters, parse the encrypted gradient data from a single node, superimpose the weight update vector to update the global model parameters, and establish a temporary copy of the model to be tested. S102: Call the temporary model copy to be tested and the standard verification dataset, input the defective image samples in the verification dataset into the model copy in batches to perform forward inference operation, obtain the classification probability distribution vector of the output layer, extract the predicted category index corresponding to the peak probability, perform a consistency comparison operation between the predicted category index and the real category label associated with the sample, count the total number of consistent samples in the comparison result and calculate the percentage value of the total number of verification sets, and calculate the node target recognition accuracy. S103: Based on the node target recognition accuracy, obtain the basic recognition accuracy value of the global model after the previous round of aggregation, perform the subtraction difference operation between the node target recognition accuracy value and the basic recognition accuracy value, obtain the difference parameter representing the contribution of a single node, update the numerical index of the overall performance optimization of the model, and use the numerical index as a quantitative basis for measuring the node data quality and training contribution to obtain the marginal performance contribution.

[0007] As a further aspect of the present invention, the specific steps of S2 are as follows: S201: Call the marginal performance contribution degree, obtain the contribution judgment threshold stored in the central control unit in advance, perform a difference comparison operation between the marginal performance contribution degree value and the contribution judgment threshold value, determine the data value level of the real-time node according to the positive and negative sign attribute of the operation result, mark the case greater than the contribution judgment threshold as high value level, mark the case less than or equal to the contribution judgment threshold as low value level, and generate a contribution priority discrimination identifier. S202: Based on the contribution priority discrimination identifier, perform parameter mapping operation, extract 16-bit floating-point format for high-value level identifier as quantization precision parameter and match low sparsity truncation threshold that retains multiple features, extract four-bit integer format for low-value level identifier as quantization precision parameter and match high sparsity truncation threshold that retains few features, combine and bind the extracted precision parameter and truncation threshold to obtain transmission parameter set; S203: Perform protocol frame encapsulation processing on the transmission parameter set, fill the data segment position of the communication control instruction with the quantization precision parameter and sparsity truncation threshold, add the network addressing header information of the target edge node and the instruction integrity check code, construct a binary control data stream with hardware executability in accordance with the federated learning communication interaction protocol specification, and construct differentiated communication interaction instructions.

[0008] As a further aspect of the present invention, the specific steps of S3 are as follows: S301: Obtain the original gradient value matrix generated by the local edge computing node in the real-time training round. Based on the differentiated communication interaction command, parse the defined sparsity truncation threshold parameter, perform absolute value operation on the elements in the original gradient value matrix, perform numerical comparison operation between the operation result and the sparsity truncation threshold, remove redundant gradient terms whose absolute value is less than the sparsity truncation threshold, retain the gradient values ​​and corresponding matrix coordinate indices that exceed the sparsity truncation threshold, and generate a sparse gradient feature subset. S302: Obtain the quantization bit width configuration parameters, determine the discretization level range, calculate the quantization scaling factor, map the floating-point values ​​in the sparse gradient feature subset to discrete integers of a specified bit width, convert the discrete integers into a computer-recognizable bit stream sequence, and at the same time maintain the correspondence of gradient coordinate indices to obtain the gradient quantized binary stream. S303: For the gradient quantized binary stream, perform data encapsulation and security processing operations, use the public key generated by the asymmetric encryption algorithm to perform encryption operations on the main body of the binary stream, construct a transmission frame structure including encrypted payload, sparse matrix dimension information and integrity check bits, serialize the transmission frame structure into a payload form that conforms to the network transmission protocol, and establish a compressed encrypted transmission data packet.

[0009] As a further aspect of the present invention, the specific steps of S4 are as follows: S401: Using the compressed and encrypted data packet, obtain the marginal performance contribution of all participating nodes in this round, construct a numerical sequence list including the contribution indicators of all nodes, perform cumulative summation on all values, calculate the gain of the entire industrial federation network in the real-time training round, use the scalar value obtained by summation as the denominator of the normalization calculation, and obtain the total cumulative value of group contribution. S402: Call the marginal performance contribution of the node to be calculated and the total value of the group contribution, perform division normalization operation, calculate the percentage share of the single node contribution value in the overall group gain, define the calculated floating-point value as a quantitative indicator to measure the key role of the node in the global aggregation process, and calculate the node contribution ratio coefficient. S403: Based on the node contribution ratio coefficient, establish an index mapping relationship with the corresponding compressed encrypted transmission data packet source node, convert the node contribution ratio coefficient into a scalar multiplier for weighted update of model parameters, configure the application priority of the scalar multiplier in the aggregation matrix operation, define the mathematical weight of node data in the global model update direction, and generate a global aggregation weight factor.

[0010] As a further aspect of the present invention, the summation result is used as the denominator basis for the normalization calculation, ensuring that the normalization result is in percentage form; The operation of calculating the contribution value of a single node includes performing a division normalization operation based on the cumulative total value of the group contribution and the marginal performance contribution of the node to be calculated, wherein the quotient value obtained by the division operation is not greater than 1. The index mapping relationship is specifically a list sorted by node contribution, arranged from largest to smallest, to ensure that the model parameters of the nodes are updated in a weighted manner according to their contribution. The priority of the weighted update of the model parameters is limited by setting a threshold of 0.05, which means that when the node contribution ratio is lower than the threshold, the update priority is reduced.

[0011] As a further aspect of the present invention, the method further includes step S5: S5: Based on the global aggregation weight factor, perform multiplication weighting operation, perform cumulative summation operation on the weighted node gradient data, and use the summation value to adjust the real-time model connection weight state to construct a global model for industrial defect identification. The global model for industrial defect identification includes defect feature convolution kernels, a global connection weight matrix, and a classification decision logic layer.

[0012] As a further aspect of the present invention, the specific steps of S5 are as follows: S501: Using the global aggregation weight factor, combined with the compressed and encrypted data packets uploaded by the node, the server-side private key is used to perform decryption operations on the data packets, stripping the security protocol layer, parsing the decrypted binary payload and calling the quantization parameters to perform inverse quantization mapping operations, and filling the discrete values ​​back into the corresponding coordinate positions in the multidimensional tensor structure according to the sparse index information, performing zero-padding operations on the unfilled areas, restoring the complete dimension of the matrix, and generating the restored node gradient matrix. S502: Call the restored node gradient matrix and the global aggregation weight factor, perform tensor multiplication operation, adjust the magnitude of the single node gradient value, perform bit-by-bit summation operation on the weighted matrix of the participating nodes, merge the differential feature update directions fed back by the industrial nodes, and obtain the global aggregation update gradient. S503: Based on the global aggregated update gradient, read the global model basic parameter set of the real-time training round, set the learning rate hyperparameter of the model iteration, perform vector subtraction update operation between the basic parameter set and the global aggregated update gradient, and write the updated value into the connection weight storage unit of the neural network to complete the global adjustment of the model parameter space and construct a global model for industrial defect identification.

[0013] A data-secure training system for large industrial models under a federated learning architecture includes: The contribution evaluation module obtains a standard verification dataset of defect features of multiple types of industrial products, constructs a temporary model copy and loads a single encrypted gradient data, performs a recognition accuracy test on the temporary model copy, calculates the difference between the accuracy value obtained from the test and the basic accuracy value of the global model, and obtains the marginal performance contribution. The communication strategy module calls the marginal performance contribution and the preset contribution judgment threshold to perform a numerical comparison operation. For nodes whose marginal performance contribution is greater than the contribution judgment threshold, a 16-bit floating-point number width and a low sparsity truncation threshold that retains multiple features are matched. For nodes whose marginal performance contribution is less than the contribution judgment threshold, a 4-bit integer width and a high sparsity truncation threshold that retains fewer features are matched to construct differentiated communication interaction instructions. The data compression module acquires the original gradient value matrix generated by local training, parses the sparsity configuration item and bit width configuration item in the differentiated communication interaction instruction, filters out non-critical elements in the original gradient value matrix whose absolute value is less than the sparsity configuration item, and constructs a compressed and encrypted transmission data packet. The weight calculation module performs a normalization operation based on the compressed and encrypted transmission data packet and the marginal performance contribution, and calculates the proportion of the marginal performance contribution of a single node to the sum of the marginal performance contributions of all nodes to obtain the global aggregate weight factor. The model aggregation module calls the global aggregation weight factor to perform multiplication and weighting operations on the gradient data in the compressed and encrypted transmission data packet, performs accumulation and summation operations on the weighted node gradient data, adjusts the real-time model connection weight state, and constructs a global model for industrial defect identification.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, a differentiated communication and aggregation mechanism based on marginal performance contribution feedback is constructed. The standard validation set is used to accurately quantify the actual gain of each node's parameter updates on model performance. The quantization accuracy and sparsity ratio of each node's communication are dynamically adjusted according to the contribution level, realizing the intelligent tilting of bandwidth resources towards high-value features. While significantly reducing the overall communication load, the high-fidelity retention of key industrial defect features is ensured. Combined with a weighted aggregation strategy based on contribution value rather than sample quantity, the global model parameter update direction is accurately focused on the effective gradient that improves the ability to identify difficult defects. Under the premise of completely blocking the risk of original data going out of domain, the communication bottleneck is broken and the accuracy of the industrial large model in capturing complex fault features is greatly improved. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the steps of the present invention; Figure 2 This is a detailed schematic diagram of S1 of the present invention; Figure 3 This is a detailed schematic diagram of S2 of the present invention; Figure 4 This is a detailed schematic diagram of S3 of the present invention; Figure 5 This is a detailed schematic diagram of S4 of the present invention; Figure 6 This is a detailed schematic diagram of S5 of the present invention; Figure 7 This is a system module diagram of the present invention. Detailed Implementation

[0017] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0018] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0019] Please see Figure 1 This invention provides a method for secure training of industrial large-scale models under a federated learning architecture, comprising the following steps: S1: Obtain a standard verification dataset of defect features of multiple types of industrial products, construct a temporary model copy, use the standard verification dataset to perform a recognition accuracy test on the temporary model copy, calculate the difference between the accuracy value obtained from the test and the basic accuracy value of the global model, and obtain the marginal performance contribution. S2: Call the marginal performance contribution and perform a numerical comparison operation with the preset contribution judgment threshold. For nodes whose marginal performance contribution is greater than the contribution judgment threshold, match a 16-bit floating-point number width with a low sparsity truncation threshold that retains multiple features. For nodes whose marginal performance contribution is less than the contribution judgment threshold, match a 4-bit integer width with a high sparsity truncation threshold that retains fewer features. Construct differentiated communication interaction instructions. S3: Obtain the original gradient numerical matrix generated by local training, parse the sparsity configuration item and bit width configuration item in the differentiated communication interaction instruction, remove non-critical elements in the original gradient numerical matrix whose absolute value is less than the sparsity configuration item, call the bit width configuration item to map the retained gradient elements to binary codes of the corresponding length, and construct compressed encrypted transmission data packets. S4: Based on the compressed and encrypted data packets, perform normalization operations in conjunction with the marginal performance contribution, calculate the proportion of the marginal performance contribution of a single node to the sum of the marginal performance contributions of all nodes, establish a mapping relationship between the proportion and the aggregation weight identifier for the node, and obtain the global aggregation weight factor. S5: Based on the global aggregation weight factor, perform multiplication weighting operation, perform cumulative summation operation on the weighted node gradient data, use the summation value to adjust the real-time model connection weight state, and construct a global model for industrial defect identification. Marginal performance contribution includes accuracy gain and defect identification increment index; differentiated communication interaction instructions include quantization accuracy parameters and gradient retention ratio; compressed encrypted transmission data packets include sparse location index, binary quantization payload, and integrity check code; global aggregation weight factors include contribution normalization coefficient and node contribution ratio coefficient; and the global model for industrial defect identification includes defect feature convolution kernel, global connection weight matrix, and classification decision logic layer.

[0020] Please see Figure 2 The specific steps of S1 are as follows: S101: Obtain a standard verification dataset of defect features of multiple types of industrial products, read the network topology configuration parameters of the real-time global model, initialize an independent model running container in the server memory according to the network topology configuration parameters, parse the encrypted gradient data from a single node, superimpose the weight update vector to update the global model parameters, and establish a temporary copy of the model to be tested. High-resolution industrial cameras and laser contour sensor arrays deployed at the end of the production line are used to collect surface image data of various industrial products. The images cover four common defect types: scratches, dents, cracks, and discoloration. All raw images undergo image denoising preprocessing, specifically median filtering to remove salt-and-pepper noise and Z-score normalization to map pixel values ​​to a distribution range with a mean of 0 and a standard deviation of 1. Expert annotation is used to assign pixel-level or bounding box-level labels to the preprocessed images, constructing a standard validation dataset of 50,000 samples. The dataset is stratified according to defect type and severity. Simultaneously, the server receives encrypted gradient data from various industrial edge nodes via an encrypted communication channel. This data is generated locally by each node based on its private data. The server reads the network topology configuration parameters of the real-time global model, which define a deep convolutional neural network architecture. This architecture includes an input layer with dimensions set to 224 x 224 x 3 to adapt to RGB three-channel industrial images. Connected to the input layer is the first convolutional layer. The network consists of three layers: a first layer with 64 3x3 convolutional kernels, a stride of 1, and a ReLU activation function to introduce non-linear features; a second layer with 2x2 max pooling to reduce feature dimensionality; three stacked residual blocks in the middle, each consisting of two convolutional layers with 128 kernels each and skip connections, designed to address the vanishing gradient problem in deep networks; and a fully connected layer with 1024 neurons at the end, and a Softmax classification layer whose output dimension corresponds to the total number of defect categories. Configure parameters, allocate contiguous memory space in the high-speed video memory area of ​​the server, initialize an independent model running container, parse the encrypted gradient data from a single node, decrypt the data using a preset private key to restore the gradient values, extract the weight update vector containing the partial derivatives of the connection weights of each layer, and fill the floating-point values ​​in the weight update vector into the corresponding connection weight positions in the model running container one by one through tensor mapping operations, based on the layer index and coordinate information carried in the gradient data. Establish a temporary model copy in the server memory that reflects the latest training state of the edge node.

[0021] S102: Call the temporary model copy to be tested and the standard validation dataset, input the defective image samples in the validation dataset into the model copy in batches to perform forward inference operation, obtain the classification probability distribution vector of the output layer, extract the predicted category index corresponding to the peak probability, perform consistency comparison operation between the predicted category index and the real category label associated with the sample, count the total number of consistent samples in the comparison result and calculate the percentage value of the total validation set, and calculate the node target recognition accuracy. Defective image samples from the validation dataset are batched into the model replica in batches of 32. The model replica performs forward inference on each input image, with the signal sequentially passing through convolutional layers for feature extraction, pooling layers for dimensionality reduction, and fully connected layers for feature integration. The output layer generates a classification probability distribution vector, where each element represents the confidence score of the input image belonging to a specific defect category. The index of the element with the largest value is extracted as the predicted category index for that sample. This predicted category index is then compared with the true class associated with that sample in the standard validation dataset. A consistency comparison operation is performed on the labels. If the two values ​​are the same, it is determined that the identification is correct; otherwise, it is determined that the identification is incorrect. The entire validation dataset is traversed, and the total number of samples with consistent comparison results is counted. The total number of samples is divided by the total number of samples in the validation set to calculate the percentage value. The percentage value is defined as the node target identification accuracy. If the standard validation dataset contains 1000 samples, and the model copy correctly identifies 850 of them, then by dividing 850 by 1000, the node target identification accuracy is calculated to be 85%. Table 1 shows the parameter configuration and inference result statistics of each level of the model during a certain validation process. Table 1: Configuration and Validation Statistics of Temporary Model Replica Levels to be Tested

[0022] As shown in Table 1, the model architecture clarifies the processing logic at each level. Finally, inference is performed on the validation set based on this architecture, and the resulting node target recognition accuracy will serve as the core basis for subsequent evaluation of the node's contribution.

[0023] S103: Based on the node target recognition accuracy, obtain the basic recognition accuracy value of the global model after the previous round of aggregation, perform the subtraction difference operation between the node target recognition accuracy value and the basic recognition accuracy value, obtain the difference parameter representing the contribution of a single node, update the numerical index of the overall performance optimization of the model, and use the numerical index as a quantitative basis for measuring the node data quality and training contribution to obtain the marginal performance contribution. Retrieve the baseline recognition accuracy of the global model on the same standard validation dataset after the previous aggregation from the persistent storage unit. Perform a difference operation, that is, subtract the baseline recognition accuracy from the node's target recognition accuracy. The operation aims to quantify the gain or loss of the gradient update uploaded by the edge node on the model performance. The obtained difference is a numerical indicator characterizing the extent of the optimization of the overall model performance by the single node parameter update. This numerical indicator is directly used as a quantitative basis for measuring the node's data quality and training contribution, defined as the marginal performance contribution. If the difference is positive, it indicates that the update of the node has improved the model performance; if the difference is negative or zero, it indicates that the node's data has noise or distribution bias. Assuming that the baseline recognition accuracy of the global model in the previous round is 82.5%, and the target recognition accuracy of the current node is 85.0%, then by subtracting 85.0% from 82.5%, the difference is 2.5 percentage points, that is, the marginal performance contribution of the node is 0.025.

[0024] Please see Figure 3 The specific steps of S2 are as follows: S201: Call the marginal performance contribution, obtain the contribution judgment threshold pre-stored in the central control unit, perform a difference comparison operation between the marginal performance contribution value and the contribution judgment threshold value, determine the data value level of the real-time node based on the positive or negative sign attribute of the operation result, mark the case greater than the contribution judgment threshold as high value level, mark the case less than or equal to the contribution judgment threshold as low value level, and generate a contribution priority discrimination identifier. The pre-stored contribution judgment threshold is retrieved from the configuration register of the central control unit. This threshold is not arbitrarily set, but is based on the contribution distribution data of the training rounds and is set to 50% of the average contribution value through statistical analysis to ensure the dynamic adaptability of the selection criteria. In this implementation, the contribution judgment threshold is set to 0.01 through experimental statistics. A numerical comparison operation is performed, comparing the marginal performance contribution value of the current node with the contribution judgment threshold value. The data value level of the real-time node is determined based on the positive or negative sign attribute of the operation result: if the result of the marginal performance contribution value minus the contribution judgment threshold is positive, that is, the marginal performance contribution value is greater than 0.01, the node is marked as high value level, indicating that the node data has a significant positive effect on model optimization; if the operation result is negative or zero, that is, the marginal performance contribution value is less than or equal to 0.01, the node is marked as low value level. A corresponding contribution priority discrimination label is generated based on this judgment result. For example, a binary bit "1" is used to represent a high value level and "0" is used to represent a low value level.

[0025] S202: Based on the contribution priority identification, perform parameter mapping operation, extract 16-bit floating-point format for high-value level identification as quantization precision parameter and match it with a low sparsity truncation threshold that retains multiple features, extract four-bit integer format for low-value level identification as quantization precision parameter and match it with a high sparsity truncation threshold that retains few features, combine and bind the extracted precision parameter and truncation threshold to obtain the transmission parameter set; For high-value levels marked "1", a 16-bit floating-point format is extracted as the quantization precision parameter. This format conforms to the IEEE 754 half-precision standard, preserving a high data dynamic range. Simultaneously, a low sparsity truncation threshold, such as 0.001, is applied. This low sparsity truncation threshold aims to preserve multiple features, allowing more gradients with smaller amplitudes to pass through the filter, maximizing the retention of detailed information from high-value nodes. For low-value levels marked "0", a four-bit integer format is extracted as the quantization precision parameter. This format compresses the data to only 16 discrete quantization levels. Simultaneously, a high sparsity truncation threshold, such as 0.005, is applied. This high sparsity truncation threshold aims to preserve multiple features, allowing more gradients with smaller amplitudes to pass through the filter, maximizing the retention of detailed information from high-value nodes. By retaining only a few features, i.e., significantly eliminating gradient terms with small magnitudes, the communication bandwidth consumption is significantly reduced. This is achieved by combining and binding the extracted precision parameters with truncation thresholds. For example, for a high-value node with a marginal performance contribution of 0.025, its transmission parameter set is {precision: 16-bit floating point, sparsity threshold: 0.001}; while for a low-value node with a marginal performance contribution of 0.005, its transmission parameter set is {precision: 4-bit integer, sparsity threshold: 0.005}. This differentiated configuration ensures that resources are tilted towards high-value data to obtain the transmission parameter set.

[0026] S203: Perform protocol frame encapsulation processing on the transmission parameter set, fill the quantization precision parameter and sparsity truncation threshold into the data segment position of the communication control command, add the network addressing header information of the target edge node and the command integrity check code, construct a binary control data stream with hardware executability according to the federated learning communication interaction protocol specification, and construct differentiated communication interaction commands. In the control command data segment of the communication protocol frame, the specific values ​​of the quantization precision parameter (e.g., representing 16-bit or 4-bit encoding) and the sparsity truncation threshold are filled in respectively. The network addressing header information of the target edge node is added, which includes the node's IP address and port number to ensure accurate delivery of the command. The integrity check code of the command content is calculated, and a 32-bit check value is generated using the cyclic redundancy check algorithm and appended to the end of the frame to prevent transmission errors. According to the federated learning communication interaction protocol specification, the above fields are concatenated to construct a binary control data stream with hardware executable capability and generate differentiated communication interaction commands.

[0027] Please see Figure 4 The specific steps of S3 are as follows: S301: Obtain the original gradient value matrix generated by the local edge computing node in the real-time training round. Based on the differentiated communication interaction command, parse the defined sparsity truncation threshold parameter, perform absolute value operation on the elements in the original gradient value matrix, perform numerical comparison operation between the operation result and the sparsity truncation threshold, remove redundant gradient terms with absolute values ​​less than the sparsity truncation threshold, retain the gradient values ​​and corresponding matrix coordinate indices that exceed the sparsity truncation threshold, and generate a sparse gradient feature subset. The original gradient value matrix generated by the local edge computing node during real-time training rounds using local private data through the backpropagation algorithm is obtained. The dimension of this matrix is ​​consistent with the dimension of the model parameters. Based on the received differential communication interaction instructions, the sparsity truncation threshold parameter defined therein is parsed. If the parsed sparsity truncation threshold is 0.005, the absolute value operation is performed on each element in the original gradient value matrix one by one. The absolute value result is compared with the sparsity truncation threshold of 0.005. If the absolute value of a gradient element is less than 0.005, it is determined to be a redundant gradient term and is removed from the matrix or set to zero. If the absolute value of a gradient element is greater than or equal to 0.005, it is retained. Only the gradient values ​​exceeding the sparsity truncation threshold and their corresponding matrix coordinate indices (such as row and column numbers) are retained to generate a sparse gradient feature subset.

[0028] S302: Obtain the quantization bit width configuration parameters, determine the discretization level range, calculate the quantization scaling factor, map the floating-point values ​​in the sparse gradient feature subset to discrete integers of the specified bit width, convert the discrete integers into a computer-recognizable bit stream sequence, and at the same time maintain the correspondence of gradient coordinate indices to obtain the gradient quantized binary stream. If the configuration parameter is set to a 4-bit integer format, the discretization level range is 0 to 15. The maximum and minimum absolute values ​​in the sparse gradient feature subset are scanned, and the quantization scaling factor is calculated. This factor is equal to the difference between the maximum and minimum absolute values ​​divided by the total number of discretization levels (i.e., 15). Each floating-point gradient value in the subset is mapped to a discrete integer of a specified bit width. The mapping logic is as follows: the gradient value is divided by the quantization scaling factor, and the result is rounded to the nearest integer. If the quantization scaling factor is 0.01 and a gradient value is 0.052, it is mapped to the integer 5. The discrete integer is converted into a 4-bit bitstream sequence that can be recognized by the computer. At the same time, the correspondence of the gradient coordinate index must be maintained. The coordinate index is stored using a compressed row storage format, resulting in a gradient quantized binary stream containing compressed weights and index information.

[0029] S303: For gradient quantized binary streams, perform data encapsulation and security processing operations, use the public key generated by the asymmetric encryption algorithm to perform encryption operations on the main body of the binary stream, construct a transmission frame structure including encrypted payload, sparse matrix dimension information and integrity check bits, serialize the transmission frame structure into a payload form that conforms to the network transmission protocol, and establish a compressed encrypted transmission data packet. Asymmetric encryption is used to exchange session keys, and then symmetric encryption algorithms (such as AES) are used to encrypt the data body. Encryption operations are performed on the binary stream body to convert plaintext data into unreadable ciphertext, preventing data from being stolen or tampered with during transmission. A transmission frame structure is constructed, which sequentially includes: frame header synchronization word, encrypted payload (i.e., encrypted gradient data), sparse matrix dimension information (used by the receiving end to reconstruct the matrix shape), and integrity check bits. Serialization technology is used to convert the transmission frame structure into a payload form that conforms to the TCP / IP network transmission protocol, and a compressed encrypted transmission data packet is established.

[0030] Please see Figure 5 The specific steps of S4 are as follows: S401: Utilize compressed and encrypted data transmission to obtain the marginal performance contribution of all participating nodes in this round, construct a numerical sequence list including the contribution indicators of all nodes, perform cumulative summation on all values, calculate the gain magnitude of the entire industrial federation network in the real-time training round, use the summed scalar value as the denominator of the normalization calculation, and obtain the total cumulative value of the group contribution. By utilizing the identification information carried in the compressed and encrypted data packets uploaded by each node, the marginal performance contribution of the participating nodes in this round is obtained. A numerical sequence list containing the contribution indicators of all nodes is constructed. The values ​​in the list are summed to calculate the gain of the entire industrial federation network in the real-time training round. The scalar value obtained by summing is used as the denominator for normalization calculation to obtain the total value of the group contribution. Assuming that there are three nodes A, B and C participating in the training in the network, their marginal performance contributions are 0.025, 0.015 and 0.010, respectively, Table 2 shows the contribution data of these three nodes. Table 2: Statistics on the Marginal Performance Contribution of Nodes in the Federated Network

[0031] As shown in Table 2, by summing these three values, namely 0.025 plus 0.015 plus 0.010, the total cumulative value of the group contribution is calculated to be 0.050. This total value reflects the overall contribution of the node in this round to the improvement of model performance.

[0032] S402: Call the marginal performance contribution of the node to be calculated and the total sum of the group contribution, perform division normalization operation, calculate the percentage share of the single node contribution value in the overall group gain, define the floating-point value obtained by the calculation as a quantitative indicator to measure the key role of the node in the global aggregation process, and obtain the node contribution ratio coefficient. The percentage share of a single node's contribution in the overall group gain is calculated by dividing its marginal performance contribution by the total accumulated group contribution. The resulting floating-point value is defined as a quantitative indicator measuring the node's criticality in the global aggregation process, i.e., the node contribution ratio coefficient. Continuing with the previous example, for node A, its marginal performance contribution of 0.025 is divided by the total accumulated group contribution of 0.050, resulting in 0.5. Similarly, the coefficient for node B is 0.015 divided by 0.050, which equals 0.3; and the coefficient for node C is 0.010 divided by 0.050, which equals 0.2. The sum of these three coefficients (0.5, 0.3, 0.2) is strictly equal to 1, achieving normalization of the contribution. The advantage of this calculation logic is that it dynamically allocates aggregation weights based on the actual data quality of the nodes, rather than simply using average aggregation, thus enhancing the model's sensitivity to high-quality data and obtaining the node contribution ratio coefficient.

[0033] S403: Based on the node contribution ratio coefficient, establish an index mapping relationship with the corresponding compressed encrypted transmission data packet source node, convert the node contribution ratio coefficient into a scalar multiplier for weighted update of model parameters, configure the application priority of the scalar multiplier in the aggregation matrix operation, define the mathematical weight of node data in the global model update direction, and generate a global aggregation weight factor. The node contribution ratio coefficient is directly converted into a scalar multiplier for weighted updating of model parameters. The application priority of this scalar multiplier in the aggregation matrix operation is configured to ensure that in subsequent aggregation steps, this coefficient is used first to adjust the magnitude of the gradient of the corresponding node. The mathematical weight of the node data in the global model update direction is defined. That is, this coefficient determines the step size ratio of the global model parameters moving in the gradient direction of the node. For example, if the global aggregation weight factor of node A is determined to be 0.5, it means that the gradient information of node A will account for 50% of the weight in the global model update, dominating the direction of model optimization, and generating the global aggregation weight factor.

[0034] Please see Figure 6 The specific steps of S5 are as follows: S501: It adopts a global aggregated weight factor, combined with the compressed and encrypted data packets uploaded by the node, uses the server-side private key to perform decryption operation on the data packets, strips the security protocol layer, parses the decrypted binary payload and calls the quantization parameters to perform inverse quantization mapping operation, fills the discrete values ​​back into the corresponding coordinate positions in the multidimensional tensor structure according to the sparse index information, performs zero-padding operation on the unfilled areas, restores the complete dimension of the matrix, and generates the restored node gradient matrix. Using the private key pre-stored on the server (paired with the public key on the node), the encrypted payload of the data packet is decrypted, the security protocol layer is stripped away, and the gradient quantized binary stream is restored. The decrypted binary payload is parsed, the quantization bit width and scaling factor are extracted, and the quantization parameters are called to perform the dequantization mapping operation. The discrete integers are multiplied by the quantization scaling factor to restore them to approximate floating-point gradient values. Based on the sparse index information (such as row and column coordinates) carried in the binary stream, the restored discrete values ​​are filled back into the corresponding coordinate positions in the multidimensional tensor structure. Since a sparse matrix is ​​transmitted, for coordinate positions that do not appear in the index, i.e., unfilled areas, zero-padding is performed to uniformly set the values ​​at those positions to 0. Through this process, the complete dimensions of the matrix (e.g., 224x224x64) are restored, generating the restored node gradient matrix.

[0035] S502: Call the restored node gradient matrix and global aggregation weight factor, perform tensor multiplication operation, adjust the magnitude of the single node gradient value, perform bit-by-bit summation operation on the weighted matrix of participating nodes, merge the differential feature update directions fed back by industrial nodes, and obtain the global aggregation update gradient. The magnitude of the gradient value of a single node is adjusted by multiplying the value of each gradient element in the matrix by the global aggregation weight factor corresponding to that node. This achieves weighted processing of data of different quality. For the weighted matrix of participating nodes, a bitwise summation operation is performed. That is, for any element at the same coordinate position in the matrix, the weighted gradient values ​​of the node at that position are added together. If at position (i, j), the weighted gradient of node A is 0.01, that of node B is 0.005, and that of node C is -0.002, then the combined gradient value is 0.013. By merging the differential feature update directions fed back by industrial nodes, the global aggregation update gradient is obtained.

[0036] S503: Based on the global aggregation update gradient, read the global model basic parameter set of the real-time training round, set the learning rate hyperparameter of the model iteration, perform vector subtraction update operation between the basic parameter set and the global aggregation update gradient, and write the updated value into the connection weight storage unit of the neural network to complete the global adjustment of the model parameter space and build a global model for industrial defect identification. The system reads the global model's basic parameter set from the real-time training rounds, i.e., the connection weights and biases of the current global network. It sets the learning rate hyperparameter for model iteration, determined through grid search experiments (e.g., 0.01), to control the step size of parameter updates and prevent oscillations. It then performs a vector subtraction update operation between the basic parameter set and the global aggregated update gradient. The logic is: the new model parameter equals the old model parameter minus (learning rate multiplied by the global aggregated update gradient). For example, if the old value of a weight is 0.5, the learning rate is 0.01, and the corresponding global aggregated update gradient is 0.1, then the updated weight value is 0.5 minus 0.001, i.e., 0.499. The updated value is then written into the neural network's connection weight storage unit, completing the global adjustment of the model parameter space and constructing the industrial defect recognition global model optimized through this round of federated learning iterations.

[0037] Please see Figure 7 A data-secure training system for large industrial models under a federated learning architecture, including: The contribution evaluation module obtains a standard verification dataset of defect features of various industrial products, constructs a temporary model copy, performs a recognition accuracy test on the temporary model copy, calculates the difference between the accuracy value obtained from the test and the basic accuracy value of the global model, and obtains the marginal performance contribution. The communication strategy module calls the marginal performance contribution and performs a numerical comparison operation with the preset contribution judgment threshold. For nodes whose marginal performance contribution is greater than the contribution judgment threshold, it matches a 16-bit floating-point number with a low sparsity truncation threshold that retains multiple features. For nodes whose marginal performance contribution is less than the contribution judgment threshold, it matches a 4-bit integer with a high sparsity truncation threshold that retains fewer features, thus constructing differentiated communication interaction instructions. The data compression module obtains the original gradient numerical matrix generated by local training, parses the sparsity configuration item and bit width configuration item in the differentiated communication interaction command, filters the non-critical elements in the original gradient numerical matrix whose absolute value is less than the sparsity configuration item, and constructs a compressed and encrypted transmission data packet. The weight calculation module performs a normalization operation based on the compressed and encrypted transmission data packet and the marginal performance contribution, and calculates the proportion of the marginal performance contribution of a single node to the sum of the marginal performance contributions of all nodes to obtain the global aggregate weight factor. The model aggregation module calls the global aggregation weight factor to perform multiplication and weighting operations on the gradient data in the compressed and encrypted transmission data packet, performs accumulation and summation operations on the weighted node gradient data, adjusts the real-time model connection weight state, and constructs a global model for industrial defect identification.

[0038] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of protection of the described technical solutions.

Claims

1. A data-secure training method for large industrial models under a federated learning architecture, characterized in that, Includes the following steps: S1: Obtain a standard verification dataset of defect features of multiple types of industrial products, construct a temporary model copy, perform a recognition accuracy test on the temporary model copy, and obtain the marginal performance contribution. S2: Call the marginal performance contribution and perform a numerical comparison operation with the preset contribution judgment threshold. For nodes whose marginal performance contribution is greater than the contribution judgment threshold, match a 16-bit floating-point number width with a low sparsity truncation threshold that retains multiple features. For nodes whose marginal performance contribution is less than the contribution judgment threshold, match a 4-bit integer width with a high sparsity truncation threshold that retains fewer features. Construct differentiated communication interaction instructions. S3: Obtain the original gradient value matrix generated by local training, parse the sparsity configuration item and bit width configuration item in the differentiated communication interaction instruction, remove non-key elements in the original gradient value matrix whose absolute value is less than the sparsity configuration item, and construct a compressed encrypted transmission data packet. S4: Based on the compressed and encrypted transmission data packet, calculate the proportion of the marginal performance contribution of a single node to the sum of the marginal performance contributions of all nodes, and obtain the global aggregation weight factor.

2. The method for secure training of industrial large-scale models under a federated learning architecture according to claim 1, characterized in that, The marginal performance contribution includes the accuracy gain value and the defect identification increment index; the differentiated communication interaction instructions include the quantization accuracy parameter and the gradient retention ratio; the compressed encrypted transmission data packet includes the sparse position index, the binary quantization payload, and the integrity check code; and the global aggregation weight factor includes the contribution normalization coefficient and the node contribution ratio coefficient.

3. The method for secure training of industrial large-scale models under a federated learning architecture according to claim 1, characterized in that, The specific steps of S1 are as follows: S101: Obtain a standard verification dataset of defect features of multiple types of industrial products, read the network topology configuration parameters of the real-time global model, initialize an independent model running container in the server memory according to the network topology configuration parameters, parse the encrypted gradient data from a single node, superimpose the weight update vector to update the global model parameters, and establish a temporary copy of the model to be tested. S102: Call the temporary model copy to be tested and the standard verification dataset, input the defective image samples in the verification dataset into the model copy in batches to perform forward inference operation, obtain the classification probability distribution vector of the output layer, extract the predicted category index corresponding to the peak probability, perform a consistency comparison operation between the predicted category index and the real category label associated with the sample, count the total number of consistent samples in the comparison result and calculate the percentage value of the total number of verification sets, and calculate the node target recognition accuracy. S103: Based on the node target recognition accuracy, obtain the basic recognition accuracy value of the global model after the previous round of aggregation, perform the subtraction difference operation between the node target recognition accuracy value and the basic recognition accuracy value, obtain the difference parameter representing the contribution of a single node, update the numerical index of the overall performance optimization of the model, and use the numerical index as a quantitative basis for measuring the node data quality and training contribution to obtain the marginal performance contribution.

4. The method for secure training of industrial large-scale models under the federated learning architecture according to claim 3, characterized in that, The specific steps of S2 are as follows: S201: Call the marginal performance contribution degree, obtain the contribution judgment threshold stored in the central control unit in advance, perform a difference comparison operation between the marginal performance contribution degree value and the contribution judgment threshold value, determine the data value level of the real-time node according to the positive and negative sign attribute of the operation result, mark the case greater than the contribution judgment threshold as high value level, mark the case less than or equal to the contribution judgment threshold as low value level, and generate a contribution priority discrimination identifier. S202: Based on the contribution priority discrimination identifier, perform parameter mapping operation, extract 16-bit floating-point format for high-value level identifier as quantization precision parameter and match low sparsity truncation threshold that retains multiple features, extract four-bit integer format for low-value level identifier as quantization precision parameter and match high sparsity truncation threshold that retains few features, combine and bind the extracted precision parameter and truncation threshold to obtain transmission parameter set; S203: Perform protocol frame encapsulation processing on the transmission parameter set, fill the data segment position of the communication control instruction with the quantization precision parameter and sparsity truncation threshold, add the network addressing header information of the target edge node and the instruction integrity check code, construct a binary control data stream with hardware executability in accordance with the federated learning communication interaction protocol specification, and construct differentiated communication interaction instructions.

5. The method for secure training of industrial large-scale models under the federated learning architecture according to claim 4, characterized in that, The specific steps for S3 are as follows: S301: Obtain the original gradient value matrix generated by the local edge computing node in the real-time training round. Based on the differentiated communication interaction command, parse the defined sparsity truncation threshold parameter, perform absolute value operation on the elements in the original gradient value matrix, perform numerical comparison operation between the operation result and the sparsity truncation threshold, remove redundant gradient terms whose absolute value is less than the sparsity truncation threshold, retain the gradient values ​​and corresponding matrix coordinate indices that exceed the sparsity truncation threshold, and generate a sparse gradient feature subset. S302: Obtain the quantization bit width configuration parameters, determine the discretization level range, calculate the quantization scaling factor, map the floating-point values ​​in the sparse gradient feature subset to discrete integers of a specified bit width, convert the discrete integers into a computer-recognizable bit stream sequence, and at the same time maintain the correspondence of gradient coordinate indices to obtain the gradient quantized binary stream. S303: For the gradient quantized binary stream, perform data encapsulation and security processing operations, use the public key generated by the asymmetric encryption algorithm to perform encryption operations on the main body of the binary stream, construct a transmission frame structure including encrypted payload, sparse matrix dimension information and integrity check bits, serialize the transmission frame structure into a payload form that conforms to the network transmission protocol, and establish a compressed encrypted transmission data packet.

6. The method for secure training of industrial large-scale models under a federated learning architecture according to claim 5, characterized in that, The specific steps of S4 are as follows: S401: Using the compressed and encrypted data packet, obtain the marginal performance contribution of all participating nodes in this round, construct a numerical sequence list including the contribution indicators of all nodes, perform cumulative summation on all values, calculate the gain of the entire industrial federation network in the real-time training round, use the scalar value obtained by summation as the denominator of the normalization calculation, and obtain the total cumulative value of group contribution. S402: Call the marginal performance contribution of the node to be calculated and the total value of the group contribution, perform division normalization operation, calculate the percentage share of the single node contribution value in the overall group gain, define the calculated floating-point value as a quantitative indicator to measure the key role of the node in the global aggregation process, and calculate the node contribution ratio coefficient. S403: Based on the node contribution ratio coefficient, establish an index mapping relationship with the corresponding compressed encrypted transmission data packet source node, convert the node contribution ratio coefficient into a scalar multiplier for weighted update of model parameters, configure the application priority of the scalar multiplier in the aggregation matrix operation, define the mathematical weight of node data in the global model update direction, and generate a global aggregation weight factor.

7. The method for secure training of industrial large-scale models under a federated learning architecture according to claim 6, characterized in that, The summation result is used as the denominator for the normalization calculation, ensuring that the normalization result is in percentage form; The operation of calculating the contribution value of a single node includes performing a division normalization operation based on the cumulative total value of the group contribution and the marginal performance contribution of the node to be calculated, wherein the quotient value obtained by the division operation is not greater than 1. The index mapping relationship is specifically a list sorted by node contribution, arranged from largest to smallest, to ensure that the model parameters of the nodes are updated in a weighted manner according to their contribution. The priority of the weighted update of the model parameters is limited by setting a threshold of 0.05, which means that when the node contribution ratio is lower than the threshold, the update priority is reduced.

8. The method for secure training of industrial large-scale models under a federated learning architecture according to claim 1, characterized in that, The method further includes step S5: S5: Based on the global aggregation weight factor, perform multiplication weighting operation, perform cumulative summation operation on the weighted node gradient data, and use the summation value to adjust the real-time model connection weight state to construct a global model for industrial defect identification. The global model for industrial defect identification includes defect feature convolution kernels, a global connection weight matrix, and a classification decision logic layer.

9. The method for secure training of industrial large-scale models under a federated learning architecture according to claim 8, characterized in that, The specific steps of S5 are as follows: S501: Using the global aggregation weight factor, combined with the compressed and encrypted data packets uploaded by the node, the server-side private key is used to perform decryption operations on the data packets, stripping the security protocol layer, parsing the decrypted binary payload and calling the quantization parameters to perform inverse quantization mapping operations, and filling the discrete values ​​back into the corresponding coordinate positions in the multidimensional tensor structure according to the sparse index information, performing zero-padding operations on the unfilled areas, restoring the complete dimension of the matrix, and generating the restored node gradient matrix. S502: Call the restored node gradient matrix and the global aggregation weight factor, perform tensor multiplication operation, adjust the magnitude of the single node gradient value, perform bit-by-bit summation operation on the weighted matrix of the participating nodes, merge the differential feature update directions fed back by the industrial nodes, and obtain the global aggregation update gradient. S503: Based on the global aggregated update gradient, read the global model basic parameter set of the real-time training round, set the learning rate hyperparameter of the model iteration, perform vector subtraction update operation between the basic parameter set and the global aggregated update gradient, and write the updated value into the connection weight storage unit of the neural network to complete the global adjustment of the model parameter space and construct a global model for industrial defect identification.

10. A data security training system for large industrial models under a federated learning architecture, characterized in that: The system is used to implement the industrial large model data security training method under the federated learning architecture according to any one of claims 1-9, and the system includes: The contribution evaluation module obtains a standard verification dataset of defect features of multiple types of industrial products, constructs a temporary model copy and loads a single encrypted gradient data, performs a recognition accuracy test on the temporary model copy, calculates the difference between the accuracy value obtained from the test and the basic accuracy value of the global model, and obtains the marginal performance contribution. The communication strategy module calls the marginal performance contribution degree and the preset contribution judgment threshold to perform a numerical comparison operation. Nodes with a marginal performance contribution degree greater than the contribution judgment threshold are matched with a 16-bit floating-point number width and a low sparsity truncation threshold that retains multiple features. For nodes with a marginal performance contribution degree less than the contribution judgment threshold, they are matched with a 4-bit integer width and a high sparsity truncation threshold that retains fewer features, thus constructing differentiated communication interaction instructions. The data compression module acquires the original gradient value matrix generated by local training, parses the sparsity configuration item and bit width configuration item in the differentiated communication interaction instruction, filters out non-critical elements in the original gradient value matrix whose absolute value is less than the sparsity configuration item, and constructs a compressed and encrypted transmission data packet. The weight calculation module performs a normalization operation based on the compressed and encrypted transmission data packet and the marginal performance contribution, and calculates the proportion of the marginal performance contribution of a single node to the sum of the marginal performance contributions of all nodes to obtain the global aggregate weight factor. The model aggregation module calls the global aggregation weight factor to perform multiplication and weighting operations on the gradient data in the compressed and encrypted transmission data packet, performs accumulation and summation operations on the weighted node gradient data, adjusts the real-time model connection weight state, and constructs a global model for industrial defect identification.