A method for accurate matching analysis of data based on machine learning
Through machine learning and distributed parallel processing technology, the problem of unreasonable data segmentation and node allocation is solved, efficient and accurate processing of large-scale data is achieved, and the performance and reliability of the system are improved.
Patent Information
- Application Number
- CN202411472045.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-10-22
AI Technical Summary
The existing data matching analysis methods have irrationality in data segmentation and node allocation, resulting in inadequate integration of data parallelism and model parallelism, resulting in synchronization problems and performance bottlenecks, which cannot meet the requirements of efficiency and accuracy in the big data era.
The precise matching analysis method based on machine learning is adopted, and the sub-data set is allocated according to the performance and data characteristics of the computing nodes, and the distributed synchronization mechanism is used to merge results and adjust tasks to ensure coordinated work of each node.
It significantly improves the efficiency and speed of data processing, ensures the accuracy and reliability of analysis results, avoids processing bottlenecks and data loss, and improves the stability and reliability of the system.
Smart Images

Figure CN119473575B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data matching, and specifically provides a method for accurate data matching analysis based on machine learning. Background Art
[0002] With the rapid development of information technology and the advent of the big data era, data has become an indispensable core resource in enterprise operation, scientific research, and government decision-making. The scale, variety, and generation speed of data are increasing at an unprecedented rate, which poses extremely high requirements for data processing capabilities. The traditional centralized data processing mode is limited by computing resources, storage capacity, and network bandwidth, and it is difficult to meet the real-time processing needs of massive data. Therefore, the distributed computing architecture has gradually become the mainstream choice for processing big data.
[0003] Currently, in terms of data segmentation, existing data matching analysis methods often adopt simple random segmentation and segmentation based on fixed rules, lacking in-depth analysis of data characteristics, which easily leads to data imbalance between sub-datasets, affecting the efficiency and accuracy of subsequent processing. Moreover, in terms of node allocation, existing technologies usually only consider a single performance indicator of computing nodes, resulting in an unreasonable allocation scheme, making the combination of data parallelism and model parallelism not tight enough, and causing synchronization problems and performance bottlenecks.
[0004] In summary, there are obvious deficiencies in the accurate matching analysis of large-scale data in the existing technology, which cannot meet the requirements of high efficiency, accuracy, and reliability of data processing in the big data era. Therefore, there is an urgent need for a method for accurate data matching analysis based on machine learning, which can make full use of distributed parallel processing technology, overcome the shortcomings of the existing technology, and achieve fast and accurate processing of large-scale data. Summary of the Invention
[0005] The purpose of the present invention is to make up for the deficiencies of the existing technology and provide a method for accurate data matching analysis based on machine learning. It can use intelligent data segmentation algorithms and feature extraction methods to better adapt to different types and structures of data. Whether it is text, numerical, or image data, the system can perform effective preprocessing and feature engineering according to its characteristics, thereby improving the support ability for diverse data sources. In addition, the machine learning-based method allows the model to self-optimize as new data is continuously input, which enables the system to continuously improve its performance to cope with changing data environments and business requirements.
[0006] To solve the above technical problems, the present invention provides the following technical solution: A method for accurate data matching analysis based on machine learning, and the specific steps of this method are as follows:
[0007] S100. Split the data to be processed into multiple sub - data sets, and the size of each sub - data set is divided according to the performance metrics and data characteristics of the computing nodes;
[0008] Perform real - time monitoring on the performance metrics of each computing node. The monitored metrics include CPU utilization U cpu , memory occupancy rate U mem , network bandwidth B, and disk I / O speed S io . Define the computing node performance index P, that is where w1, w2, w3, and w4 are the weight coefficients of CPU utilization, memory occupancy rate, network bandwidth, and disk I / O speed respectively, and w1 + w2 + w3 + w4 = 1, B max is the maximum value of the network bandwidth, and S io,max is the maximum value of the disk I / O speed;
[0009] For each sub - data set of the data characteristics, analyze its size S data , data type complexity C type and processing difficulty D proc . Define the allocation index A, that is where α is a weight parameter used to balance the influence of the computing node performance P and the sub - data set size A on the allocation. When α≈1, the allocation is more inclined to allocate the sub - data set to the computing node with a high performance index P value. When α≈0, the allocation pays more attention to the matching between the size of the sub - data set and the computing node performance;
[0010] S200. Allocate each sub - data set to multiple computing nodes, and each computing node independently performs operations such as data pre - processing, feature extraction, and model training;
[0011] S300. On each computing node, process the allocated sub - data set, where:
[0012] The pre - processing stage includes data cleaning and normalization operations;
[0013] In the feature extraction stage, extract effective features according to the characteristics of the data and the target of matching analysis;
[0014] In the model training stage, use the extracted features to train the machine learning model;
[0015] S400. Through a distributed synchronization mechanism, merge and summarize the results on each computing node;
[0016] S500. After merging and summarizing, obtain the final data matching analysis result.
[0017] Furthermore, during the S200 node allocation process for each sub-dataset, all computing nodes are traversed, the allocation index A between each computing node and the sub-dataset is calculated, and the computing node with the largest allocation index A is selected as the allocation target for the sub-dataset.
[0018] Furthermore, the S200 performs real-time performance monitoring on each computing node, including CPU utilization U cpu , memory occupancy U mem , network bandwidth B, and disk I / O speed S io . At the same time, the size S data , data type complexity C type , and processing difficulty D proc of each sub-dataset are analyzed to form a sub-dataset feature vector, and the sub-dataset feature vector is matched with the computing node performance metric P. Specifically:
[0019] Sub-datasets with high processing difficulty and complex data types are allocated to computing nodes with high performance metric values P;
[0020] Small-scale sub-datasets are allocated to computing nodes with low performance metric values P and idle resources to fully utilize computing resources.
[0021] Furthermore, during the model training process of the S300, a distributed training algorithm is adopted, namely a method that combines data parallelism and model parallelism. Among them:
[0022] In data parallelism, the data is split into multiple small batches and allocated to different computing nodes for training. Each computing node independently updates the model parameters;
[0023] In model parallelism, different parts of the model are allocated to different computing nodes for parallel computing.
[0024] Furthermore, for data parallelism with N computing nodes and a total dataset D, the dataset is split into K small batches, and the size of each small batch is |D| / K. In data parallelism, each computing node i updates the model parameters based on the small batch data it is allocated. That is, if the model parameter is θ and the loss function is L(θ), the parameter update formula on computing node i is: where, represents the model parameter of computing node i at the t-th iteration, η is the learning rate, is the gradient of the loss function at the current parameter. After one round of iteration is completed, the parameters of each computing node are synchronized, and the method of parameter averaging is adopted, that is: enabling all computing nodes to use the synchronized parameters for training in the new round of iteration.
[0025] Furthermore, in the model parallelism, the model can be divided into M parts and distributed to M computing nodes respectively. For an input sample x, the forward propagation process of the model is expressed as: y = f M (f M-1 (…f2(f1(x)))), where f i represents the computing function of the i-th part of the model. During the backpropagation process, each computing node i calculates the gradient of the part it is responsible for and passes the gradient to the previous computing node for further gradient calculation to obtain the gradient of the entire model. That is, if the loss function is L(y), the gradient of the model is calculated by the chain rule: Each computing node i is responsible for calculating and where θ i is the model parameter that the computing node i is responsible for. When updating the parameters, each computing node i updates the parameters according to the gradient it calculates:
[0026] Furthermore, during the model training process, the S300 combines data parallelism and model parallelism. It divides the dataset into multiple small batches and distributes them to different computing nodes for data parallel training. At the same time, it divides the model into multiple parts and distributes them to different computing nodes for model parallel computing. In each round of iteration, each computing node first performs parameter update of data parallelism, and then performs gradient calculation and parameter update of model parallelism. Through the parameter synchronization mechanism, the parameters of each computing node are merged and updated. That is, initialize the model parameter θ, divide the dataset into small batches and distribute them to different computing nodes. For each round of iteration t, perform data parallel update and model parallel gradient calculation. The data parallel update updates the parameters of data parallelism for each computing node i according to the small batch data it is assigned: The model parallel gradient calculation performs model parallel gradient calculation for each computing node i: The parameter synchronization and update synchronize the parameters of each computing node: Each computing node updates according to the synchronized parameters:
[0027] Furthermore, the S400 synchronization mechanism includes a distributed message queue and a distributed file system to ensure data consistency and result accuracy among various computing nodes. A synchronization monitoring and error handling mechanism is set up to monitor the synchronization status of each computing node in real time. When a node fails or experiences synchronization delays, tasks are immediately reallocated and backup nodes are started. At the same time, data during the synchronization process is verified and validated to ensure data integrity and accuracy.
[0028] Furthermore, the distributed message queue and distributed file system in the synchronization mechanism specifically include:
[0029] After each computing node completes part of the task, the distributed message queue sends the intermediate results to the distributed message queue. That is, the message queue uses message middleware. Each computing node acts as a producer of messages and sends the results to the central node. At the same time, the central node is set as the consumer of messages to read the intermediate results of each computing node from the message queue.
[0030] Distributed file system: At different stages of model training, the model parameters and intermediate results are saved to the distributed file system. Each computing node can read and update the model parameters from the distributed file system at any time. In the result merging stage, the parallel reading and writing functions of the distributed file system are used to merge the final results on each computing node into a unified file.
[0031] Compared with the prior art, the data precise matching and analysis method based on machine learning has the following
[0032] Beneficial effects:
[0033] 1. Through the distributed parallel processing technology, the present invention divides the data to be processed into multiple sub-datasets and distributes these sub-datasets to multiple computing nodes for parallel processing. Each computing node can independently process the sub-dataset assigned to it, thus making full use of the parallelism of computing resources, significantly improving the efficiency and speed of data processing. At the same time, the present invention also adopts an optimized synchronization mechanism to ensure the coordinated work among various computing nodes, avoiding processing bottlenecks caused by waiting or communication delays, and further enhancing the overall processing performance.
[0034] Second, in the result merging stage, the present invention adopts a distributed synchronization mechanism to summarize and merge the processing results of each computing node through a message middleware or other means. This synchronization mechanism can not only ensure data consistency among various computing nodes, but also effectively avoid problems of data loss or duplication, thereby guaranteeing the accuracy and reliability of the analysis results. In addition, the present invention also introduces a load balancing strategy to dynamically adjust task allocation according to the actual performance of computing nodes to ensure load balancing of each node, further improving the reliability and stability of the system.
[0035] Other advantages, objectives and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. Brief Description of the Drawings
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0037] Figure 1 It is an operation flowchart of a data precise matching analysis method based on machine learning. Detailed Embodiments
[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0039] Embodiment 1
[0040] This embodiment details the application of the data precise matching analysis method based on machine learning in the personalized recommendation system of an e-commerce platform. Through this method, efficient processing and precise analysis of user behavior data are achieved, providing personalized product recommendations for users.
[0041] In specific implementation, first, the e-commerce platform collects the behavioral data of users' browsing history, purchase records, and search keywords, as well as the attribute information and category information of products. It cleans and preprocesses this data, removes noise and abnormal data, eliminates duplicate records and obviously incorrect data, and performs data normalization operations to map the data to a unified numerical range to ensure the quality and consistency of the data. According to the performance metrics and data characteristics of the computing nodes, the preprocessed data is segmented into multiple sub-datasets. The performance metrics of each computing node are monitored in real time. The monitored metrics include the CPU utilization rate U cpu 、the memory occupancy rate U mem 、the network bandwidth B, and the disk I / O speed S io . Define the computing node performance index P, that is where w1, w2, w3, and w4 are the weight coefficients of the CPU utilization rate, memory occupancy rate, network bandwidth, and disk I / O speed respectively, and w1 + w2 + w3 + w4 = 1, B max is the maximum value of the network bandwidth, S io,max is the maximum value of the disk I / O speed. For each sub-dataset, analyze its size S data 、the data type complexity C type and the processing difficulty D proc . Define the allocation index A, that is where α is a weight parameter used to balance the impact of the computing node performance P and the sub-dataset size A on the allocation. When α ≈ 1, the allocation is more inclined to allocate the sub-dataset to the computing node with a high performance index P value. When α ≈ 0, the allocation pays more attention to the matching of the sub-dataset size and the computing node performance. According to the allocation index A, the sub-dataset is allocated to the appropriate computing node.
[0042] Then, the performance metrics of the computing nodes are monitored in real time to ensure an accurate understanding of the status of each node. Analyze the characteristics of the sub-datasets, including size, data type complexity, and processing difficulty. For sub-datasets with high processing difficulty and complex data types, allocate them to the computing nodes with a high performance index P value to give full play to the advantages of high-performance nodes. For small-scale sub-datasets, allocate them to the computing nodes with a low performance index P value and idle resources to make full use of the computing resources and avoid resource waste. On each computing node, clean the sub-datasets allocated to it, remove duplicate data and fill in missing values, and perform data normalization to map the data to a unified numerical range for subsequent feature extraction and model training. According to the characteristics of e-commerce data and the goal of personalized recommendation, extract users' behavioral characteristics, such as browsing time, purchase frequency, and categories of concerned products, as well as product characteristics, such as price, brand, and sales volume. Adopt an effective feature selection algorithm to screen out the features that have a greater impact on the recommendation results and improve the accuracy and efficiency of the model.
[0043] Subsequently, using the extracted features, a distributed training algorithm is adopted to train the machine learning model. For N computing nodes, the total dataset is D. The dataset is divided into K small batches, and the size of each small batch is |D| / K. In data parallelism, each computing node i updates the model parameters based on the small batch data assigned to it. That is, the model parameters are θ and the loss function is L(θ). Then the parameter update formula on computing node i is: where, represents the model parameters of computing node i at the t-th iteration, η is the learning rate, is the gradient of the loss function under the current parameters. After one round of iteration is completed, the parameters of each computing node are synchronized, and the method of parameter averaging is adopted, that is: so that all computing nodes use the synchronized parameters for training in the new round of iteration. The model can be divided into M parts and assigned to M computing nodes respectively. For an input sample x, the forward propagation process of the model is expressed as: y = f M (f M-1 (…f2(f1(x)))), where f i represents the calculation function of the i-th part of the model. In the backpropagation process, each computing node i calculates the gradient of the part it is responsible for and passes the gradient to the previous computing node for further gradient calculation to obtain the gradient of the entire model. That is, the loss function is L(y), then the gradient of the model is calculated by the chain rule: Each computing node i is responsible for calculating and where θ i are the model parameters that computing node i is responsible for. When updating the parameters, each computing node i updates the parameters according to the gradient it calculates: The dataset is divided into multiple small batches and assigned to different computing nodes for data parallel training. At the same time, the model is divided into multiple parts and assigned to different computing nodes for model parallel computing. In each round of iteration, each computing node first performs parameter update in data parallelism, and then performs gradient calculation and parameter update in model parallelism. The parameters of each computing node are merged and updated through the parameter synchronization mechanism. That is, the model parameters θ are initialized, the dataset is divided into small batches and assigned to different computing nodes. For each round of iteration t, data parallel update and model parallel gradient calculation are performed. The data parallel update performs data parallel parameter update for each computing node i according to the small batch data assigned to it: The model parallel gradient calculation performs model parallel gradient calculation for each computing node i: The parameter synchronization and update synchronize the parameters of each computing node: Each computing node updates according to the synchronized parameters: And the results on each computing node are merged and summarized through a distributed message queue and a distributed file system.
[0044] Among them, after each computing node completes part of the tasks, the distributed message queue sends the intermediate results to the distributed message queue. The message queue uses highly reliable message middleware. Each computing node acts as a producer of messages and sends the results to the central node. At the same time, the central node is set as the consumer of messages to read the intermediate results of each computing node from the message queue; in different stages of model training, the distributed file system saves the model parameters and intermediate results to the distributed file system. Each computing node can read and update the model parameters from the distributed file system at any time. In the result merging stage, using the parallel reading and writing functions of the distributed file system, the final results on each computing node are merged into a unified file. A synchronization monitoring and error handling mechanism is set up to monitor the synchronization status of each computing node in real time. Once a node fails or the synchronization is delayed, tasks are immediately reallocated and backup nodes are started. At the same time, the data during the synchronization process is verified and validated to ensure the integrity and accuracy of the data.
[0045] Finally, according to the results obtained from model training, a personalized product recommendation list is generated for the user. When the user accesses the e-commerce platform, the system uses the trained model to perform real-time analysis based on the user's historical behavior data, predicts the products that the user may be interested in, and recommends these products to the user to improve the user's shopping experience and the sales volume of the platform.
[0046] In summary, through the application of this embodiment in the personalized recommendation system of the e-commerce platform, the data precise matching and analysis method based on machine learning can give full play to its advantages, realize the efficient processing and precise analysis of large-scale data, provide more personalized and precise product recommendation services for users through reasonable data segmentation, node allocation and parallel processing, combined with an effective synchronization mechanism, improve the user satisfaction and the competitiveness of the e-commerce platform. At the same time, the distributed parallel processing and synchronization mechanism of this method ensure the stability and reliability of the system and can handle the processing requirements of the massive data of the e-commerce platform.
[0047] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in all respects, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims concerned.
Claims
1. A data precise matching and analysis method based on machine learning, characterized in that The specific steps of this method are as follows: S100, split the data to be processed into multiple sub-datasets, and the size of each sub-dataset is divided according to the performance metrics of the computing nodes and the data characteristics; Monitor the performance metrics of each computing node in real time. The monitored metrics include CPU utilization U cpu , memory occupancy U mem , network bandwidth B, and disk I / O speed S io . Define the computing node performance index P, that is where w1, w2, w3, and w4 are the weight coefficients of CPU utilization, memory occupancy, network bandwidth, and disk I / O speed respectively, and w1 + w2 + w3 + w4 = 1, B max is the maximum value of network bandwidth, and S io,max is the maximum value of disk I / O speed; For each sub - dataset, analyze its size S data , data - type complexity C type and processing difficulty D proc , and define the allocation index A, that is where α is a weight parameter used to balance the influence of the computing - node performance P and the sub - dataset size S data on the allocation. When α≈1, the allocation tends to assign the sub - dataset to the computing node with a high performance - index P value. When α≈0, the allocation pays more attention to the matching of the sub - dataset size and the computing - node performance; S200, allocate each sub-dataset to multiple computing nodes, and each computing node independently performs operations such as data preprocessing, feature extraction, and model training; S300, on each computing node, process the allocated sub-dataset, where: The preprocessing stage includes data cleaning and normalization operations; In the feature extraction stage, effective features are extracted according to the characteristics of the data and the objectives of matching analysis; In the model training stage, the extracted features are used to train the machine learning model; S400, through a distributed synchronization mechanism, merge and summarize the results on each computing node; S500, after merging and summarizing, obtain the final data matching analysis result.
2. The data precise matching and analysis method based on machine learning according to claim 1, wherein During the node allocation process of S200, for each sub-dataset, traverse all computing nodes, calculate its allocation index A with each computing node, and select the computing node with the largest allocation index A as the allocation target for this sub-dataset.
3. The data precise matching and analysis method based on machine learning according to claim 2, characterized in that The S200 performs real-time performance monitoring on each computing node, including the CPU utilization U cpu , the memory occupancy rate U mem , the network bandwidth B, and the disk I / O speed S io of the metrics. At the same time, analyze the size S data , the data type complexity C type and the processing difficulty D proc of the characteristics, form a sub-dataset feature vector, and match the sub-dataset feature vector with the computing node performance metric P. Specifically: For sub-datasets with high processing difficulty and complex data types, allocate them to computing nodes with high performance metric values P; For small-scale sub-datasets, allocate them to computing nodes with low performance metric values P and idle resources to make full use of computing resources.
4. The data precise matching and analysis method based on machine learning according to claim 1, characterized in that In the model training process of S300, a distributed training algorithm is adopted, that is, a method combining data parallelism and model parallelism, where: Data parallelism is to split the data into multiple small batches, allocate them to different computing nodes for training, and each computing node independently updates the model parameters; Model parallelism is to allocate different parts of the model to different computing nodes for parallel computing.
5. The data precise matching and analysis method based on machine learning according to claim 4, characterized in that For the data parallelism of N computing nodes, the total data set is D. The data set is divided into K small batches, and the size of each small batch is |D| / K, where |D| represents the total amount of data in data set D, and K is the number of small batches. In data parallelism, each computing node i updates the model parameters based on the small batch data assigned to it. That is, if the model parameter is θ and the loss function is L(θ), the parameter update formula on computing node i is: where represents the model parameter of computing node i at the t-th iteration, η is the learning rate, is the gradient of the loss function under the current parameters. After one round of iteration is completed, the parameters of each computing node are synchronized, and the method of parameter averaging is adopted, that is: so that all computing nodes use the synchronized parameters for training in the new round of iteration.
6. The data precise matching and analysis method based on machine learning according to claim 4, wherein The model parallelism divides the model into M parts and distributes them to M computing nodes respectively. For an input sample x, the forward propagation process of the model is expressed as: y = f M (f M-1 (…f2(f1(x)))). Where f M represents the calculation function of the M-th part of the model, and y represents the output result of the forward propagation of the model. During the backpropagation process, each computing node i calculates the gradient of its responsible part and passes the gradient to the previous computing node for further gradient calculation to obtain the gradient of the entire model. That is, if the loss function is L(y), the gradient of the model is calculated by the chain rule: Each computing node i is responsible for calculating and where θ i is the model parameter responsible for computing node i, and θ is the model parameter. When updating the parameters, each computing node i updates the parameters according to the gradient it calculates: where represents the model parameter of computing node i at the t-th iteration, η is the learning rate, is the gradient of the loss function with respect to the model parameter θ i .
7. The data precise matching and analysis method based on machine learning according to claim 6, characterized in that During the model training process, S300 combines data parallelism and model parallelism. It divides the dataset into multiple small batches and distributes them to different computing nodes for data parallel training. At the same time, the model is divided into multiple parts and distributed to different computing nodes for model parallel computing. In each iteration, each computing node first performs parameter updates for data parallelism, and then performs gradient calculations and parameter updates for model parallelism. Through the parameter synchronization mechanism, the parameters of each computing node are merged and updated, that is, initialize the model parameters θ, divide the dataset into small batches and distribute them to different computing nodes. For each iteration t, perform data parallel updates and model parallel gradient calculations. The data parallel update performs parameter updates for data parallelism for each computing node i according to the small batch data assigned to it: The model parallel gradient calculation performs gradient calculations for model parallelism for each computing node i: The parameter synchronization and update synchronizes the parameters of each computing node: Each computing node updates according to the synchronized parameters:
8. The data precise matching and analysis method based on machine learning according to claim 1, wherein, The synchronization mechanism of S400 includes a distributed message queue and a distributed file system to ensure data consistency and result accuracy among computing nodes, and a synchronization monitoring and error handling mechanism is set up to monitor the synchronization status of each computing node in real time. When a node fails and there is a synchronization delay, tasks are immediately reallocated and backup nodes are started. At the same time, the data during the synchronization process is verified and validated to ensure data integrity and accuracy.
9. The data precise matching and analysis method based on machine learning according to claim 8, wherein In the synchronization mechanism, the distributed message queue and the distributed file system specifically include: The distributed message queue sends the intermediate results to the distributed message queue after each computing node completes part of the tasks. That is, the message queue uses message middleware. Each computing node acts as a producer of messages and sends the results to the central node. At the same time, the central node is set as the consumer of messages to read the intermediate results of each computing node from the message queue; Distributed file system: In different stages of model training, save the model parameters and intermediate results to the distributed file system. Each computing node can read and update the model parameters from the distributed file system at any time. In the result merging stage, use the parallel reading and writing functions of the distributed file system to merge the final results on each computing node into a unified file.
Citation Information
Patent Citations
Data distribution method based on node data processing capability and node operation load
CN109600359A
Big data file analysis processing method and system in cloud computing environment
CN118535577A