A multi-modal AI model training method based on a distributed network architecture
The distributed network architecture addresses data synchronization, consistency, and resource management issues in multi-modal AI training, enhancing efficiency and performance through innovative protocols and techniques.
Patent Information
- Application Number
- CN202411190744.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-08-28
AI Technical Summary
The prior art has shortcomings in multimodal data synchronization, data consistency, computing resource allocation, fault-tolerant processing and parallel computing, which affects the efficiency and performance of multimodal AI model training.
The distributed file system HDFS and gRPC protocols are used for data synchronization, the improved Paxos consistency algorithm ensures data consistency, the Kubernetes resource management framework is deployed for dynamic resource allocation, combined with the Prometheus monitoring system to achieve fault-tolerant recovery, and the TensorFlow Distributed framework is used for parallel training, and multimodal data fusion is carried out through the attention mechanism.
Real-time synchronization and consistency of multimodal data is realized, the utilization of computing resources is optimized, the stability and speed of model training is improved, and the comprehensive performance of multimodal AI models is improved.
Smart Images

Figure CN118981503B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and distributed computing, and particularly to a multi-modal AI model training method based on a distributed network architecture. Background Art
[0002] With the development of artificial intelligence technology, the processing and analysis of multi-modal data have gradually become a research hotspot. A multi-modal AI model needs to simultaneously process heterogeneous data from different data sources, such as images, texts, and audios, which poses higher requirements for model training. However, there are many deficiencies in the existing technologies when dealing with multi-modal data and model training, mainly manifested in the following aspects:
[0003] 1. Data synchronization problem: Existing distributed systems mostly focus on the processing of single-modal data and lack effective solutions for the synchronization of multi-modal data. Since multi-modal data comes from different sources, it is necessary to maintain data synchronization in a distributed environment to ensure the consistency and effectiveness of model training. However, in a multi-node environment, it is a great challenge to achieve real-time synchronization of multi-modal data, and problems such as data delay and out-of-sync often occur, affecting the model training effect.
[0004] 2. Data consistency problem: Multi-modal data is prone to inconsistency during transmission and processing, especially in a distributed environment, and it is difficult to guarantee data consistency between different nodes. Existing consistency algorithms have performance bottlenecks and complexity problems when dealing with multi-modal data, resulting in data inconsistency, which in turn affects the training effect and performance of the model.
[0005] 3. Computational resource allocation problem: Existing resource management strategies are mostly static allocations and cannot flexibly cope with changes in computational loads. In a distributed environment, how to efficiently allocate computational resources to make full use of the computational capabilities of each node and accelerate the model training process is an urgent problem to be solved. Existing technologies have deficiencies in resource dynamic scheduling and load balancing and are difficult to achieve the optimal utilization of computational resources.
[0006] 4. Fault tolerance handling problem: In a distributed environment, node failures and network delays are inevitable, but existing fault tolerance mechanisms often react slowly and have a long recovery time when dealing with failures in a large-scale distributed environment. The lack of effective fault detection and recovery mechanisms makes it difficult to guarantee the system stability and the continuity of the training process.
[0007] 5. Parallel computing and model training problem: Although existing distributed training frameworks exist, they still have limitations in dealing with multi-modal data. Existing parallel computing technologies and model training methods are difficult to achieve efficient parallel processing when facing large-scale, multi-source heterogeneous data, and the training speed and model performance are limited.
[0008] 6. Issues of Multimodal Data Processing and Fusion: The processing and fusion of multimodal data have always been difficult points in the training of AI models. Existing technologies have limitations in data fusion methods and are difficult to achieve efficient and accurate comprehensive processing of multimodal data, which affects the comprehensive performance and application effects of the models.
[0009] Therefore, how to provide a multimodal AI model training method based on a distributed network architecture is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention
[0010] An object of the present invention is to propose a multimodal AI model training method based on a distributed network architecture. Through innovative distributed data synchronization protocols, consistency guarantee mechanisms, dynamic resource allocation strategies, fault tolerance and recovery mechanisms, as well as advanced parallel computing and multimodal data processing and fusion technologies, the present invention solves many defects in the prior art, significantly improves the efficiency, performance and reliability of multimodal AI model training, and provides strong technical support for the training of AI models for large-scale, multi-source heterogeneous data.
[0011] A multimodal AI model training method based on a distributed network architecture according to an embodiment of the present invention includes the following steps:
[0012] S1. In a multi-node environment, store multimodal data through the distributed file system HDFS, and use the gRPC protocol to transfer multimodal data between nodes to achieve data synchronization;
[0013] S2. Perform preliminary preprocessing on the multimodal data, specifically including denoising, normalization and enhancement of image data, word segmentation, stop word removal and word vectorization processing of text data, and denoising, frame splitting and feature extraction of audio data;
[0014] S3. Use an improved Paxos consistency algorithm to ensure data consistency among multiple nodes, and ensure data consistency in the case of network partitioning or node failures through a distributed consensus mechanism, and use the distributed database Apache Cassandra for data storage;
[0015] S4. Deploy the Kubernetes resource management framework to monitor the usage of computing resources of each node in real time, dynamically adjust the allocation of computing tasks according to the node load conditions, and balance the task loads of each node through the Nginx load balancing technology;
[0016] S5. In a distributed environment, use the Prometheus distributed monitoring system to monitor the node status in real time, automatically detect node failures, and trigger task reallocation and data backup when a failure is detected to ensure the rapid recovery of the system;
[0017] S6. Adopt the TensorFlow Distributed distributed training framework, combine data parallelism and model parallelism technologies, distribute large-scale data and model parameters to multiple computing nodes for parallel training, and accelerate the model training process;
[0018] S7. Further process and extract features from the multi-modal data;
[0019] S8. Use the attention mechanism to fuse data features of different modalities, generate comprehensive feature representations, and improve the training effect and comprehensive performance of the multi-modal AI model.
[0020] Furthermore, the specific steps of S1 include:
[0021] S11. In a multi-node environment, store multi-modal data in the distributed file system HDFS. The multi-modal data includes image data, text data, and audio data. Each node stores and accesses data through the distributed file system HDFS;
[0022] S12. Store and manage image data through the distributed file system HDFS, including block storage and redundant storage of image data to ensure high availability and reliability of data in case of node failures;
[0023] S13. For text data, perform distributed storage and index management through the distributed file system HDFS to ensure data reading efficiency in a large-scale text data environment, and use the inverted index technology to improve the retrieval speed of text data;
[0024] S14. For audio data, adopt the distributed file system HDFS for distributed storage, and ensure efficient storage and transmission of audio data in a distributed node environment through the sharding and multi-copy mechanism of audio data;
[0025] S15. In the process of data transmission between multiple nodes, use the gRPC protocol to achieve efficient data transmission, including:
[0026] Define the gRPC service and message format to ensure unity and standardization in the process of data transmission between multiple nodes;
[0027] Use the streaming transmission function of gRPC to transmit multi-modal data in real time, specifically including the streaming transmission of image data, text data, and audio data;
[0028] During the transmission process, use data compression technology to reduce the amount of data transmitted and improve data transmission efficiency, specifically including JPEG compression of image data, gzip compression of text data, and FLAC compression of audio data;
[0029] Establish an efficient data transmission channel between nodes, leveraging the multiplexing and load balancing capabilities of gRPC to ensure the stability and efficiency of data transmission;
[0030] To ensure the security of data transmission, use the TLS protocol to encrypt the gRPC data transmission channel and prevent data from being eavesdropped on and tampered with during transmission;
[0031] During data transmission, monitor the transmission status in real-time, record transmission logs, and use log analysis tools to detect and handle anomalies during the transmission process;
[0032] After data transmission is completed, perform data verification and confirmation at each node to ensure the integrity and consistency of the transmitted data. Use checksum codes to verify the transmitted data and determine whether the data transmission is successful based on the verification results.
[0033] Furthermore, the improved Paxos consensus algorithm of S3 specifically includes:
[0034] S31. Run Paxos instances in a multi-node environment, including three roles: proposers, acceptors, and learners. Optimize the message passing mechanism through batch message processing and message compression techniques to reduce communication overhead. The formula for the message compression technique is:
[0035]
[0036] where C msg represents the size of the compressed message, S i represents the size of the i-th message, and N represents the number of messages; this formula compresses the messages to reduce the size of each message and thereby reduce the burden on network communication. The size of the compressed message is the average calculated by logarithmically scaling all message sizes.
[0037] By optimizing the message passing mechanism through batch message processing and message compression techniques, the communication overhead is reduced, the efficiency of message passing is improved, and the consumption of network bandwidth is reduced.
[0038] S32. The proposer receives requests from clients and generates proposals. A proposal contains a unique proposal number n and a proposal value v, and multiple proposals are merged into one message for sending, which improves the message passing efficiency, reduces the number of single-proposal processing times, and reduces the network load. The proposer sends Prepare messages to multiple acceptors. The formula for batch processing is:
[0039]
[0040] where B msg represents the batch-processed message, P jDenote the j-th proposal, ∈ represents the coefficient, and m represents the number of proposals; Proposal batch processing reduces the number of messages by combining multiple proposals into one message. ∈ in the formula is a small positive number used to adjust the logarithmic ratio of the proposal size.
[0041] S33. After the receiver receives the Prepare message, if the number n is greater than the maximum number it has responded to, it accepts the Prepare message and promises not to accept proposals with numbers less than n. At the same time, it sends a Promise message to the proposer. The Promise message contains the number of the proposal with the maximum number accepted by the receiver and the corresponding proposal value.
[0042] S34. After the proposer receives the Promise messages from the majority of receivers, it immediately generates a proposal with number n and proposal value v and sends an Accept request to the receivers. The message processing speed is improved by multi-threaded parallel processing. The formula for multi-threaded parallel processing is:
[0043]
[0044] where, T process represents the total time of parallel processing, T i represents the processing time of the i-th thread, k represents the number of threads, and α represents the constant coefficient.
[0045] Through multi-threaded parallel processing, the message processing speed is improved. α in the formula is used to adjust the square root part of the thread processing time, further optimizing the total processing time.
[0046] S35. After the receiver receives the Accept request, if the number n is greater than or equal to the maximum number it has responded to, it accepts the proposal and sends an Accepted message to the proposer. The Accepted message contains the number of the proposal and the proposal value. A dynamic priority mechanism is adopted to ensure that important proposals reach a consensus first. The formula for the dynamic priority mechanism is:
[0047]
[0048] where, P priority represents the proposal priority, L represents the delay of the proposal, T represents the importance of the proposal, R represents the resource consumption of the proposal, and β is a small positive number to avoid the denominator being zero.
[0049] The dynamic priority mechanism calculates the priority of the proposal by comprehensively considering the delay, importance, and resource consumption of the proposal to ensure that important proposals are processed first. β in the formula avoids the situation where the denominator is zero.
[0050] S36. After the proposer receives the Accepted messages from the majority of acceptors, the proposal is considered to have reached a consensus. The proposer broadcasts the proposal value to all learners and reduces the broadcast latency through the hierarchical broadcast and intelligent broadcast strategies. The formula for hierarchical broadcast is:
[0051]
[0052] where H broadcast represents the total time of hierarchical broadcast, B l represents the broadcast time of the l-th layer, and h represents the number of layers;
[0053] Hierarchical broadcast reduces the total broadcast time through a hierarchical broadcast strategy. The broadcast time of each layer in the formula is calculated logarithmically, further optimizing the broadcast efficiency.
[0054] S37. After the learner receives the notification from the proposer, it records the proposal value and applies the value to ensure data consistency in the distributed system. An efficient consistency check technology is adopted to quickly check data consistency before each write operation. The formula for efficient consistency check is:
[0055]
[0056] where E check represents the total time of consistency check, H(D i ) represents the hash value calculation time of the i-th data block, n represents the number of data blocks, and γ represents a constant coefficient;
[0057] Consistency check optimizes the check time through hash value calculation and adjustment coefficient to ensure data consistency before write operations. γ in the formula is used to optimize the square root part of the hash value calculation time.
[0058] S38. In the case of network partitioning or node failure, the acceptor and the proposer will renegotiate the proposal according to the steps of the improved Paxos consistency algorithm, and introduce a fault node isolation and dynamic priority mechanism to ensure that the system can quickly reach a consensus and maintain consistency after fault recovery;
[0059] S39. The distributed database Apache Cassandra is used for data storage. Combining the incremental update strategy and the improved Paxos consistency algorithm to achieve strong consistency storage. Before each write operation, a consensus is reached among multiple nodes through an optimized consensus mechanism to ensure the consistency of write operations on all replicas. The formula for the incremental update strategy is:
[0060]
[0061] where Δ updateRepresents the amount of data for incremental update, δ i Represents the size of the data updated for the i-th time, and m represents the number of updates;
[0062] The incremental update strategy reduces the overhead of data transmission and storage by only transmitting and storing the changed part of the data. The square part of the changed data volume is considered in the formula, further optimizing the update strategy.
[0063] Furthermore, the specific steps of S4 are as follows:
[0064] S41. In a multi-node environment, install and configure the Kubernetes resource management framework, and deploy the Kubernetes node agent on each computing node to ensure that each node can be managed and scheduled by the Kubernetes master node;
[0065] S42. Use the resource monitoring module of Kubernetes to collect and monitor the CPU, memory, and network bandwidth resource usage of each node in real time, and summarize the monitoring data to the Kubernetes master node for use by the resource scheduling algorithm;
[0066] S43. Design and implement a dynamic priority scheduling algorithm. The algorithm predicts the load change trend of each node based on real-time monitoring data and historical resource usage, and preferentially allocates resources to nodes with lighter loads to improve resource utilization. The priority P i (t) is calculated by the following formula:
[0067]
[0068] where, P i (t) represents the priority of node i at time t, which is used to determine the priority order of resource allocation. CPU i (t) represents the CPU usage rate of node i at time t. MEM i (t) represents the memory usage rate of node i at time t. α, β, γ represent weight coefficients, which are used to balance the influence of CPU usage rate, memory usage rate, and historical load change trend on the priority. H i (t) represents the historical load change trend of node i at time t, which represents the resource usage of the node in the past period of time and is calculated by the exponentially weighted moving average of historical data:
[0069]
[0070] where, λ represents the smoothing coefficient, and the value range is 0 < λ ≤ 1;
[0071] S44. Configure the Nginx load balancer. By setting reverse proxy and load balancing rules, evenly distribute user requests and data traffic to different computing nodes to ensure that the task load of each node is evenly distributed and avoid overloading a single node.
[0072] S45. Design and implement a load balancing optimization algorithm based on machine learning. Utilize the historical load data and current status of nodes to dynamically adjust the load balancing rules of Nginx; use a linear regression model to predict the load situation L of node i after time t: i (t + Δt):
[0073] L i (t + Δt) = θ0 + θ1·CPU i (t) + θ2·MEM i (t) + θ3·Net i (t);
[0074] Where, L i (t + Δt) represents the predicted load situation of node i at time t + Δt, θ0, θ1, θ2, θ3 represent the parameters of the linear regression model, CPU i (t) represents the CPU usage rate of node i at time t, MEM i (t) represents the memory usage rate of node i at time t, Net i (t) represents the network bandwidth usage rate of node i at time t; the load balancing rules are dynamically adjusted according to the predicted load situation, and select the node j with the lightest load to process new tasks:
[0075]
[0076] Where, j represents the node with the lightest load selected;
[0077] S46. When it is detected that a certain node has a resource bottleneck or fails, utilize the automatic task rescheduling function of Kubernetes to re - allocate the affected tasks to other nodes with sufficient resources.
[0078] S47. Introduce a scheduling decision record mechanism based on blockchain technology. Record the decisions of each task scheduling and resource allocation on the blockchain. The hash value H of the blockchain record is calculated as follows:
[0079] H = SHA256(T i ||P i (t)||L i (t + Δt));
[0080] Among them, H represents the hash value recorded by the blockchain, SHA256 represents the hash function, Ti represents the task ID, Pi(t) represents the priority of node i at time t, and L i (t + Δt) represents the predicted load condition of node i at time t + Δt. The || represents the string concatenation operator, which is used to concatenate multiple values into a string for hash calculation.
[0081] Furthermore, the S7 specifically includes: using a convolutional neural network to extract features from image data; using a long short-term memory network to extract features from text data; using Mel spectrogram to transform audio data and using a convolutional neural network to extract features.
[0082] The beneficial effects of the present invention are:
[0083] 1. Efficient data synchronization: The present invention realizes the real-time synchronization of multimodal data between different nodes through the distributed file system HDFS and the efficient data transmission protocol gRPC, significantly reducing data latency and asynchronous problems, and ensuring data consistency and effectiveness during the model training process.
[0084] 2. Strong consistency data guarantee: By using the improved Paxos consistency algorithm, a solid data consistency guarantee mechanism is established among multiple nodes. Through the distributed consensus mechanism, data consistency is ensured in the case of network partitioning or node failures, and the Apache Cassandra distributed database is used for data storage, completely solving the consistency problem in the process of multimodal data processing and transmission, and avoiding model training errors caused by data inconsistency.
[0085] 3. Dynamic resource allocation: By deploying the Kubernetes resource management framework, the computing resource usage of each node is monitored in real time, and the allocation of computing tasks is dynamically adjusted according to the load conditions. Combining with the Nginx load balancing technology, it ensures that the computing resources of each node are optimally utilized, significantly improving the efficiency and performance of model training.
[0086] 4. Reliable fault tolerance and recovery mechanism: In the distributed environment, the present invention uses the Prometheus distributed monitoring system to monitor the node status in real time, automatically detects node failures and triggers task reallocation and data backup, ensuring that the system can quickly recover in case of failures, guaranteeing the stability and continuity of the training process, and greatly improving the reliability of the system.
[0087] 5. Efficient parallel training: By adopting the TensorFlow Distributed distributed training framework and combining data parallelism and model parallelism techniques, large-scale data and model parameters are distributed to multiple computing nodes for parallel training, effectively accelerating the model training process and improving the training speed and model performance.
[0088] 6. Precise multi-modal data processing and fusion: Through the preliminary preprocessing and feature extraction of image, text, and audio data, and combining the attention mechanism to fuse the data features of different modalities, a comprehensive feature representation is generated. The fusion technology of the present invention improves the efficiency and accuracy of multi-modal data comprehensive processing, and significantly improves the training effect and comprehensive performance of the AI model. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0090] Figure 1 is a flowchart of a multi-modal AI model training method based on a distributed network architecture proposed by the present invention;
[0091] Figure 2 is a schematic diagram of a multi-node data consistency guarantee mechanism of a multi-modal AI model training method based on a distributed network architecture proposed by the present invention;
[0092] Figure 3 is a schematic diagram of dynamic allocation of computing resources and load balancing of a multi-modal AI model training method based on a distributed network architecture proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0093] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.
[0094] Refer to Figures 1-3 , a multi-modal AI model training method based on a distributed network architecture, includes the following steps:
[0095] S1. In a multi-node environment, store multi-modal data through the distributed file system HDFS, and use the gRPC protocol to transfer multi-modal data between nodes to achieve data synchronization;
[0096] S2. Perform preliminary preprocessing on the multi-modal data, specifically including denoising, normalization, and enhancement of image data, word segmentation, stop word removal, and word vectorization processing of text data, and denoising, frame splitting, and feature extraction of audio data;
[0097] S3. Use the improved Paxos consensus algorithm to ensure data consistency among multiple nodes. Through the distributed consensus mechanism, ensure data consistency in the case of network partitions or node failures, and use the distributed database Apache Cassandra for data storage;
[0098] S4. Deploy the Kubernetes resource management framework to monitor the computing resource usage of each node in real time. Dynamically adjust the computing task allocation according to the node load conditions, and balance the task loads of each node through the Nginx load balancing technology;
[0099] S5. In a distributed environment, use the Prometheus distributed monitoring system to monitor the node status in real time, automatically detect node failures, and when a failure is detected, trigger task reallocation and data backup to ensure the rapid recovery of the system;
[0100] S6. Adopt the TensorFlow Distributed distributed training framework, combine data parallelism and model parallelism technologies, distribute large-scale data and model parameters to multiple computing nodes for parallel training, and accelerate the model training process;
[0101] S7. Further process and extract features from the multi-modal data;
[0102] S8. Use the attention mechanism to fuse the data features of different modalities to generate a comprehensive feature representation, and improve the training effect and comprehensive performance of the multi-modal AI model.
[0103] In this embodiment, the S1 specifically includes:
[0104] S11. In a multi-node environment, store the multi-modal data in the distributed file system HDFS. The multi-modal data includes image data, text data, and audio data, and each node stores and accesses data through the distributed file system HDFS;
[0105] S12. Store and manage the image data through the distributed file system HDFS, including block storage and redundant storage of the image data to ensure the high availability and reliability of the data in the case of node failures;
[0106] S13. For text data, perform distributed storage and index management through the distributed file system HDFS to ensure the data reading efficiency in a large-scale text data environment, and use the inverted index technology to improve the retrieval speed of text data;
[0107] S14. For audio data, use the distributed file system HDFS for distributed storage. Through the sharding and multi-copy mechanism of audio data, ensure the efficient storage and transmission of audio data in a distributed node environment;
[0108] S15. During the data transmission between multiple nodes, use the gRPC protocol to achieve efficient data transmission, including:
[0109] Define the gRPC service and message format to ensure the unity and standardization during the data transmission between multiple nodes;
[0110] Utilize the streaming transmission function of gRPC to transmit multi-modal data in real time, specifically including the streaming transmission of image data, text data, and audio data;
[0111] During the transmission process, use data compression technology to reduce the data transmission volume and improve the data transmission efficiency, specifically including JPEG compression for image data, gzip compression for text data, and FLAC compression for audio data;
[0112] Establish an efficient data transmission channel between nodes, and utilize the multiplexing and load balancing functions of gRPC to ensure the stability and efficiency of data transmission;
[0113] To ensure the security of data transmission, use the TLS protocol to encrypt the gRPC data transmission channel to prevent data from being eavesdropped and tampered with during the transmission process;
[0114] During the data transmission process, monitor the transmission status in real time, record the transmission log, and use the log analysis tool to detect and handle the anomalies during the transmission process;
[0115] After the data transmission is completed, perform data verification and confirmation on each node to ensure the integrity and consistency of the transmitted data. Use the checksum to verify the transmitted data, and judge whether the data transmission is successful through the verification result.
[0116] In this embodiment, the improved Paxos consistency algorithm of S3 specifically includes:
[0117] S31. Run Paxos instances in a multi-node environment, including three roles: proposer, acceptor, and learner. Optimize the message passing mechanism through batch processing of messages and message compression technology to reduce communication overhead. The formula of the message compression technology is:
[0118]
[0119] Among them, C msg represents the size of the compressed message, S irepresents the size of the i-th message, and N represents the number of messages; this formula reduces the size of each message by compressing the messages, so as to reduce the burden of network communication. The size of the compressed message is the average value calculated by logarithmically scaling all message sizes.
[0120] S32. The proposer receives requests from the client and generates proposals. The proposal contains a unique proposal number n and a proposal value v, and combines multiple proposals into one message for sending to improve the message delivery efficiency. The proposer sends Prepare messages to multiple acceptors. The formula for batch processing is as follows:
[0121]
[0122] where B msg represents the message after batch processing, P j represents the j-th proposal, ∈ represents a coefficient, and m represents the number of proposals; Proposal batch processing reduces the number of messages by combining multiple proposals into one message. ∈ in the formula is a small positive number used to adjust the logarithmic ratio of the proposal size.
[0123] S33. After receiving the Prepare message, if the number n is greater than the maximum number it has responded to, the acceptor accepts the Prepare message and promises not to accept proposals with numbers less than n, and at the same time replies a Promise message to the proposer. The Promise message contains the number of the proposal with the maximum number accepted by the acceptor and the corresponding proposal value;
[0124] S34. After receiving the Promise messages from the majority of acceptors, the proposer immediately generates a proposal with the number n and the proposal value v, and sends an Accept request to the acceptors. The message processing speed is improved through multi-threaded parallel processing. The formula for multi-threaded parallel processing is as follows:
[0125]
[0126] where T process represents the total time of parallel processing, T i represents the processing time of the i-th thread, k represents the number of threads, and α represents a constant coefficient;
[0127] Through multi-threaded parallel processing, the message processing speed is improved. α in the formula is used to adjust the square root part of the thread processing time, further optimizing the total processing time.
[0128] S35. After the acceptor receives the Accept request, if the number n is greater than or equal to the maximum number it has responded to, it accepts the proposal and replies to the proposer with an Accepted message. The Accepted message contains the proposal number and the proposal value. A dynamic priority mechanism is adopted to ensure that important proposals reach a consensus first. The formula of the dynamic priority mechanism is:
[0129]
[0130] Among them, P priority represents the proposal priority, L represents the proposal delay, T represents the importance of the proposal, R represents the resource consumption of the proposal, and β is a small positive number to avoid the denominator being zero;
[0131] The dynamic priority mechanism calculates the priority of the proposal by comprehensively considering the delay, importance, and resource consumption of the proposal to ensure that important proposals are processed first. The β in the formula avoids the situation where the denominator is zero.
[0132] S36. After the proposer receives the Accepted messages from the majority of acceptors, the proposal is considered to have reached a consensus. The proposer broadcasts the proposal value to all learners and reduces the broadcast delay through hierarchical broadcast and intelligent broadcast strategies. The formula of hierarchical broadcast is:
[0133]
[0134] Among them, H broadcast represents the total time of hierarchical broadcast, B l represents the broadcast time of the l-th layer, and h represents the number of layers;
[0135] Hierarchical broadcast reduces the total broadcast time through a hierarchical broadcast strategy. The broadcast time of each layer in the formula is calculated logarithmically, further optimizing the broadcast efficiency.
[0136] S37. After the learner receives the notification from the proposer, it records the proposal value and applies the value to ensure data consistency in the distributed system. An efficient consistency check technology is adopted to quickly check data consistency before each write operation. The formula of efficient consistency check is:
[0137]
[0138] Among them, E check represents the total time of consistency check, H(D i ) represents the hash value calculation time of the i-th data block, n represents the number of data blocks, and γ represents a constant coefficient;
[0139] Consistency checking optimizes the checking time through hash value calculation and adjustment coefficient, ensuring data consistency before write operations. γ in the formula is used to optimize the square root part of the hash value calculation time.
[0140] S38. In the case of network partitioning or node failures, the acceptor and the proposer will renegotiate the proposal according to the steps of the improved Paxos consistency algorithm, and introduce a faulty node isolation and dynamic priority mechanism to ensure that the system can quickly reach a consensus and maintain consistency after fault recovery;
[0141] S39. Use the distributed database Apache Cassandra for data storage, and combine the incremental update strategy and the improved Paxos consistency algorithm to achieve strong consistency storage. Before each write operation, a consensus is reached among multiple nodes through an optimized consensus mechanism to ensure the consistency of the write operation on all replicas. The formula for the incremental update strategy is:
[0142]
[0143] where, Δ update represents the amount of incrementally updated data, δ i represents the size of the data updated at the i-th time, and m represents the number of updates;
[0144] The incremental update strategy reduces the overhead of data transmission and storage by only transmitting and storing the changed part of the data. The square part of the changed data volume is considered in the formula, further optimizing the update strategy.
[0145] In this embodiment, the said S4 specifically includes:
[0146] S41. In a multi-node environment, install and configure the Kubernetes resource management framework, and deploy a Kubernetes node agent on each computing node to ensure that each node can be managed and scheduled by the Kubernetes master node;
[0147] S42. Use the resource monitoring module of Kubernetes to collect and monitor the CPU, memory, and network bandwidth resource usage of each node in real time, and summarize the monitoring data to the Kubernetes master node for use by the resource scheduling algorithm;
[0148] S43. Design and implement a dynamic priority scheduling algorithm. The algorithm predicts the load change trend of each node based on real-time monitoring data and historical resource usage, and preferentially allocates resources to nodes with lighter loads to improve resource utilization. The priority P i (t) is calculated by the following formula:
[0149]
[0150] Among them, P i (t) represents the priority of node i at time t, which is used to determine the priority order of resource allocation. CPU i (t) represents the CPU usage rate of node i at time t. MEM i (t) represents the memory usage rate of node i at time t. α, β, and γ represent weight coefficients, which are used to balance the influence of CPU usage rate, memory usage rate, and historical load change trend on the priority. H i (t) represents the historical load change trend of node i at time t, which represents the resource usage situation of the node in the past period of time and is calculated by the exponentially weighted moving average of historical data:
[0151]
[0152] Among them, λ represents the smoothing coefficient, and the value range is 0 < λ ≤ 1;
[0153] S44. Configure the Nginx load balancer. By setting the reverse proxy and load balancing rules, evenly distribute the user requests and data traffic to different computing nodes to ensure that the task loads of each node are evenly distributed and avoid overloading a single node;
[0154] S45. Design and implement a load balancing optimization algorithm based on machine learning. Utilize the historical load data and current state of the nodes to dynamically adjust the load balancing rules of Nginx; Use a linear regression model to predict the load situation L i (t + Δt) of node i after time t:
[0155] L i (t + Δt) = θ0 + θ1·CPU i (t) + θ2·MEM i (t) + θ3·Net i (t);
[0156] Among them, L i (t + Δt) represents the predicted load situation of node i at time t + Δt. θ0, θ1, θ2, and θ3 represent the parameters of the linear regression model. CPU i (t) represents the CPU usage rate of node i at time t. MEM i (t) represents the memory usage rate of node i at time t. Net i (t) represents the network bandwidth usage rate of node i at time t. The load balancing rules are dynamically adjusted according to the predicted load situation, and the node j with the lightest load is selected to process new tasks:
[0157]
[0158] Among them, j represents the node with the lightest load selected;
[0159] S46. When it is detected that a certain node has a resource bottleneck or a failure, use the automatic task rescheduling function of Kubernetes to reallocate the affected tasks to other nodes with sufficient resources;
[0160] S47. Introduce a scheduling decision record mechanism based on blockchain technology, record the decisions of each task scheduling and resource allocation on the blockchain, and calculate the hash value H of the blockchain record as follows:
[0161] H = SHA256(T i ||P i (t)||L i (t + Δt));
[0162] Among them, H represents the hash value of the blockchain record, SHA256 represents the hash function, T i represents the task ID, P i (t) represents the priority of node i at time t, and L i (t + Δt) represents the predicted load condition of node i at time t + Δt, and || represents the string concatenation operator, which is used to concatenate multiple values into a string for hash calculation.
[0163] In this embodiment, the S7 specifically includes: using a convolutional neural network to extract features from image data; using a long short-term memory network to extract features from text data; using Mel spectrogram to convert audio data, and using a convolutional neural network to extract features.
[0164] Example 1:
[0165] To verify the feasibility of the present invention in implementation, the present invention is applied to an Internet company. The company needs to build a multi-modal AI model to improve the accuracy of its recommendation system and the user experience. This model needs to process a large amount of heterogeneous data from different data sources, including images browsed by users, text comments submitted, and audio data input by voice. Due to the huge amount of data and the scattered sources, the traditional centralized data processing method cannot meet the real-time processing and training requirements, so it is decided to adopt a multi-modal AI model training method based on a distributed network architecture.
[0166] To achieve efficient training of a multi-modal AI model, the system first deploys the distributed file system HDFS in a multi-node environment to store multi-modal data from different data sources. The data includes millions of product images browsed by users, hundreds of thousands of text reviews, and tens of thousands of hours of user voice input. The gRPC protocol is used to transfer this multi-modal data between nodes to achieve real-time data synchronization.
[0167] After the data synchronization is completed, the system performs preliminary preprocessing on this data. The image data is denoised through Gaussian filtering and normalized using the Min-Max normalization method. At the same time, data augmentation techniques such as random cropping and rotation are used to increase the diversity of the data. The text data is tokenized using the NLTK library, stop words are removed, and then the Word2Vec model is used to convert the text into word vectors. The audio data is first subjected to wavelet transform to remove noise, then framed, with each frame having a length of 25ms, and finally Mel-frequency cepstral coefficients (MFCCs) are extracted as features.
[0168] To ensure data consistency between multiple nodes, the system utilizes an improved Paxos consistency algorithm to ensure data consistency in the event of network partitions or node failures through a distributed consensus mechanism, and uses the Apache Cassandra distributed database for data storage to ensure data consistency and persistence.
[0169] In terms of resource management, the system deploys the Kubernetes resource management framework to monitor the usage of computing resources on each node in real-time and dynamically adjust the distribution of computing tasks according to the node load. Through the Nginx load balancing technology, it ensures that the computing resources of each node are optimally utilized, avoiding overloading or idling of some nodes.
[0170] In a distributed environment, the system uses the Prometheus distributed monitoring system to monitor the node status in real-time, automatically detect node failures and trigger task reallocation and data backup. When a node fails, the system can quickly reallocate the tasks on that node to other normally running nodes and restore the data of the failed node from the backup data to ensure the stability of the system and the continuity of the training process.
[0171] Using the TensorFlow Distributed distributed training framework, the system combines data parallelism and model parallelism techniques to distribute large-scale data and model parameters to multiple computing nodes for parallel training, accelerating the model training process. During the training process, the system further processes and extracts features from the multi-modal data. The image data extracts features through a convolutional neural network (CNN), the text data extracts features through a long short-term memory network (LSTM), and the audio data extracts features through Mel-spectrum conversion and a convolutional neural network.
[0172] Finally, the system uses the attention mechanism to fuse the data features of different modalities, generates a comprehensive feature representation, and inputs it into the AI model for training and optimization. In this way, the system achieves efficient multi-modal AI model training, significantly improving the training effect and comprehensive performance of the model.
[0173] Table 1 Comparison table of comprehensive performance data between the traditional method and the method of the present invention
[0174]
[0175]
[0176] Table 1 fully demonstrates the significant improvement of the performance indicators of the present invention compared with the traditional method in practical applications. These data fully prove the technical advantages and beneficial effects of the present invention in aspects such as data synchronization, data consistency guarantee, computing resource utilization, fault detection and recovery, model training speed, and recommendation system performance.
[0177] In summary, the present invention significantly improves the efficiency, performance, and reliability of multi-modal AI model training through an efficient data synchronization and consistency guarantee mechanism, a dynamic resource allocation strategy, a reliable fault tolerance and recovery mechanism, and advanced parallel computing and multi-modal data processing and fusion technologies, solves many defects in the prior art, and provides strong technical support for the training of AI models for large-scale, multi-source heterogeneous data.
[0178] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A method for training a multi-modal AI model based on a distributed network architecture, characterized in that, It includes the following steps: S1. In a multi-node environment, store multi-modal data through the distributed file system HDFS, and use the gRPC protocol to transfer multi-modal data between nodes to achieve data synchronization; S2. Perform preliminary preprocessing on the multi-modal data, specifically including denoising, normalization, and enhancement of image data, word segmentation, stop word removal, and word vectorization processing of text data, and denoising, frame segmentation, and feature extraction of audio data; S3. Use the improved Paxos consistency algorithm to ensure data consistency among multiple nodes, ensure data consistency in the case of network partitioning or node failures through a distributed consensus mechanism, and use the distributed database Apache Cassandra for data storage; S4. Deploy the Kubernetes resource management framework to monitor the usage of computing resources of each node in real time, dynamically adjust the computing task allocation according to the node load conditions, and balance the task load of each node through the Nginx load balancing technology; S5. In a distributed environment, use the Prometheus distributed monitoring system to monitor the node status in real time, automatically detect node failures, and trigger task reallocation and data backup when a failure is detected; S6. Adopt the TensorFlow Distributed distributed training framework, combine data parallelism and model parallelism technologies, distribute large-scale data and model parameters to multiple computing nodes, and perform parallel training to accelerate the model training process; S7. Perform further processing and feature extraction on the multi-modal data; S8. Use the attention mechanism to fuse the data features of different modalities to generate a comprehensive feature representation, and improve the training effect and comprehensive performance of the multi-modal AI model.
2. The multimodal AI model training method based on a distributed network architecture according to claim 1, wherein, The specific content of S1 includes: S11. In a multi-node environment, distribute and store multi-modal data in the distributed file system HDFS. The multi-modal data includes image data, text data, and audio data, and each node stores and accesses data through the distributed file system HDFS; S12. Store and manage image data through the distributed file system HDFS, including block storage and redundant storage of image data to ensure the high availability and reliability of data in the case of node failures; S13. For text data, perform distributed storage and index management through the distributed file system HDFS to ensure the data reading efficiency in a large-scale text data environment, and use the inverted index technology to improve the retrieval speed of text data; S14. For audio data, use the distributed file system HDFS for distributed storage, and ensure the efficient storage and transmission of audio data in a distributed node environment through the sharding and multi-copy mechanism of audio data; S15. In the process of data transmission between multiple nodes, use the gRPC protocol to achieve efficient data transmission, including: Define the gRPC service and message format to ensure unity and standardization during data transmission between multiple nodes; utilize the streaming transmission function of gRPC to transmit multi-modal data in real time, specifically including the streaming transmission of image data, text data, and audio data; During the transmission process, use data compression technology to reduce the amount of data transmitted and improve data transmission efficiency, specifically including JPEG compression for image data, gzip compression for text data, and FLAC compression for audio data; Establish an efficient data transmission channel between nodes, and utilize the multiplexing and load balancing functions of gRPC to ensure the stability and efficiency of data transmission; To ensure the security of data transmission, adopt the TLS protocol to encrypt the gRPC data transmission channel to prevent data from being eavesdropped and tampered with during transmission; During the data transmission process, monitor the transmission status in real time, record the transmission log, and use log analysis tools to detect and handle anomalies during the transmission process; After the data transmission is completed, perform data verification and confirmation on each node to ensure the integrity and consistency of the transmitted data. Use check codes to verify the transmitted data, and judge whether the data transmission is successful based on the verification result.
3. A method for training a multimodal AI model based on a distributed network architecture according to claim 1, characterized in that, The improved Paxos consistency algorithm of S3 specifically includes: S31. Run Paxos instances in a multi-node environment, including three roles: proposer, acceptor, and learner. Optimize the message passing mechanism by batch processing messages and message compression technology to reduce communication overhead. The formula for the message compression technology is: Among them, C msg represents the size of the compressed message, S i represents the size of the i-th message, and N represents the number of messages; S32. The proposer receives requests from clients and generates proposals. A proposal contains a unique proposal number n and a proposal value v, and multiple proposals are merged into one message for sending. The proposer sends Prepare messages to multiple acceptors. The formula for batch processing is: Among them, B msg represents the message after batch processing, P j represents the j-th proposal, ∈ represents the coefficient, and m represents the number of proposals; S33. After receiving the Prepare message, if the number n is greater than the maximum number it has responded to, the acceptor accepts the Prepare message and promises not to accept proposals with numbers less than n. At the same time, the acceptor replies with a Promise message to the proposer. The Promise message contains the number of the proposal with the maximum number accepted by the acceptor and the corresponding proposal value; S34. After receiving Promise messages from a majority of acceptors, the proposer immediately generates a proposal with number n and proposal value v, and sends an Accept request to the acceptors. Improve the message processing speed through multi-threaded parallel processing. The formula for multi-threaded parallel processing is: Among them, T process represents the total time of parallel processing, T i represents the processing time of the i-th thread, k represents the number of threads, and α represents a constant coefficient; S35. After receiving the Accept request, if the number n is greater than or equal to the maximum number it has responded to, the acceptor accepts the proposal and replies with an Accepted message to the proposer. The Accepted message contains the number and proposal value of the proposal. Adopt a dynamic priority mechanism. The formula for the dynamic priority mechanism is: Among them, P priority represents the proposal priority, L represents the proposal delay, T represents the proposal importance, R represents the proposal resource consumption, and β is a small positive number to avoid a zero denominator; S36. After receiving Accepted messages from a majority of acceptors, the proposal is considered to have reached a consensus. The proposer broadcasts the proposal value to all learners, and reduces the broadcast delay through hierarchical broadcast and intelligent broadcast strategies. The formula for hierarchical broadcast is: Among them, H broadcast represents the total time of hierarchical broadcasting, B l represents the broadcasting time of the l-th layer, and h represents the number of layers; S37. After the learner receives the notice from the proposer, it records the proposal value and applies this value. By adopting an efficient consistency check technology, it quickly checks data consistency before each write operation. The formula for the efficient consistency check is as follows: Among them, E check represents the total time for consistency check, H(D i ) represents the hash value calculation time of the i-th data block, n represents the number of data blocks, and γ represents a constant coefficient; S38. In the case of network partitioning or node failure, the acceptor and the proposer will renegotiate the proposal according to the steps of the improved Paxos consistency algorithm, and introduce a fault node isolation and dynamic priority mechanism; S39. Use the distributed database Apache Cassandra for data storage. Combine the incremental update strategy and the improved Paxos consistency algorithm to achieve strong consistency storage. Before each write operation, reach a consensus among multiple nodes through an optimized consensus mechanism. The formula for the incremental update strategy is as follows: Among them, Δ update represents the amount of incrementally updated data, δ i represents the size of the data updated for the i-th time, and m represents the number of updates.
4. A method for training a multimodal AI model based on a distributed network architecture according to claim 1, characterized in that S4 specifically includes: S41. In a multi-node environment, install and configure the Kubernetes resource management framework, and deploy the Kubernetes node agent on each computing node; S42. Use the resource monitoring module of Kubernetes to collect and monitor the CPU, memory, and network bandwidth resource usage of each node in real time, and summarize the monitoring data to the Kubernetes master node; S43. Design and implement a dynamic priority scheduling algorithm. Based on real-time monitoring data and historical resource usage, the algorithm predicts the load change trend of each node, preferentially allocates resources to nodes with lighter loads, and improves resource utilization. The priority P i (t) is calculated by the following formula: Among them, P i (t) represents the priority of node i at time t, and CPU i (t) represents the CPU usage rate of node i at time t, and MEM i (t) represents the memory usage rate of node i at time t. α, β, and γ represent weight coefficients, and H i (t) represents the historical load change trend of node i at time t, which is calculated by the exponentially weighted moving average of historical data: Among them, λ represents the smoothing coefficient, and its value range is 0 < λ ≤ 1; S44. Configure the Nginx load balancer. By setting reverse proxy and load balancing rules, evenly distribute user requests and data traffic to different computing nodes; S45. Design and implement a load balancing optimization algorithm based on machine learning, which dynamically adjusts the load balancing rules of Nginx by using the historical load data and current status of nodes; use a linear regression model to predict the load condition Li(t + Δt) of node i after time t i (t + Δt): L i (t + Δt) = θ0 + θ1·CPU i (t) + θ2·MEM i (t) + θ3·Net i (t); Among them, L i (t + Δt) represents the predicted load condition of node i at time t + Δt, θ0, θ1, θ2, θ3 represent the parameters of the linear regression model, and CPU i (t) represents the CPU usage rate of node i at time t, and MEM i (t) represents the memory usage rate of node i at time t, and Net i (t) represents the network bandwidth usage rate of node i at time t; the load balancing rule is dynamically adjusted according to the predicted load condition, and the node j with the lightest load is selected to process the new task: Among them, j represents the node with the lightest load selected; S46. When it is detected that a certain node has a resource bottleneck or failure, use the automatic task rescheduling function of Kubernetes to reassign the affected tasks to other nodes with sufficient resources; S47. Introduce a scheduling decision record mechanism based on blockchain technology, record the decisions of each task scheduling and resource allocation on the blockchain, and calculate the hash value H of the blockchain record as follows: H = SHA256(T i ||P i (t)||L i (t + Δt)); Among them, H represents the hash value recorded by the blockchain, SHA256 represents the hash function, T i represents the task ID, P i (t) represents the priority of node i at time t, L i (t + Δt) represents the predicted load situation of node i at time t + Δt, and || represents the string concatenation operator, which is used to concatenate multiple values into a string for hash calculation.
5. A method for training a multi-modal AI model based on a distributed network architecture according to claim 1, characterized in that, S7 specifically includes: using a convolutional neural network to extract features from image data; using a long short-term memory network to extract features from text data; using Mel spectrogram to transform audio data, and using a convolutional neural network to extract features.
Citation Information
Patent Citations
Distributed database automatic operation and maintenance method and system based on artificial intelligence
CN112579391A
Support system for designing an artificial intelligence application, executable on distributed computing platforms
US20210064346A1