Digital twin hydraulic engineering large language model reasoning method based on tensor parallel processing

By adopting a tensor parallel processing architecture in the water conservancy digital twin system, the calculation task decomposition and parallel processing of the large language model are optimized, and the problems of high computational complexity and slow response speed are solved, and efficient processing of large-scale water conservancy data and rapid prediction of dam safety risks are achieved.

CN119918680AActive Publication Date: 2025-05-02CHANGJIANG RIVER SCI RES INST CHANGJIANG WATER RESOURCES COMMISSION +2

Patent Information

Application Number
CN202411920140.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-02
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The large language model has high computational complexity and slow inference response in the water conservancy digital twin system. Especially when processing massive, extensive and multi-source water conservancy data, traditional computing architectures are prone to bottlenecks, resulting in resource waste and processing delays.

Method used

Using a digital twin water conservancy large language model inference architecture based on tensor parallel processing, through gradient calculation optimization, self-attention tensor dimensionality reduction blocking, low-dimensional layer normalization parallelism, and overall weak expansion resource scheduling processing, large-scale computing tasks are decomposed into multiple tensor computing tasks and allocated to multiple computing nodes for parallel processing.

Benefits of technology

It significantly improves the inference efficiency of large-scale water conservancy models, shortens reaction time, improves real-time monitoring and early warning capabilities, and provides more accurate and efficient solutions for the safety management of dams and other water conservancy facilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005207764370000031
    Figure BDA0005207764370000031
  • Figure BDA0005207764370000041
    Figure BDA0005207764370000041
  • Figure BDA0005207764370000042
    Figure BDA0005207764370000042
Patent Text Reader

Abstract

The invention provides a digital twin hydraulic engineering large language model reasoning method based on tensor parallel processing. The method sequentially comprises the following steps: S1, gradient calculation optimization; s2, carrying out self-attention tensor dimension reduction blocking; s3, normalizing and parallelizing a low-dimensional layer; s4, inter-layer overall weak extension resource scheduling processing is carried out on the low-dimensional layer after middle-layer normalization processing in the step S3; and S5, converting an intermediate tensor result of a decoder into a specific predicted value. According to the method, the parameters and the input data of the large-scale language model are subjected to tensor segmentation and are distributed to the computing nodes for parallel processing, so that the reasoning efficiency of the large-scale model is remarkably improved. By using the method, the calculation efficiency and the real-time performance of the hydraulic engineering large language model can be greatly improved, the calculation nodes can work simultaneously by processing the reasoning task of the large-scale model in parallel, the calculation burden of a single node is remarkably reduced, larger-scale hydraulic engineering data can be processed, and the calculation efficiency and the real-time performance of the hydraulic engineering large language model are improved. And efficient and accurate risk prediction capability can be maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and water conservancy project digital twin technology, and in particular to a digital twin water conservancy large language model reasoning architecture and method based on tensor parallel processing. Technical Background

[0002] As an innovative engineering management technology, digital twins provide a new observation and management perspective for water conservancy projects by virtualizing the physical environment. In the monitoring of water conservancy facilities such as dams, digital twin technology can collect and simulate various physical parameters (such as water level, flow, pressure, etc.) in real time to create a high-precision virtual model. Combined with a large language model, this virtual environment can achieve detailed analysis and prediction of actual facilities. The introduction of digital twins not only provides an intuitive reference for the operation and maintenance of facilities, but also plays an important role in preventive maintenance and disaster warning, and improves the ability to respond to extreme weather events and other potential risks.

[0003] In the digital twin of water conservancy projects, the application of tensor parallel technology has greatly improved the efficiency of data analysis of large language models. Tensor parallelism maximizes the use of computing resources by breaking down complex computing tasks into multiple subtasks that can be processed simultaneously, thereby accelerating the data processing process. In complex water conservancy facilities such as dams, real-time processing of huge amounts of sensor data is key to ensuring facility safety. With the help of tensor parallelism, the response time can be significantly shortened while ensuring computing accuracy. This fast computing capability is essential for timely prediction and handling of potential risks.

[0004] Although large language models have powerful capabilities in data analysis, they still face challenges in computing efficiency. Especially in dam safety risk prediction, the model needs to process massive, extensive and multi-source data. In this environment, traditional computing architectures are prone to bottlenecks, resulting in waste of resources and processing delays. This delay may lead to inaccurate predictions when responding to emergencies. In order to overcome this problem, the tensor parallel architecture is applied, and its high-dimensional data processing capabilities that are unique to other technologies can be used to optimize equipment and algorithms, effectively improving the reasoning efficiency of large language models. Through the adoption of a distributed architecture, tensor parallelism not only significantly improves the speed of large-scale real-time computing, but also improves the allocation and use efficiency of resources. Ultimately, this improvement provides a more accurate and efficient solution for the safety management of dams and other water conservancy facilities in terms of real-time monitoring and early warning capabilities. Summary of the invention

[0005] Aiming at the problems of high computational complexity and slow reasoning response speed of large language models in water conservancy digital twin systems, the present invention proposes a digital twin water conservancy large language model reasoning method based on tensor parallel processing. This method applies tensor parallel architecture and utilizes its high-dimensional data processing capabilities that are unique to other technologies to optimize equipment and algorithms, effectively improving the reasoning efficiency of large language models; through the adoption of distributed architecture, tensor parallelism can not only significantly improve the speed of large-scale real-time computing, but also improve the allocation and use efficiency of resources; ultimately, this improvement in real-time monitoring and early warning capabilities provides a more accurate and efficient solution for the safety management of dams and other water conservancy facilities.

[0006] To achieve the above technical objectives, the present invention provides a digital twin water conservancy large language model reasoning method based on tensor parallel processing, which specifically includes the following steps:

[0007] S1. Gradient calculation optimization: collect water conservancy data in real time as the input of the digital twin water conservancy language model, optimize the gradient calculation process through the SUMMA algorithm, and distribute the processing matrix so that each computing node can independently complete its own computing task. The optimized gradient information obtained can realize parallel processing of water conservancy data collected in real time;

[0008] S2. Dimensionality reduction and partitioning of self-attention tensor; The gradient information optimized in step S1 is used as the basis for the self-attention mechanism, and the self-attention mechanism is implemented using the Transformer model. Long sequence data is processed in parallel by row and column partitioning to achieve the purpose of dimensionality reduction and partitioning of the gradient information optimized in step S1;

[0009] S3. Normalize and parallelize the low-dimensional layer: normalize and calculate the low-dimensional layer of the reduced-dimensional blocks in step S2 using vector parallelization;

[0010] S4. Perform inter-layer overall weak expansion resource scheduling for the low-dimensional layer after the layer normalization processing in step S3; after the layer normalization calculation in step S3 is stable, distribute the computing load to different computing nodes, so that each computing node is responsible for the calculation of a specific layer, and each node completes the calculation of the specific layer and generates the inference intermediate tensor result corresponding to the layer;

[0011] S5. The decoder is used based on the fully connected layer and the Softmax layer to convert the intermediate tensor results of the model into specific prediction values.

[0012] A further technical solution of the present invention: the water conservancy data in step S1 is collected in real time by a variety of sensors arranged in water conservancy facilities, specifically including water level, flow velocity, flow, seepage pressure, and surface movement data.

[0013] A further technical solution of the present invention is as follows: In the step S1, the gradient calculation process is optimized by the SUMMA algorithm. A two-dimensional matrix multiplication scheme of an extensible general matrix multiplication algorithm is adopted. In the context of deep learning, the system performs matrix multiplication C=A×B output by the objective function, and its differential, i.e., gradient, is calculated by the chain method to form multiple sub-matrices. These sub-matrices are then assigned to different computing nodes, and each node independently calculates the gradient of the sub-matrix it is responsible for. The calculated gradient will serve as an important basis for subsequent model updates. The specific calculation process is as follows:

[0014] Let C' be the gradient of the objective function with respect to C, A' and B' be the gradients with respect to the input matrices A and B respectively, then:

[0015]

[0016] Where C′ is the gradient of the objective function with respect to C;

[0017] A′ and B′ are the gradients with respect to the input matrices A and B respectively, represents partial derivative.

[0018] A further technical solution of the present invention is as follows: In the step S2, the dimensionality reduction is first performed through linear projection, and the input hidden vector is mapped into a query (Q), a key (K), and a value (V). The attention weight is calculated through matrix multiplication and multiplied with the value V to obtain the output of the self-attention module; wherein the original hidden vector X is first projected into a query Q, a key K, and a value V respectively through three independent linear layers:

[0019] Q=XW Q +b Q

[0020] K=XW K +b K

[0021] V=XW V +b V

[0022] Among them, W Q , W K and W V is the learnable weight matrix, b Q , b K and b V is bias;

[0023] After using the linear transformation, the query and the key are dot-producted to get the attention score, and then the attention weight A is obtained by scaling and softmax transformation:

[0024]

[0025] Among them, A is the calculated attention weight matrix; Q is the query matrix, which represents the input query information;

[0026] K is the key matrix, which represents the input key information; K T is the transpose of the bond matrix, which serves as the basis for calculating the similarity measure;

[0027] d k is the dimension of the key, and the scaling operation helps avoid extreme values ​​in the softmax;

[0028] The softmax function converts the input into a probability distribution so that its sum is 1. The attention weight A obtained in this way is multiplied by the value V, and the weighted sum is calculated to obtain the output of the self-attention module:

[0029] Output=AV

[0030] Finally, the resulting output is resized back to the hidden size and processed through another linear layer to produce the final output of the self-attention module:

[0031] Final Output=Output W O +b O

[0032] Among them, W O is the weight matrix of the second linear layer; b O is its bias.

[0033] A further technical solution of the present invention is as follows: after the self-attention module completes the calculation, the redundant low-dimensional layer of step S2 is normalized in step S3, and each processor independently calculates the mean and variance of the input data and performs a normalization operation; the calculation of the layer normalization can be expressed as:

[0034]

[0035] in, represents the normalized output; Y represents the input data tensor; μ represents the mean E[Y];

[0036] σ 2 Represents the variance Var[Y]; δ represents a small constant to prevent division by zero errors; it is usually used to prevent division by zero errors and maintain numerical stability;

[0037] The mean μ and variance σ 2 The calculation formula is:

[0038]

[0039] σ 2 =E[Y 2]-μ 2

[0040] Where m is the number of samples;

[0041] E[Y 2 ] is the expected value of the square of the input data, which represents the average value of the squares of all elements in the input data vector Y;

[0042] Y j is the jth element in the input sample.

[0043] A further technical solution of the present invention is as follows: for the calculation of the mean and variance, the processor aggregates the calculation results and uses the all-reduce function to summarize the data of all computing nodes to obtain accurate global mean and variance information. For the gradient calculation of Y, the function description is as follows:

[0044]

[0045] in, represents partial derivative.

[0046] A further technical solution of the present invention is as follows: In the step S4, the computing load is distributed to different computing nodes so that each node is responsible for computing a specific layer.

[0047] Assuming the total number of layers is L, if it is assigned to N GPUs, each GPU will be responsible for L layers. i It can be expressed as:

[0048]

[0049] When processing time series data, the input of multiple time steps is divided into small batches and assigned to different GPUs for independent calculations. If the number of time steps is set to T, then each GPU processes time steps T. j It can be expressed as:

[0050]

[0051] A further technical solution of the present invention: In the S5 step, a high-speed communication mechanism is used to integrate the intermediate tensor results of all computing nodes to form a complete global tensor, ensuring that the intermediate results of all nodes are synchronized and consistent; the global tensor contains complete water conservancy feature data; the fully connected layer maps the input global tensor to the output space through a linear transformation, extracts basic information related to the safety value prediction, and outputs a feature vector related to the safety value prediction, and the Softmax layer probabilistically processes the feature vector output by the fully connected layer to generate a probability distribution for each safety level; based on the probability distribution generated by the Softmax layer, the safety level is divided in combination with the following classification rules, and finally a risk level prediction result is generated; wherein, the safety level classification rules for the probability P of the risk level are as follows: P>0.6, predicted as low risk; 0.3≤P≤0.6, predicted as medium risk; P<0.3, predicted as high risk.

[0052] The present invention decomposes the computing task of a large language model into multiple tensor computing tasks through a tensor parallel processing architecture, and distributes them to multiple computing nodes for parallel processing. First, the system splits the parameters and input data of the large language model into small-scale tensors and distributes them to different computing nodes. Each node independently completes its own reasoning calculation, and then ensures the consistency and synchronization of the intermediate results through high-speed communication between nodes. Finally, the system summarizes, corrects and post-processes the reasoning results of all nodes, and outputs the final dam safety risk prediction results. During the entire reasoning process, the system adjusts the computing load in real time through dynamic load balancing and node monitoring to ensure efficient and stable operation of the system. Inter-layer overall weak expansion resource scheduling is an important strategy to improve the training and reasoning efficiency of large-scale deep learning models. Through reasonable resource monitoring, dynamic adjustment, mixed precision calculation and efficient communication, the utilization rate of computing resources and the overall performance of the model can be greatly improved.

[0053] Through the tensor parallel processing method, the system realizes efficient parallel reasoning of large-scale water conservancy language models, significantly improving the calculation speed and real-time performance of model reasoning. The final system can output dam safety risk prediction results in a timely manner while processing massive monitoring data, ensuring rapid response and accurate assessment of dam safety hazards. Compared with the traditional serial reasoning method, the present invention not only reduces the computational burden of a single node through multi-node parallel computing, but also improves the overall computing efficiency of the system, providing a more economical and efficient solution for the digital twin technology of water conservancy projects.

[0054] Beneficial effects of the present invention:

[0055] (1) The present invention significantly improves the reasoning efficiency of large-scale water conservancy models by dividing the computational tasks of a large language model into multiple tensors for parallel processing. The system effectively reduces the computational burden of a single node by working in parallel with multiple computing nodes, and achieves rapid prediction of dam safety risks through high-speed communication and synchronization mechanisms. This method ensures a balance between real-time and accuracy, and significantly improves the reasoning response speed and system stability.

[0056] (2) The present invention effectively improves the ability of large language models to process large-scale data through multi-node parallel computing. By utilizing a dynamic load balancing mechanism, it can continuously optimize the computing resource utilization of each node to ensure efficient and stable operation of the system. Compared with the prior art, the present invention can not only significantly improve the reasoning speed, but also reduce the consumption of computing resources and improve the real-time response capability, thus providing a more economical and efficient solution for large-scale water conservancy data processing and dam safety risk management.

[0057] (3) The present invention specifically uses tensor parallel technology to split model parameters and input data into small-scale tensors, and performs parallel reasoning in a distributed manner, thereby achieving efficient water conservancy large language model reasoning response. This method can not only achieve efficient parallel processing of model reasoning, but also ensure the efficiency and stability of the system through dynamic load balancing and inter-node communication optimization, effectively solving the problems of slow reasoning speed and high resource usage in the prior art. DETAILED DESCRIPTION

[0058] The present invention is further described below in conjunction with embodiments. The technical solutions shown in the following embodiments are specific solutions of the present invention and are not intended to limit the scope of the invention claimed for protection. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0059] The embodiment provides a digital twin water conservancy large language model reasoning method based on tensor parallel processing, which can significantly improve the reasoning efficiency of the large language model in the water conservancy digital twin system and realize real-time processing and efficient response of large-scale water conservancy data; the specific steps are as follows:

[0060] S1. Gradient calculation optimization: Real-time water conservancy data is collected as the input of the digital twin water conservancy language model. The gradient calculation process is optimized through the SUMMA algorithm, and the matrix is ​​processed in a distributed manner so that each node can complete its own calculation task independently. The optimized gradient information obtained can realize parallel processing of real-time collected water conservancy data.

[0061] 64 GPUs are used in the implementation, each GPU is equipped with 32GB video memory, and the nodes are connected by a 10Gbps high-speed network. The real-time collected water conservancy data, including sensors measuring water level, flow velocity, flow, seepage pressure, surface movement and other data, reaches GB level, the sequence length is set to 512, and the batch size is 128. These real-time collected physical data are used as the input of the large language model. The data stream is not only huge, but also needs to meet extremely high requirements in terms of real-time performance. In order to meet the needs of efficient processing of large-scale sensor data streams, the system adopts gradient calculation optimization technology. Specifically, the gradient calculation process is optimized through the SUMMA algorithm to make it more suitable for distributed environments and achieve efficient and real-time computing capabilities. In order to meet the needs of efficient processing of large-scale sensor data streams, the system adopts gradient calculation optimization technology. The sensor data is decomposed into small batches and input into the large language model. Each GPU calculates the average distribution of sub-matrix gradients. For example, GPU1 is responsible for the 1st to 16th columns of the gradient matrix; GPU2 is responsible for the 17th to 32nd columns of the gradient matrix, and so on. The optimized gradient information is used to update the model parameters to ensure that the model can capture the latest dynamic information. The method involves a two-dimensional matrix multiplication scheme that is a scalable general matrix multiplication algorithm, which is particularly suitable for large-scale parallel computing. In the context of deep learning, the system uses the chain method to calculate the differential (gradient) of the matrix multiplication C=A×B output by the objective function. Multiple sub-matrices are formed, and then these sub-matrices are assigned to different computing nodes. Each node independently calculates the gradient of the sub-matrix it is responsible for, reducing the global computing bottleneck. Let C′ be the gradient of the objective function with respect to C, A′ and B′ are the gradients with respect to the input matrices A and B respectively, then:

[0062]

[0063] Where C′ is the gradient of the objective function with respect to C;

[0064] A′ and B′ are the gradients with respect to the input matrices A and B respectively; represents partial derivative.

[0065] In the context of water conservancy facility monitoring, gradient calculation is used to optimize the model's prediction of the safety status of dams or levees. Specifically, the system divides the matrix output by the objective function (such as C = A × B), and through the chain rule, the calculated gradient will serve as an important basis for subsequent model updates. This parallel processing reduces the global computational bottleneck, so that each processor only needs to focus on the calculation of the sub-matrix it manages. After the calculation is completed, each processor aggregates the local calculation results to a centralized location, or merges all partial results through repeated all-reduce processes to obtain global gradient information. In this way, the distribution of the obtained gradient matrix is ​​consistent with the partial division of the original matrices A and B. By adopting the SUMMA algorithm, the complexity of gradient calculation is significantly reduced, allowing large-scale parallel processing and effective use of computing resources, thereby significantly improving the speed and efficiency of model training. The gradient information optimized by step S1 will provide the basis for the self-attention mechanism in step S2, ensuring that the model can perform efficient reasoning with the latest parameter state when performing further calculations.

[0066] Table 1 Model segmentation parameter settings

[0067] parameter value describe Number of tensor splits 4 Split the model's tensor into 4 parts Number of computing nodes 4 Each compute node is responsible for a tensor part Network bandwidth 10Gbps Ensure high-speed communication between nodes Memory per node 32GB Available memory per compute node Partitioning strategy Even Partition Evenly partition tensors to achieve load balancing Data Locality high Optimizing data locality to reduce latency

[0068] S2. Dimensionality reduction and block partitioning of the self-attention tensor; after completing the gradient calculation, the system enters the implementation stage of the self-attention mechanism, and improves the computational efficiency of the model by dimensionality reduction and block partitioning. Dimensionality reduction first undergoes linear projection to map the input hidden vector to query (Q), key (K), and value (V). The attention weight is calculated by matrix multiplication and multiplied by the value V. Each GPU calculates the assigned part of the matrix. For example, GPU1 is responsible for the first 128 time steps; GPU2 is responsible for the 129-256 time steps, thereby obtaining the output of the self-attention module. The original hidden vector X is first projected into query Q, key K, and value V through three independent linear layers:

[0069] Q=XW Q +b Q

[0070] K=XW K +b K

[0071] V=XW V +b V

[0072] Among them, W Q , W K and W V is a learnable weight matrix; b Q , b K and b V is bias.

[0073] After using the linear transformation, the query and the key are dot-producted to get the attention score, and then the attention weight A is obtained by scaling and softmax transformation:

[0074]

[0075] Among them, A is the calculated attention weight matrix; Q is the query matrix, which represents the input query information;

[0076] K is the key matrix, which represents the input key information; K T is the transpose of the bond matrix, which serves as the basis for calculating the similarity measure;

[0077] d k is the dimension of the key, and the scaling operation helps avoid extreme values ​​in the softmax;

[0078] The softmax function transforms the input into a probability distribution so that its sum is 1.

[0079] The obtained attention weight A is multiplied by the value V, and the weighted sum is calculated to obtain the output of the self-attention module:

[0080] Output=AV

[0081] Finally, the resulting output is resized back to the hidden size and processed through another linear layer to produce the final output of the self-attention module:

[0082] Final Output=Output W O +b O

[0083] Among them, W O is the weight matrix of the second linear layer; b O is its bias.

[0084] The output of the self-attention module is obtained through the above steps. By dividing the tensor into small-scale sub-tensors and processing them in parallel, the processing capability of long sequence data can be improved. The self-attention module and MLP each have two linear layers, which facilitates row and column partitioning (tensor partitioning), greatly improving the computational efficiency and parallel processing capability. The model can not only process long sequence data efficiently, but also make full use of the advantages of parallel computing on multi-GPU systems, thereby improving the overall performance and efficiency of the model. This is the method of dimensionality reduction and block division. The self-attention mechanism relies on the optimized gradient calculated in step S1, so that the model can update the attention weight according to the latest gradient information. In addition, the low-dimensional layer normalization in step S3 is performed after the output of the self-attention module to ensure that the model remains stable during processing. This process focuses on resource scheduling after differentiation.

[0085] Table 2 Parallel reasoning parameter settings

[0086]

[0087] S3. Parallel normalization of low-dimensional layers; the low-dimensional layers of the dimensionality reduction blocks in step S2 are normalized using vector parallelization; after the self-attention module is calculated, the redundant low-dimensional layers of S2 are normalized to further improve the stability and convergence speed of the model. Each processor independently calculates the mean and variance of the input data and performs normalization operations to avoid global communication delays. Through vector parallel technology, each node can effectively use local results to ensure efficient calculations; for, the calculation of layer normalization can be expressed as:

[0088]

[0089] in, represents the normalized output; Y represents the input data tensor;

[0090] μ represents the mean E[Y]; σ 2 represents the variance Var[Y];

[0091] δ represents a small constant used to prevent division by zero errors. It is usually used to prevent division by zero errors and maintain numerical stability.

[0092] Among them, the mean μ and variance σ 2 The calculation formula is:

[0093]

[0094] σ 2 =E[Y 2 ]-μ 2

[0095] Where m is the number of samples;

[0096] E[Y 2 ] is the expected value of the square of the input data, which represents the average value of the squares of all elements in the input data vector Y;

[0097] Y j is the jth element in the input sample.

[0098] Each processor computes the input X and X 2 The mean and variance of Y can be calculated to avoid unnecessary global communication; in each residual connection, the bias addition operation is broadcasted to expand the matrix to each column so that it can be executed smoothly in the forward process. During backpropagation, the gradient is restored to the processor in row 0. The designed aggregation function summarizes the local mean and variance results of each computing node. During backpropagation, the aggregation function propagates the gradient of Y to the world to obtain accurate global mean and variance information, ensuring that each node can obtain the complete global gradient. The normalized output function designed for the gradient calculation of Y is described as follows:

[0099]

[0100] in, represents partial derivative.

[0101] Through such processing, layer normalization can maintain stability during training and achieve more uniform and rapid convergence for each batch of data. The use of vector parallel technology enables each node to effectively utilize local results in large-scale parallel computing, reducing the delay of global communication. The system finally outputs the prediction results of dam safety risks. This process ensures the real-time and accuracy of the prediction, and effectively improves the reasoning efficiency of large-scale language models. The above steps are derived to show the parallel method of the feature tensor after the dimension reduction of the self-attention tensor. The low-dimensional layer normalization ensures that the output of the self-attention module in S2 can be stabilized in subsequent steps. At the same time, this process also lays the foundation for the weak expansion resource scheduling in step S4, ensuring the efficient use of computing resources.

[0102] S4. Perform inter-layer overall weak expansion resource scheduling for the low-dimensional layers after the layer normalization in step S3; after the layer normalization calculation in step S3 is stable, distribute the computing load to different computing nodes so that each node is responsible for the calculation of a specific layer. Each node completes the calculation of a specific layer and generates the corresponding inference intermediate result of the layer; the decoder converts the intermediate tensor results of the model into specific prediction values ​​based on the fully connected layer and the Softmax layer.

[0103] In this step, after the stabilization of layer normalization, the system distributes the computing tasks to different GPUs by layer. For example, GPU1 is responsible for the calculation of layers 1-4; GPU2 is responsible for layers 5-8, and so on, so that each node is responsible for the calculation of a specific layer. In the time series, the data is divided into small batches and assigned to different GPUs for independent calculation. This batch processing method can significantly speed up the overall processing speed. Weak expansion means that when the number of computing nodes is increased, the load of each node is uneven, and the overall workload and structure remain unchanged. Blocks are divided at different layers so that each GPU is only responsible for the calculation of certain layers. Assuming the total number of layers is L, if it is assigned to N GPUs, each GPU is responsible for the number of layers L. i It can be expressed as:

[0104]

[0105] This can effectively reduce the memory pressure of each node, thereby improving computing efficiency. When processing time series data, the input of multiple time steps can be divided into small batches and assigned to different GPUs for independent calculations. Set the number of time steps to T, then each GPU processes time steps T j It can be expressed as:

[0106]

[0107] Partitioning allows multiple GPUs to process time series in parallel, which significantly accelerates the overall computing process. Weakly extended resource scheduling ensures that data can be processed efficiently and achieves real-time prediction results. Finally, after multiple steps of processing, the system can quickly and accurately output the dam safety risk prediction results, providing timely decision support for water conservancy project managers. The comparison table of parameters for three common different network structures is shown in Table 3. For the three different network structures, the corresponding hidden layer sizes are set respectively, and the performance is evaluated by the following indicators: throughput is defined as the ratio of batch size to the sum of forward and backward propagation time in each iteration, while inference efficiency is the ratio of batch size to the sum of forward propagation time. Under weakly extended operating conditions, each GPU uses the same memory configuration, but due to differences in inter-GPU communication requirements, experiments using fewer GPUs usually perform better in terms of throughput and inference efficiency. In the weakly extended setting, by optimizing resource scheduling and configuration, the computational efficiency and inference ability of the model are significantly improved, which effectively promotes the application of deep learning models in large-scale distributed systems. Based on the resource naturalization at the tensor level, the total scheduling of block-parallel computation is realized through multi-pole parallelization, which further optimizes the computing performance. A performance comparison experiment was conducted on the NVIDIA A100 GPU cluster to verify the improvement effect of the optimization strategy on large-scale distributed deep learning tasks.

[0108] When using 64 GPUs, Tesseract's 4D tensor parallel throughput reaches that of Megatron-LM:

[0109]

[0110] Compared with Optimus's optimized parallelism:

[0111]

[0112] For inference, Tesseract's four-dimensional tensor parallelism reaches the parallelism of Megatron-LM giant tensor language model:

[0113]

[0114] Compared with Optimus's optimized parallelism:

[0115]

[0116] This usage effect supports Tesseract's ability to better utilize resources on GPU servers than other 1-D and 2-D tensor parallel methods. In a weakly scalable setting, optimized resource scheduling and configuration can significantly improve the model's computational efficiency and reasoning capabilities, thereby better promoting the application of deep learning models in large-scale distributed systems. The multi-polar parallelization of resource naturalization at the tensor level mentioned above is the total scheduling method for block-parallel computing.

[0117] Table 3 Horizontal comparison of different GPU models

[0118] Parallelization GPUs GPU form factor Hidden capacity Throughput Reasoning Tesseract 4 [2,2,1] 512 0.720 2.971 16 [4,4,1] 1024 0.502 2.136 64 [4,4,4] 1024 0.510 2.218 Optimus 4 [2,2] 512 0.783 3.041 16 [4,4] 1024 0.325 1.217 64 [8,8] 2048 0.302 1.142 Megatron-LM 4 [2,2,1] 512 0.765 3.114 16 [4,4,1] 1024 0.343 1.142 64 [4,4,4] 2048 0.143 0.524

[0119] The intermediate tensor results calculated by each GPU are aggregated through high-speed communication, and the aggregated tensors are passed through the final decoding layer to generate output. The decoding result is the real-time safety risk prediction value of the dam, the normalized feature tensor obtained from the inter-layer resource scheduling, and then the input low-dimensional feature tensor is mapped to the output space through linear transformation to obtain the basic information required for prediction.

[0120] Specifically, a distributed computing framework is used, through the tensor parallel mechanism, each computing node independently processes its assigned feature tensor to achieve parallel computing. A high-speed communication mechanism (such as all-reduce) is used to integrate the intermediate tensor results of all computing nodes to form a complete global tensor, ensuring that the intermediate results of all nodes are synchronized and consistent. The intermediate tensor results contain a multi-dimensional feature representation of water conservancy data that has been processed (such as gradient optimization, self-attention mechanism processing and normalization). The integrated global tensor contains complete water conservancy feature data, covering core dimensions such as future water level prediction and flow rate monitoring, providing a complete and reliable feature basis for subsequent risk prediction.

[0121] The decoder module consists of two parts: the fully connected layer and the Softmax layer. Based on the fully connected layer and the Softmax layer, key features are extracted from the integrated global tensor and the security level prediction results are generated, which are expressed as specific probability distributions or classification results. The fully connected layer maps the input global tensor to the output space through linear transformation, extracts basic information related to the security value prediction, performs dimensionality reduction operations on high-dimensional tensors, retains key features to reduce computational complexity, and generates feature representations containing basic features of risk levels of each category, providing input for the probabilistic processing of the Softmax layer. The Softmax layer probabilistically processes the feature vector output by the fully connected layer, generates a probability distribution for each security level, and uses the Softmax function to convert the linear output into a probability value ranging from 0 to 1, ensuring that the sum of the probabilities of all categories is 1, which is convenient for subsequent classification and analysis. Based on the probability distribution generated by the Softmax layer, the safety level is divided into the following classification rules: Low safety risk: P>0.6 Medium safety risk: 0.3≤P≤0.6 High safety risk: P<0.3 By setting the threshold, the model can accurately classify samples according to different probability distributions, and the output classification results can be used for risk warning and management in actual water conservancy projects.

[0122] Predict the future trend of key parameters of water conservancy facilities based on input time series, including water level: water level change curve in the next 24 hours. Flow rate: flow rate change trend in different basins. Seepage pressure: pressure anomaly detection of potential leakage areas.

[0123] The present invention proposes a large language model reasoning architecture based on tensor parallel processing. By dividing the parameters and input data of the large language model into tensors and distributing them to multiple computing nodes for parallel processing, the reasoning efficiency of the large-scale model is significantly improved. This tensor parallel architecture effectively solves the problems of high computational complexity and slow response speed in the existing single-node reasoning architecture, and is particularly suitable for large-scale water conservancy data and dam safety risk prediction scenarios. The use of the present invention can greatly improve the computational efficiency and real-time performance of the large water conservancy language model. By parallelizing the reasoning tasks of the large-scale model, each computing node can work simultaneously, significantly reducing the computational burden of a single node.

[0124] The present invention is applied to water conservancy, and can not only process larger-scale water conservancy data, but also ensure that the system maintains efficient and accurate risk prediction capabilities while responding quickly. It is particularly suitable for dam safety management applications that require real-time monitoring and high response speed. The system can cope with the challenges of large-scale water conservancy data processing, realize real-time analysis and prediction of dam monitoring data, and provide efficient and accurate risk assessment; multi-node parallel processing ensures the efficiency and stability of the system, avoids single-point computing bottlenecks, and makes the entire system more flexible and economical. The present invention enables the digital twin system of water conservancy projects to better adapt to complex water conservancy environments and provides strong technical support for dam safety management.

[0125] The above is only one embodiment of the present invention, and its description is relatively specific and detailed, but it cannot be understood as limiting the scope of the present invention. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the attached claims.

Claims

1. A digital twin water conservancy large language model reasoning method based on tensor parallel processing, characterized in that: The specific steps include: S1. Gradient calculation optimization: collect water conservancy data in real time as the input of the digital twin water conservancy language model, optimize the gradient calculation process through the SUMMA algorithm, and distribute the processing matrix so that each computing node can independently complete its own computing task. The optimized gradient information obtained can realize parallel processing of water conservancy data collected in real time; S2. Dimensionality reduction and partitioning of self-attention tensor; The gradient information optimized in step S1 is used as the basis for the self-attention mechanism, and the self-attention mechanism is implemented using the Transformer model. Long sequence data is processed in parallel by row and column partitioning to achieve the purpose of dimensionality reduction and partitioning of the gradient information optimized in step S1; S3. Normalize and parallelize the low-dimensional layer: normalize and calculate the low-dimensional layer of the dimension reduction block in step S2 using vector parallelization; S4. Perform inter-layer overall weak extension resource scheduling for the low-dimensional layer after the layer normalization processing in step S3; After the layer normalization calculation is stable in step S3, the calculation load is distributed to different computing nodes, so that each computing node is responsible for the calculation of a specific layer and generates the inference intermediate tensor result corresponding to the specific layer; S5. The decoder is used based on the fully connected layer and the Softmax layer to convert the intermediate tensor results of the model into specific prediction values.

2. According to claim 1, a digital twin water conservancy large language model reasoning method based on tensor parallel processing is characterized by: The water conservancy data in step S1 are collected in real time through a variety of sensors arranged in water conservancy facilities, and specifically include water level, flow velocity, flow, seepage pressure, and surface movement data.

3. A digital twin water conservancy large language model reasoning method based on tensor parallel processing according to claim 1 or 2, characterized in that: In the step S1, the gradient calculation process is optimized by the SUMMA algorithm. A two-dimensional matrix multiplication scheme of an extensible general matrix multiplication algorithm is adopted. In the context of deep learning, the system multiplies the matrix output by the objective function C=A×B, and its differential, i.e., the gradient, is calculated using the chain method to form multiple sub-matrices. These sub-matrices are then assigned to different computing nodes. Each node independently calculates the gradient of the sub-matrix it is responsible for. The calculated gradient will serve as an important basis for subsequent model updates. The specific calculation process is as follows: Let C' be the gradient of the objective function with respect to C, A' and B' be the gradients with respect to the input matrices A and B respectively, then: Where C′ is the gradient of the objective function with respect to C; A′ and B′ are the gradients with respect to the input matrices A and B respectively, and θ represents the partial derivative.

4. A digital twin water conservancy large language model reasoning method based on tensor parallel processing according to claim 1 or 2, characterized in that: In the S2 step, the dimensionality reduction is first performed through linear projection to map the input hidden vector into query (Q), key (K) and value (V), and the attention weight is calculated through matrix multiplication and multiplied with the value V to obtain the output of the self-attention module; the original hidden vector X is first projected into query Q, key K and value V through three independent linear layers: Q=XW Q +b Q K=XW K +b K V=XW V +b V Among them, W Q , W K and W V is the learnable weight matrix, b Q 、b K and b V is bias; After using the linear transformation, the query and the key are dot-producted to get the attention score, and then the attention weight A is obtained by scaling and softmax transformation: Among them, A is the calculated attention weight matrix; Q is the query matrix, which represents the input query information; K is the key matrix, which represents the input key information; K T is the transpose of the bond matrix, which serves as the basis for calculating the similarity measure; d k is the dimension of the key, and the scaling operation helps avoid extreme values ​​in the softmax; The softmax function converts the input into a probability distribution so that its sum is 1; the attention weight A obtained in this way is multiplied by the value V, and the weighted sum is calculated to obtain the output of the self-attention module: Output=AV Finally, the resulting output is resized back to the hidden size and processed through another linear layer to produce the final output of the self-attention module: FinalOutput=Output W O +b O Among them, W O is the weight matrix of the second linear layer; b O is its bias.

5. A digital twin water conservancy large language model reasoning method based on tensor parallel processing according to claim 1 or 2, characterized in that: After the self-attention module is calculated, the S3 step normalizes the redundant low-dimensional layer of the S2 step. Each processor independently calculates the mean and variance of the input data and performs normalization. The calculation of the layer normalization can be expressed as: in, represents the normalized output; Y represents the input data tensor; μ represents the mean E[Y]; σ 2 Represents the variance Var[Y]; δ represents a small constant to prevent division by zero errors; it is usually used to prevent division by zero errors and maintain numerical stability; The mean μ and variance σ 2 The calculation formula is: s 2 =E[Y 2 ]-m 2 Where m is the number of samples; E[Y 2 ] is the expected value of the square of the input data, which represents the average value of the squares of all elements in the input data vector Y; Y j is the jth element in the input sample.

6. According to claim 5, a digital twin water conservancy large language model reasoning method based on tensor parallel processing is characterized by: For the calculation of mean and variance, the processor aggregates the calculation results and uses the all-reduce function to summarize the data of all computing nodes to obtain accurate global mean and variance information. For the gradient calculation of Y, the function description is as follows: in, represents partial derivative.

7. A digital twin water conservancy large language model reasoning method based on tensor parallel processing according to claim 1 or 2, characterized in that: In the step S4, the computing load is distributed to different computing nodes so that each node is responsible for computing a specific layer, as follows: Assuming the total number of layers is L, if it is assigned to N GPUs, each GPU will be responsible for L layers. i It can be expressed as: When processing time series data, the input of multiple time steps is divided into small batches and assigned to different GPUs for independent calculations. If the number of time steps is set to T, then each GPU processes time steps T. j It can be expressed as:

8. A digital twin water conservancy large language model reasoning method based on tensor parallel processing according to claim 1 or 2, characterized in that: In the S5 step, a high-speed communication mechanism is used to integrate the intermediate tensor results of all computing nodes to form a complete global tensor, ensuring that the intermediate results of all nodes are synchronized and consistent; the global tensor contains complete water conservancy feature data; the fully connected layer maps the input global tensor to the output space through a linear transformation, extracts basic information related to the safety value prediction, and outputs the feature vector related to the safety value prediction. The Softmax layer probabilistically processes the feature vector output by the fully connected layer to generate a probability distribution for each safety level; based on the probability distribution generated by the Softmax layer, the safety level is divided in combination with the following classification rules, and finally a risk level prediction result is generated; wherein, the safety level classification rules for the probability P of the risk level are as follows: P>0.6, predicted as low risk; 0.3≤P≤0.6, predicted as medium risk; P<0.3, predicted as high risk.

Citation Information

Patent Citations

  • Edge end collaborative Transform reasoning method based on hybrid model parallelism

    CN117436530A

  • Image generation acceleration method based on consumer graphics card

    CN117788645A

  • Intelligent scheduling system for digital twin pump station group

    CN119047757A

  • Optimizing algorithms for hardware devices

    US20240127045A1

Cited By

  • High-dimensional tensor decomposition-oriented parallel computing system and implementation method thereof

    CN120255963A

  • Large language model distributed reasoning method based on intra-layer parallelism and communication quantization

    CN121562780A

  • Distributed inference method for large language model based on intra-layer parallelism and communication quantization

    CN121562780B

  • Tensor parallel reasoning method supporting large language model and artificial intelligence acceleration device

    CN122175014A