Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

6 results about "Data parallelism" patented technology

Data parallelism is parallelization across multiple processors in parallel computing environments. It focuses on distributing the data across different nodes, which operate on the data in parallel. It can be applied on regular data structures like arrays and matrices by working on each element in parallel. It contrasts to task parallelism as another form of parallelism.

Cross-data center large model training system architecture and resource allocation method and system

PendingCN121957895AImplement collaborative trainingEfficient collaborative utilizationResource allocationBiological modelsWide areaData center
The invention provides a system architecture for cross-data center large model training and a resource allocation method and system, and belongs to the technical field of cross-wide area distributed large model training. According to the method, large-scale model cooperative training across multiple data centers can be realized, and the bottleneck that the computing power of a single data center is limited is broken through. Through unified modeling and scheduling of calculation, memory and network resources, task loads can be intelligently allocated according to hardware performance and network bandwidth of different data centers, and efficient collaborative utilization of computing power resources is realized. The training task of the super-large-scale model can be rapidly completed in the heterogeneous computing power environment, and the training time is remarkably shortened. The provided flexible parallelism degree allocation method can be adaptive to different task and resource conditions, the proportion of data parallelism, model parallelism and pipeline parallelism is automatically adjusted, the parallelism efficiency is improved, and the communication overhead is reduced. A training time estimation function is integrated, the overall time delay and resource requirements can be predicted before task execution, and a basis is provided for scheduling decision making.
Owner:BEIJING JIAOTONG UNIV

A high-performance multi-party secure computing training method and system based on GPU

The application relates to a GPU-based high-performance multi-party secure computation training method and system, and particularly relates to the field of multi-party secure computation protocols.The application aims to provide a multi-party secure computation training framework with higher parallelism, so as to realize parallelism between different layers of a neural network in a manner of combining data parallelism and model parallelism, and improve the data throughput speed of a training process.The method is a multi-party secure computation training system based on a pipeline flow training method, as shown in Figure 1, the method is designed according to the characteristics that the bottlenecks of linear computation network layers and nonlinear computation network layers in the MPC model training process are calculation and communication respectively, a pipeline flow training method is designed, parallelism between sub-networks is realized, and an optimal sub-network segmentation algorithm is realized to balance the training load between each sub-network.The application provides a multi-party secure computation training framework with higher parallelism, parallelism between different layers of a neural network is realized in a manner of combining data parallelism and model parallelism, and the data throughput speed of a training process is greatly improved.
Owner:HARBIN INST OF TECH

Methods and devices for umbilical blood flow ultrasound image measurement and parallel processing

This application relates to a method and apparatus for measuring and parallel processing umbilical blood flow ultrasound images. The method includes: accelerating the acquisition of a standard cross-section using data parallelism and convolution acceleration; identifying the region of interest (ROI), X-axis region, and umbilical blood flow spectral envelope of the standard cross-section using digital image processing technology, thereby calculating the peaks and troughs of the umbilical blood flow spectrum and locating a continuous and stable umbilical blood flow spectrum. Further identification of scale points allows for the rapid location of velocity and time scales, enabling efficient calculation of the conversion ratios of the Y-axis velocity scale and the X-axis time scale. Based on this, the calculated continuous and stable umbilical blood flow spectrum, the conversion ratio of the Y-axis velocity scale, and the conversion ratio of the X-axis time scale are used to automatically and accurately measure the correlation coefficient of the umbilical blood flow spectrum. This process requires no manual intervention, significantly improving the calculation efficiency of the umbilical blood flow spectrum correlation coefficient.
Owner:HUNAN UNIV

Distributed processing architecture

Embodiments of the present disclosure include techniques for processing neural networks. Various forms of parallelism can be achieved using a topology of combined processor sequences. In one embodiment, the present disclosure includes a computer system comprising a plurality of processor groups, each processor group comprising a plurality of processors. A plurality of network switches are coupled to a subset of the plurality of processor groups. A subset of the processors in a processor group can be configured to form a sequence, and the network switches can be configured to form at least one sequence across one or more of the plurality of processor groups to perform neural network computations. Various configurations for creating Hamiltonian cycles are disclosed to support data parallelism, pipeline parallelism, layer parallelism, or a combination thereof.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

A multi-cluster hybrid parallel training strategy hierarchical optimization method and system

This invention relates to the field of artificial intelligence technology and discloses a hierarchical optimization method and system for multi-cluster hybrid parallel training strategies. The method includes the following steps: S1: Acquire multi-cluster environment information and AI training task information, and construct a multi-dimensional feature space containing static and dynamic features; S2: Based on the multi-dimensional feature space, determine the applicable execution scale of the task through an AI task training hierarchy partitioning model, wherein the execution scale includes multi-cluster training, single-cluster training, or single-machine training; S3: For tasks determined to be multi-cluster training, generate a set of hybrid parallel training candidate schemes containing data parallelism, tensor parallelism, and pipeline parallelism; S4: Based on a multi-dimensional cost model, select the hybrid parallel scheme with the lowest training cost from the set of candidate schemes for execution; S5: Monitor the training process in real time, and dynamically execute cross-cluster load migration and resource reallocation when an abnormal risk is detected.
Owner:GUANGZHOU RADIO & TELEVISION RES INST CO LTD +2

Large-scale language model automatic parallel training method and system based on parallel monte carlo algorithm

The application discloses a large-scale language model automatic parallel training method and system based on a parallel Monte Carlo algorithm, and the method steps comprise the following steps: in step S01, a large-scale language model and a device topology structure of a computing cluster are acquired; in step S02, the acquired large-scale language model is converted into an operator sequence and a corresponding calculation graph, and a model representation of the model is generated; in step S03, a unified parallel strategy search space containing at least three dimensions of data parallelism, model parallelism and pipeline parallelism is constructed according to the operator sequence, the calculation graph, the model representation and the device topology structure; in step S04, an optimal parallel training strategy is searched in the unified parallel strategy search space by using a parallel Monte Carlo tree search algorithm; and in step S05, the large-scale language model is trained according to the optimal parallel training strategy. The application can realize automatic parallel training of the large-scale language model, reduce the time and resources required for training, and improve the training efficiency.
Owner:CHINESE PEOPLES LIBERATION ARMY INFORMATION SUPPORT CORPS ENGINEERING UNIVERSITY