Artificial intelligence reasoning acceleration method based on special data processor
By configuring hardware resources on the data processor and adopting parallel data network protocol processing and dynamic resource scheduling technology, the problems of low efficiency and resource waste in massive heterogeneous data processing are solved, and efficient network protocol processing and dynamic expansion capabilities are achieved.
Patent Information
- Application Number
- CN202411351106.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-06-03
AI Technical Summary
When processing massive heterogeneous data, the network protocol processing efficiency is low, unable to match the high-performance computing speed, and lacks dynamic scalability and compatibility, resulting in waste of computing resources and reduced inference speed.
An artificial intelligence inference acceleration method based on a dedicated data processor is designed. By configuring the hardware resources of the data processor, parallel data network protocol processing of network virtual technology is realized, and a circular dynamic polling algorithm and inference resource intelligent scheduling system are adopted to dynamically allocate hardware resources and expand computing architecture to adapt to the sudden changes in computing resource requirements.
It improves the network protocol processing efficiency for massive heterogeneous data, realizes dynamic expansion of computing power, adapts to the sudden large data set selection needs, reduces the burden on the CPU, and realizes efficient data reception and distribution inference.
Smart Images

Figure CN120087471A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processor acceleration, and more specifically to an artificial intelligence reasoning acceleration method based on a dedicated data processor. Background Art
[0002] With the rapid development of big data and artificial intelligence (AI) technology, the amount of data generated by various applications is growing exponentially, and the data formats are becoming increasingly diverse. Traditionally, data reception and analysis tasks mainly rely on general-purpose central processing units (CPUs) for processing, including network protocol processing, computing acceleration, etc. However, as the performance growth rate of the CPU and the growth of data volume have shown a significant scissors gap phenomenon, facing massive and diverse data, the CPU often shows obvious resource bottlenecks when performing data analysis. In particular, when there is a sudden change in the amount of data being processed, the existing processing system needs to prepare enough computing resources in advance, which leads to a waste of resources.
[0003] To this end, many manufacturers have begun to study dedicated parallel architectures, including graphics processing unit (GPU) acceleration, FPGA acceleration, neural network processor (NPU) acceleration, etc. Although these AI reasoning accelerations have made significant progress, there are still several problems that need to be solved:
[0004] Lack of efficient network protocol processing technology for massive heterogeneous data: In the face of massive heterogeneous reasoning data input, existing solutions focus on the computational reasoning of the model, but are inefficient in parsing and processing data in different formats, and cannot match the high-performance computing speed, resulting in a decrease in the overall reasoning speed. How to design a corresponding data processor architecture to efficiently process diverse data formats is an urgent problem to be solved.
[0005] Lack of compatibility and dynamic scalability support for heterogeneous architectures: On the one hand, the current amount of inference data has increased dramatically, and the amount of processed data has suddenly changed. Without the support of a dynamically scalable architecture, it will cause a huge waste of subsequent computing resources. On the other hand, with the continuous development of AI technology, new deep learning models and algorithms are constantly emerging, and the types of models provided are diverse. An intelligent dynamic scheduling service is also needed to realize the scheduling of computing resources, thereby achieving efficient reasoning of various model algorithms.
[0006] To this end, the present invention proposes an artificial intelligence inference acceleration method and system based on a dedicated data processor that can adapt to sudden changes in computing resource requirements and meet the computing requirements of specific tasks. The system improves the network protocol processing efficiency for massive heterogeneous data. At the same time, it realizes a parallel computing architecture with dynamically scalable computing power and adaptable to sudden large dataset selection requirements, reducing the burden on the CPU. Through this technical means, efficient data reception and distribution inference are achieved, thereby realizing artificial intelligence inference acceleration based on the data processor. Summary of the Invention
[0007] To overcome the above defects of the prior art, the present invention provides an artificial intelligence inference acceleration method based on a dedicated data processor.
[0008] The technical solution of the present invention is as follows:
[0009] An artificial intelligence inference acceleration method based on a dedicated data processor includes the following steps:
[0010] S1: Configure the hardware resources of the data processor;
[0011] S2: Parallel data network protocol processing based on network virtual technology
[0012] Obtain a service request input by an external object into the data processor, make a first confirmation of the type of the service request, verify whether the type of the service request belongs to the working range of the data processor, and selectively respond to or ignore the service request; a network virtual machine Hypervisor will be designed on the data processor to achieve efficient reception and processing of network transmission data; design a network traffic monitoring module to monitor massive heterogeneous data to be processed in real time, and design a circular dynamic polling algorithm. Use a circular data receiver to receive data in sequence. When the circular receiver is completely occupied, the receiver will be expanded, and the expansion will be carried out on a scale of 2N, where N is the size of the currently awakened data receiver, so as to achieve efficient processing and reception of network data; at the same time, hypervisor will also monitor the usage status of the circular receiver. When the usage rate of the receiver in M rounds of data reception is N / 2, the size of the circular receiver will be dynamically reduced to N / 2, and data reception relocation will be performed before the reduction, moving all reception points to between [0, N / 2);
[0013] S3: Dynamic scheduling of inference resources based on data IO and resource matching
[0014] After the data processor confirms that the received service request is within its scope of work and responds to the service request, it further breaks down the content of the service request for a second confirmation. It checks whether the content of the service request includes an inference link of machine learning. According to the requirements and scale of the inference link of machine learning, it selectively allocates and enables hardware resources and performs corresponding processing. When it is confirmed for the second time that the service request includes an inference link, dynamic scheduling of inference resources will be carried out to meet the requirements of parallel processing of data requests while efficiently utilizing computing resources. For this purpose, we have developed an intelligent scheduling system for inference resources. The scheduling center will monitor the current inference resources and the status of the loaded algorithms in real time, and intelligently calculate the number of resources required by the inference module, and use kubernetes to achieve dynamic expansion and scheduling of resources. The scheduling principle is as follows:
[0015] N d *V d ≤N i *V i
[0016] Among them, N d is the number of data receivers, V d is the average speed of data received by each data receiver, N i is the number of inference modules, V i is the inference speed of the inference module;
[0017] S4: The data processor returns the result of the service request after response to the external object and resets the redundant hardware resources.
[0018] The steps of the ring dynamic polling algorithm are as follows:
[0019] (1) Initialization:
[0020] Set the task set T = {T 1 , T 2 , …, T 1}, where n is the number of initial tasks; Initialize the current pointer p to point to the first task, that is, p = 1; Set a capacity threshold C, which represents the maximum allowable number of tasks in the ring structure;
[0021] (2) Polling process:
[0022] Start processing from the task pointed to by the current pointer p, that is, process task T p ; After processing the current task, move the pointer p to the next task:
[0023] p = (p mod n) + 1
[0024] If a certain task i in the task set has been completed or is unavailable, skip this task;
[0025] (3) Dynamic adjustment:
[0026] Dynamic expansion:
[0027] When the number of the task set T reaches the current capacity threshold C, double-expand the capacity of the ring structure:
[0028] C = 2 * C
[0029] Add the new task to the expanded set T and continue polling.
[0030] Dynamic reduction:
[0031] If the number of tasks in the ring structure remains low for a period of time (e.g., t time units) and the current number of tasks is less than half of the capacity threshold C, halve the capacity of the ring structure:
[0032]
[0033] Before reducing the capacity, the existing tasks must be grouped from their original positions into the new ring capacity inside. The specific operations are as follows:
[0034] Reassign all tasks to positions numbered from 1 to to ensure that all tasks are stored within the new capacity range. Adjust the pointer p to point to the first task of the new task set, avoiding pointing to an empty position.
[0035] (4) Capacity adjustment trigger conditions:
[0036] Expansion trigger condition: When the number of tasks |T| satisfies |T| ≥ C.
[0037] Reduction trigger condition: When the number of tasks |T| satisfies and there is no significant increase within t time units.
[0038] (5) Termination condition:
[0039] When the task set T is empty or all tasks have been completed, the algorithm terminates.
[0040] Configuring the hardware resources of the data processor described in step S2 means that the data processor includes a first data processing module, a standard communication interface, an on-chip shared cache, a local memory, an input / output channel, at least one second data processing module, and a network interface; the first data processing module, the input / output channel, and at least one second data processing module are all communicatively connected to the standard communication interface; the on-chip shared cache, the local memory, and the network interface are respectively communicatively connected to the first data processing module and at least one second data processing module; the on-chip shared cache is used to receive service requests input by an external object through the standard communication interface and the input / output channel, and perform a first confirmation and a second confirmation on the service requests; the first data processing module responds to the service request after the first confirmation, and at least one second data processing module selectively responds to the service request after the second confirmation; both the first data processing module and at least one second data processing module are provided with on-chip caches.
[0041] The on-chip shared cache is used to receive service requests input by an external object through the standard communication interface and the input / output channel, and perform a first confirmation and a second confirmation on the service requests. When performing the first confirmation, it verifies whether the type of the service request belongs to any one of a network communication request, a data storage request, a security request, or a data interaction analysis request. If it belongs, the service request is transferred to the first data processing module for processing. If it does not belong, the data processor ignores the service request; when the type of the service request is a data interaction analysis request and includes an inference link of machine learning, the on-chip shared cache combines with the first data processing module to further perform a second confirmation on the service request.
[0042] The on-chip shared cache combines with the first data processing module to further perform a second confirmation on the service request, which is to judge whether the first data processing module can meet the requirements and scale of the inference link of machine learning. Specifically, the following conditions need to be met simultaneously: A) The first data processing module is in a completely idle state or a partially idle state; B) The processing efficiency of the first data processing module cannot meet the real-time requirement of the inference link of machine learning for the processing result; C) The data peak value of the data operation exceeds the capacity of the on-chip cache of the first data processing module.
[0043] The real-time requirement of the inference link of machine learning for the processing result is that the response time of the first data processing module to the fixed-length data or variable-length data input by the service request does not exceed 30 seconds.
[0044] The at least one second data processing module selectively responds to the service request after the second confirmation. It comprehensively considers the sample size, the number and dimension of prediction features, and the proportion of the parameters updated in each iteration to the total parameters, and determines the number K of the second data processing modules that need to be put into parallel, which is to make Where α represents the sample size participating in the inference process of machine learning; β1 represents the number of predicted features output by the inference process of machine learning; β2 represents the dimension of the predicted features output by the inference process of machine learning; ρ represents the average value of the proportion of updated parameters in each iteration to all parameters; f 1 、f 2 、f 3 and f 4 are all scaling factors, and the value range is a real number within (0, 1]; the parameter K obtained by summing each term on the right side of the formula is rounded up to obtain the number of second data processing modules that need to be put into parallel.
[0045] Preferably, the first data processing module is a data processing unit DPU.
[0046] More preferably, the at least one second data processing module is a tensor processing unit TPU, and the on-chip cache capacity of the at least one second data processing module is greater than the on-chip cache capacity of the first data processing module.
[0047] Based on the above technical solutions, preferably, the first data processing module is one of an ARM chip, an ASCI chip, an FPGA chip, or a RISC-V chip.
[0048] On the other hand, the present invention also provides an artificial intelligence inference acceleration system based on a data processor, including a data processor and a storage unit. The storage unit stores a program that can be run by the data processor. When the program is executed, the following steps are performed:
[0049] Obtain a service request input by an external object to the data processor, perform a first confirmation on the type of the service request, verify whether the type of the service request belongs to the working range of the data processor, and selectively respond to or ignore the service request;
[0050] After the data processor confirms that the received service request belongs to the working range and responds to the service request, further subdivide the content of the service request, perform a second confirmation, confirm whether the content of the service request contains the inference process of machine learning, and selectively allocate hardware resources and execute the content of the service request according to the requirements and scale of the inference process of machine learning;
[0051] The data processor returns the result of the service request after response to the external object and resets the redundant hardware resources.
[0052] An artificial intelligence inference acceleration method and system provided by the present invention have the following beneficial effects compared with the prior art:
[0053] (1) The present invention proposes an artificial intelligence inference acceleration method and system based on a dedicated data processor that can adapt to mutations in computing resource requirements and meet the computing requirements of specific tasks. This system improves the network protocol processing efficiency for massive heterogeneous data. At the same time, it realizes a parallel computing architecture with dynamically scalable computing power that can adapt to sudden large dataset selection requirements, reducing the burden on the CPU. Through this technical means, efficient data reception and distribution inference are achieved, thereby realizing artificial intelligence inference acceleration based on the data processor.
[0054] (2) The data processor of the present application estimates the corresponding computing scale according to service requests related to external data, evaluates whether the current hardware can meet the actual requirements, and activates the integrated heterogeneous hardware resources for simultaneous processing of real-time parallel tasks to improve the computing power bottleneck. The data processor integrates a first data processing module and several second data processing modules, and the expandability of this integrated architecture of the data processor is excellent;
[0055] (3) By estimating the computing power required for service requests related to data, two confirmation links are used to confirm whether to activate the second data processing modules and the corresponding quantity required, so as to meet the variable requirements of real-time computing power. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0057] Figure 1 It is a flowchart of the steps of an artificial intelligence inference acceleration method and system based on a data processor of the present invention;
[0058] Figure 2 It is an internal structure block diagram of a data processor of an artificial intelligence inference acceleration method and system based on a data processor of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in combination with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0060] Embodiment 1
[0061] As Figure 1 AndFigure 2 As shown in the figure, the present invention provides an artificial intelligence inference acceleration method based on a data processor, specifically including the following steps:
[0062] S1: Configure the hardware resources of the data processor;
[0063] The hardware resources of the data processor in this application include a first data processing module, a standard communication interface, an on-chip shared cache, a local memory, an input / output channel, at least one second data processing module, and a network interface; the first data processing module, the input / output channel, and at least one second data processing module are all communicatively connected to the standard communication interface; the on-chip shared cache, the local memory, and the network interface are respectively communicatively connected to the first data processing module and at least one second data processing module. The on-chip shared cache is used to receive and process service requests input by external objects into the data processor. In this application, the first data processing module is a data processing unit DPU, and it can be one of an ARM chip, an ASCI chip, an FPGA chip, or a RISC-V chip.
[0064] At least one second data processing module is a tensor processing unit TPU. The first data processing module and at least one second data processing module are both provided with on-chip caches. The standard communication interface can be a PCIE interface or an AXI interface. The on-chip shared cache can be a DDR RAM. The local memory can use a solid-state storage medium. The network interface uses a high-speed Ethernet interface.
[0065] S2: Obtain the service request input by the external object into the data processor, perform a first confirmation on the type of the service request, verify whether the type of the service request belongs to the working range of the data processor, and selectively respond to or ignore the service request. A network virtual machine Hypervisor will be designed on the data processor to achieve efficient reception and processing of network transmission data; design a network traffic monitoring module to monitor the massive heterogeneous data to be processed in real time, and design a circular dynamic polling algorithm. Use a circular data receiver to receive data in sequence. When the circular receiver is completely occupied, the receiver will be expanded, and the expansion will be carried out on a scale of 2N, where N is the size of the currently awakened data receiver, so as to achieve efficient processing and reception of network data; at the same time, hypervisor will also monitor the usage status of the circular receiver. When the usage rate of the receiver in M rounds of data reception is N / 2, the size of the circular receiver will be dynamically reduced to N / 2, and data reception repositioning will be performed before the reduction, and all receiving points will be moved to between [0, N / 2);
[0066] S3: After the data processor confirms that the received service request is within its scope of work and responds to the service request, it further breaks down the content of the service request for a second confirmation. It checks whether the content of the service request contains an inference link in machine learning. According to the requirements and scale of the inference link in machine learning, it selectively allocates and enables hardware resources and performs corresponding processing. When it is secondarily confirmed that the service request contains an inference link, dynamic scheduling of inference resources will be carried out to meet the requirements of parallel processing of data requests while efficiently utilizing computing resources. For this purpose, we have developed an intelligent inference resource scheduling system. The scheduling center will monitor the current inference resources and the status of loaded algorithms in real time, and intelligently calculate the number of resources required by the inference module, and use kubernetes to achieve dynamic expansion and scheduling of resources. The scheduling principle is as follows:
[0067] N d *V d ≤N i *V i
[0068] Among them, N d is the number of data receivers, V d is the average speed of data received by each data receiver, N i is the number of inference modules, V i is the inference speed of the inference module;
[0069] The specific content of the first and second confirmations made by the data processor for the service request input by an external object is: The first data processing module responds to the service request after the first confirmation, and at least one second data processing module selectively responds to the service request after the second confirmation. Specifically, according to the above content, the on-chip shared cache is used to receive the service request input by the external object through the standard communication interface and the input and output channels, and perform the first and second confirmations on the service request.
[0070] Specifically, when the service request is first confirmed, it is verified whether the type of the service request belongs to any one of network communication requests, data storage requests, security requests, or data interaction analysis requests. If it belongs, the service request is transferred to the first data processing module for processing. If not, the data processor ignores the service request. When the type of the service request is a data interaction analysis request and includes an inference link of machine learning, the on-chip shared cache combines with the first data processing module to further perform a second confirmation on the service request. The first confirmation indicates that the service request is within the exclusive processing scope of the data processor. The second confirmation is that the current first data processing module may not be able to meet the capabilities of long-term, high-concurrency large data processing, and it is necessary to further enable a dedicated and extended second data processing module for data processing. The second data processing module focuses on data interaction analysis requests and includes the content of machine learning inference and is good at iterative operations of machine learning. In this embodiment, it is set that the on-chip cache capacity of at least one second data processing module is greater than that of the first data processing module.
[0071] Among them, the on-chip shared cache combines with the first data processing module to further perform a second confirmation on the service request, which is to judge whether the first data processing module can meet the requirements and scale of the inference link of machine learning. Specifically, the following conditions need to be met simultaneously: A) The first data processing module is in a completely idle state or a partially idle state; B) The processing efficiency of the first data processing module cannot meet the real-time requirement of the processing result for the inference link of machine learning; C) The data peak value of the data operation exceeds the capacity of the on-chip cache of the first data processing module.
[0072] In condition B), the real-time requirement of the processing result for the inference link of machine learning is that the response time of the first data processing module to the fixed-length data or variable-length data input by the service request does not exceed 30 seconds. That is, the time interval from receiving the first frame of data from the standard communication interface to completing the data processing and sending the last frame of the processing result through the standard communication interface does not exceed 30 seconds. The second confirmation determines that the type of the service request is a data interaction analysis request and includes an inference link of machine learning. If the on-chip shared cache receives the relevant data of the service request to be processed and the first data processing module evaluates that the data processing process may not be able to meet the high-speed real-time requirement, a certain number of second data processing modules will be enabled for corresponding processing to respond to the service request. When the capacity of the on-chip cache of the first data processing module in condition C) is insufficient, it is necessary to rely on the space of the on-chip shared cache for common storage and output. However, this process uses the on-chip shared cache outside the first data processing module, resulting in a large amount of communication and read / write time, which will further affect the real-time performance.
[0073] The number of the second data processing modules to be invested is determined by comprehensively considering the sample size of the inference session of machine learning in the data interaction analysis request, the number and dimensions of the prediction features, and the proportion of the parameters updated in each iteration to all the parameters: Let where α represents the sample size of the inference session of machine learning; β1 represents the number of the prediction features output by the inference session of machine learning; β2 represents the dimensions of the prediction features output by the inference session of machine learning; ρ represents the average value of the proportion of the parameters updated in each iteration to all the parameters; f1, f2, f3, and f4 are all scaling factors, and the value range is real numbers within (0, 1]; the parameter K obtained by summing up each item on the right side of the formula is rounded up, which is the number of the second data processing modules to be invested in parallel. Since the second data processing module has a large on-chip cache, multiple second data processing modules can be used for parallel processing.
[0074] When performing the inference session of machine learning in the data interaction analysis request, an input database is constructed in the on-chip shared cache in advance, and the content in the database is divided into several adjacent regions with the same elements. Each region in the input database is processed by a second data processing module according to the preset processing rules; after completing the inference session of machine learning in the data interaction analysis request, the results output by each second data processing module are combined to construct an output database, and the output database contains information related to the number and dimensions of the prediction features of the inference session of machine learning. On the one hand, the content of the output database is output to the requester through the standard communication interface, and can also be stored on-chip through the local memory.
[0075] S4: The data processor returns the result of the service request after response to the external object, and resets the redundant hardware resources.
[0076] After completing the inference session of machine learning in the data interaction analysis request, when there is no situation that requires the second data processing module to process again within a certain period of time, the hardware resources of the second data processing module will be reset.
[0077] On the other hand, the present invention also provides an artificial intelligence inference acceleration system based on a data processor, including a data processor and a storage unit. The storage unit stores a program that can be run by the data processor, and when the program is executed, the following steps are performed:
[0078] Obtain a service request input by an external object to the data processor, perform a first confirmation on the type of the service request, verify whether the type of the service request belongs to the working scope of the data processor, and selectively respond to or ignore the service request;
[0079] After the data processor confirms that the received service request is within its scope of work and responds to the service request, it further breaks down the content of the service request for a second confirmation. It checks whether the content of the service request includes an inference link of machine learning. According to the requirements and scale of the inference link of machine learning, it selectively allocates hardware resources and executes the content of the service request.
[0080] The data processor returns the result of the service request after response to the external object and resets the redundant hardware resources.
[0081] Embodiment 2
[0082] The steps of the ring-shaped dynamic polling algorithm are as follows:
[0083] (1) Initialization:
[0084] Set the task set T = {T 1 , T 2 , …, T 1}, where n is the number of initial tasks; initialize the current pointer p to point to the first task, i.e., p = 1; set a capacity threshold C, which represents the maximum allowable number of tasks in the ring structure.
[0085] (2) Polling process:
[0086] Start processing from the task pointed to by the current pointer p, i.e., process task T p ; after processing the current task, move the pointer p to the next task:
[0087] p = (p mod n) + 1
[0088] If a certain task i in the task set has been completed or is unavailable, skip this task.
[0089] (3) Dynamic adjustment:
[0090] Dynamic expansion:
[0091] When the number of the task set T reaches the current capacity threshold C, double-expand the capacity of the ring structure:
[0092] C = 2 * C
[0093] Add the new task to the expanded set T and continue polling.
[0094] Dynamic reduction:
[0095] If the number of tasks in the ring structure remains low for a period of time (e.g., t time units) and the current number of tasks is less than half of the capacity threshold C, halve the capacity of the ring structure:
[0096]
[0097] Before reducing the capacity, the existing tasks must be aggregated from their original positions to the new annular capacity as follows:
[0098] Reallocate all tasks to positions numbered from 1 to , ensuring that all tasks are stored within the new capacity range. Adjust the pointer p to point to the first task of the new task set, avoiding pointing to empty positions.
[0099] (4) Capacity adjustment trigger conditions:
[0100] Expansion trigger condition: When the number of tasks |T| satisfies |T| ≥ C.
[0101] Reduction trigger condition: When the number of tasks |T| satisfies and there is no significant increase within t time units.
[0102] (5) Termination condition:
[0103] When the task set T is empty or all tasks are completed, the algorithm terminates. Specifically, the calculation method of the circular dynamic polling algorithm:
[0104]
[0105]
[0106]
[0107]
[0108] Example 3
[0109] The calculation method for the average proportion ρ of the parameters updated in each iteration to all the parameters is as follows:
[0110]
[0111] where: T is the total number of iterations, P l is the total number of parameters in the l-th layer, is the number of parameters updated in the l-th layer in the t-th iteration, and L is the number of layers of the neural network model.
[0112] The specific calculation steps are as follows:
[0113] Calculate the number of parameters P in each layer l
[0114] For the convolutional layer, assuming the convolutional kernel size is filter_height×filter_width, the number of input channels is in_channels, the number of output channels is out_channels, and the total number of parameters is the weighted sum of the weights and biases of the convolutional kernel. The calculation formula is as follows:
[0115] P l = P conv = P weights + P bias
[0116] = filter_height×filter_width×in_channels×out_channels
[0117] + out_channels
[0118] For the fully connected layer, assuming the number of input units of the fully connected layer is input_units, the number of output units is output_units, and the total parameters are the weighted sum of the number of weights and biases. The calculation formula is as follows:
[0119] P l = P fc = P weights + P bias = input_units×outputs_units + output_units
[0120] Calculate the number of updated parameters in the l-th layer in the t-th iteration
[0121] Set the parameter update threshold δ threshold , and obtain the weight update amplitude by calculating the gradient of each parameter of the loss function during training and using the gradient descent method Bias update amplitude where η is a hyperparameter and is set to a fixed value according to the model
[0122] For the convolutional layer, the total number of updated parameters is the weighted sum of the weight update number and the bias update number. The calculation formula for the weight update number is where is the weight update amount at the total positions i, j, k, l of the convolutional kernel, and 1(·) represents the indicator function, which takes the value of 1 when the condition is satisfied and 0 otherwise; the calculation formula for the update number of the bias is where is the update amount of the m-th bias.
[0123] Total number of updated parameters:
[0124]
[0125] For the fully connected layer, the total updated parameters are also the weighted sum of the weight update quantity and the bias update quantity. The formula for calculating the weight update quantity is where is the weight update amount at positions i, j in the fully connected layer. The formula for calculating the bias update amount is where is the update amount of the k-th bias.
[0126] Total number of updated parameters:
[0127]
[0128] Total number of parameters updated per iteration:
[0129]
Claims
1. An artificial intelligence reasoning acceleration method based on a dedicated data processor, comprising the following steps: S1: configure the hardware resources of the data processor; S2: Parallel data network protocol processing based on network virtualization technology Obtain the service request input from the external object to the data processor, make the first confirmation of the type of the service request, verify whether the type of the service request belongs to the working scope of the data processor, and selectively respond to or ignore the service request; design a network virtual machine Hypervisor on the data processor to achieve efficient reception and processing of network transmission data; design a network traffic monitoring module to monitor the massive heterogeneous data to be processed in real time, and design a circular dynamic polling algorithm to use the circular data receiver to receive data in sequence. When the circular receiver is fully occupied, the receiver will be expanded to a scale of 2N, where N is the size of the current awakened data receiver, so as to achieve efficient processing and reception of network data; at the same time, the hypervisor will also monitor the usage status of the circular receiver. When the usage scale of the receiver is N / 2 in the M rounds of data reception, the size of the circular receiver will be dynamically reduced to N / 2, and the data reception will be relocated before the reduction, and all receiving points will be moved to [0, N / 2). S3: Dynamic scheduling of inference resources based on data IO and resource matching After the data processor confirms that the received service request is within the scope of work and responds to the service request, it further subdivides the content of the service request and performs a second confirmation to confirm whether the content of the service request contains the reasoning link of machine learning. According to the needs and scale of the reasoning link of machine learning, it selectively allocates hardware resources and enables and performs corresponding processing; when the second confirmation that the service request contains the reasoning link, the reasoning resources will be dynamically scheduled to meet the requirements of parallel processing of data requests while efficiently utilizing computing resources; for this purpose, we have developed an intelligent scheduling system for reasoning resources. The scheduling center will monitor the current reasoning resources and the status of the loaded algorithm in real time, and intelligently calculate the number of resources required by the reasoning module, and use kubernetes to achieve dynamic expansion and scheduling of resources; the scheduling principle is as follows: N d *V d ≤N i *V i in, N d is the number of data receivers, V d The average speed at which each data receiver receives data, N i is the number of inference modules, V i The inference speed of the inference module; S4: The data processor returns the result of the responded service request to the external object and resets the redundant hardware resources.
2. The artificial intelligence reasoning acceleration method based on a dedicated data processor according to claim 1, characterized in that: The steps of the circular dynamic polling algorithm are as follows: (1) Initialization: Set the task set T = {T1, T2, ..., T1}, where n is the number of initial tasks; initialize the current pointer p to point to the first task, that is, p = 1; Set a capacity threshold C, which represents the maximum number of tasks allowed in the ring structure; (2) Polling process: Start processing from the task pointed to by the current pointer p, that is, process task T p ; After processing the current task, move the pointer p to the next task: p=(pmod n)+1 If a task,i,in the task set is completed or unavailable, then skip the task; (3) Dynamic adjustment: Dynamic expansion: When the number of task sets T reaches the current capacity threshold C, the capacity of the ring structure is expanded exponentially: C=2*C Add the new task to the expanded set T and continue polling. Dynamic reduction: If the number of tasks in the ring structure remains low for a period of time (e.g., t time units), and the current number of tasks is less than half of the capacity threshold C, the capacity of the ring structure is reduced exponentially: Before reducing capacity, existing tasks must be aggregated from their original locations to the new ring capacity. The specific operations are as follows: Reassign all tasks to the same group numbered from 1 to The position of , ensuring that all tasks are stored within the new capacity range. Adjust the pointer p to point to the first task of the new task set to avoid pointing to an empty position. (4) Capacity adjustment trigger conditions: Extension trigger condition: when the number of tasks |T| satisfies |T|≥C. Reduction trigger condition: When the number of tasks |T| satisfies And there is no significant increase within t time unit. (5) Termination conditions: The algorithm terminates when the task set T is empty or all tasks have been completed.
3. The artificial intelligence reasoning acceleration method based on a dedicated data processor according to claim 1, characterized in that: The hardware resources of the data processor configured in step S2 are to make the data processor include a first data processing module, a standard communication interface, an on-chip shared cache, a local memory, an input / output channel, at least one second data processing module and a network interface; the first data processing module, the input / output channel and the at least one second data processing module are all communicatively connected to the standard communication interface; the on-chip shared cache, the local memory and the network interface are communicatively connected to the first data processing module and the at least one second data processing module respectively; The on-chip shared cache is used to receive service requests input by external objects through a standard communication interface and input / output channels, and to perform a first confirmation and a second confirmation on the service requests; the first data processing module responds to the service request after the first confirmation, and at least one second data processing module selectively responds to the service request after the second confirmation; the first data processing module and at least one second data processing module are both provided with an on-chip cache.
4. The artificial intelligence reasoning acceleration method based on a dedicated data processor according to claim 3 is characterized in that: The on-chip shared cache is used to receive service requests input by external objects through a standard communication interface and input and output channels, and to perform a first confirmation and a second confirmation on the service requests. During the first confirmation, it is verified whether the type of the service request belongs to any of the network communication request, data storage request, security request or data interaction analysis request. If so, the service request is transferred to the first data processing module for processing; if not, the data processor ignores the service request; when the type of the service request is a data interaction analysis request and includes the reasoning link of machine learning, the on-chip shared cache combines with the first data processing module to further perform a second confirmation on the service request.
5. The artificial intelligence reasoning acceleration method based on a dedicated data processor according to claim 3 is characterized in that: The on-chip shared cache combines with the first data processing module to further perform a second confirmation of the service request to determine whether the first data processing module can meet the needs and scale of the reasoning link of machine learning. Specifically, the following conditions must be met at the same time: A) the first data processing module is in a completely idle state or a partially idle state; B) the processing efficiency of the first data processing module cannot meet the real-time requirements of the processing results of the reasoning link of machine learning; C) the data peak value of the data operation exceeds the capacity of the on-chip cache of the first data processing module.
6. The artificial intelligence reasoning acceleration method based on a dedicated data processor according to claim 1, characterized in that: The real-time requirement of the processing results in the inference phase of the machine learning is that the response time of the first data processing module to the fixed-length data or variable-length data input by the service request does not exceed 30 seconds.
7. The artificial intelligence reasoning acceleration method based on a dedicated data processor according to claim 3 is characterized in that: The at least one second data processing module selectively responds to the service request after the second confirmation by comprehensively considering the sample size, the number and dimension of the prediction features, and the proportion of the parameters updated in each iteration to all the parameters, and determines the number K of the second data processing modules that need to be invested in parallel, which is Where α represents the sample size involved in the reasoning phase of machine learning; β1 represents the number of predictive features output by the reasoning phase of machine learning; β2 represents the dimension of the predictive features output by the reasoning phase of machine learning; ρ represents the average value of the proportion of updated parameters to all parameters in each iteration; f1, f2, f3 and f4 are all scaling factors, and their value range is real numbers in (0, 1]; the parameter K obtained by summing up the terms on the right side of the formula is rounded up, which is the number of second data processing modules that need to be invested in parallel.
8. The artificial intelligence reasoning acceleration method based on a dedicated data processor according to claim 1, characterized in that: This method is run in an artificial intelligence reasoning acceleration system based on a data processor. The system includes a data processor and a storage unit. The storage unit stores a program that can be run by the data processor. When the program is executed, the following steps are performed: Obtain the service request input into the data processor by the external object, make the first confirmation of the type of the service request, verify whether the type of the service request falls within the scope of work of the data processor, and selectively respond to or ignore the service request; After the data processor confirms that the received service request is within the scope of work and responds to the service request, it further subdivides the content of the service request and performs a second confirmation to confirm whether the content of the service request contains the reasoning link of machine learning. According to the needs and scale of the reasoning link of machine learning, it selectively allocates hardware resources and executes the content of the service request; The data processor returns the result of the responded service request to the external object and resets the redundant hardware resources.
9. The artificial intelligence reasoning acceleration method based on a dedicated data processor according to claim 7, characterized in that: The calculation method for the mean value ρ of the proportion of updated parameters to all parameters in each iteration is as follows: Where: T is the total number of iterations, P l is the total number of parameters of the lth layer, is the number of parameters updated in the lth layer in the tth iteration, and L is the number of layers in the neural network model. The specific calculation steps are as follows: Calculate the number of parameters P for each layer l For the convolution layer, assuming that the convolution kernel size is filter_height×filter_width, the number of channels is in_channels, the number of output channels is out_channels, and the total number of parameters is the weighted sum of the convolution kernel weight and bias. The calculation formula is as follows: P l =P conv =P weights +P bias =filter_height×filter_width×in_channels×out_channels+out_channels For the fully connected layer, assuming that the number of input units of the fully connected layer is input_units, the number of output units is output_units, and the total parameter is the weighted sum of the number of weights and biases, the calculation formula is as follows: P l =P fc =P weights +P bias =input_units×outputs_units+output_units Calculate the number of parameters updated in layer l at iteration t Set the parameter update threshold δ threshold , by calculating the gradient of the loss function for each parameter during training, and using the gradient descent method to obtain the weight update amplitude Bias update amplitude Where η is a hyperparameter, which is set to a fixed value according to the model For the convolutional layer, the total number of updated parameters is the weighted sum of the number of weight updates and the number of bias updates. The formula for calculating the number of weight updates is: in is the weight update amount of the total convolution kernel position i, j, k, l, 1(·) represents the indicator function, which takes the value of 1 when the condition is met, otherwise it is 0; the calculation formula for the update amount of the bias is in is the update amount of the mth bias. Total number of updated parameters: For the fully connected layer, the total update parameter is also the weighted sum of the number of weight updates and the number of bias updates. The weight update number is calculated as in is the weight update of position i,j in the fully connected layer, and the bias update is calculated as in is the update amount of the kth bias. Total number of updated parameters: The total number of parameters updated per iteration: