High-frequency statistical method and application thereof in artificial intelligence large model training or reasoning
By initializing and configuring dedicated data acquisition channels on the ASIC chip, combining shared memory and timer processing, high-frequency statistics are realized, which solves the problem that ASIC chip cannot provide high-frequency traffic statistics and improves the training efficiency of artificial intelligence large models.
Patent Information
- Application Number
- CN202510417906.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-11
AI Technical Summary
Existing ASIC chips cannot provide high-frequency traffic statistics, resulting in inaccurate data synchronization in AI big model training, affecting training efficiency.
By initializing the ASIC chip, configuring a dedicated data acquisition channel, and applying for shared memory in the operating system, using self-cycle acquisition and timer to process statistical data, high-frequency statistics at millisecond level or even submillisecond level are achieved.
Provide accurate high-frequency statistics, depict the traffic characteristics of network equipment, and improve the training efficiency of artificial intelligence large models.
Smart Images

Figure CN120295679A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computers, and particularly relates to a high-frequency statistical method and its application in the training or inference of artificial intelligence large models. Background Art
[0002] When an artificial intelligence computing power network is being trained, there will be high-frequency gradient synchronization and weight updates, as well as the transmission of large-scale data sets. There will be a fierce traffic burst within a very short time, but the traffic will be balanced after a longer duration, that is, the distribution of traffic characteristics relative to time is not uniform.
[0003] When training an artificial intelligence large model, in order to improve the training efficiency, it is necessary to pursue extreme data synchronization performance. To improve the synchronization performance, network parameter tuning must be carried out, and the best state can be achieved through the collaborative processing of the server and the network. High-frequency (corresponding to a cycle of milliseconds or sub-milliseconds) traffic statistics are used for indication and verification during the network parameter tuning process.
[0004] Traditional network devices such as switches or routers have statistical functions. The basic process is to use a timer to regularly obtain various statistical data from an ASIC (Application Specific Integrated Circuit) chip simultaneously, and through inter-process data synchronization, from the ASIC chip through the ASIC chip driver, and finally store and display various statistical data in the application. However, due to technical limitations, only second-level statistics can be provided, and high-frequency statistics cannot be provided.
[0005] In addition, the existing ASIC chip driver processes all statistical data uniformly, and there are many types of statistical data in the ASIC chip. It also takes dozens of milliseconds to transmit all statistical data at once, and then the integration processing of the statistical data is required. Therefore, even if the application uses a higher-frequency timer, the obtained statistical data may not be updated in time, that is, the obtained statistical data is inaccurate.
[0006] Therefore, in the scenario of training an artificial intelligence large model, in order to improve the data synchronization efficiency, it is urgent to perform high-frequency statistics on the traffic. Summary of the Invention
[0007] To solve the foregoing technical problems, the present invention provides a high-frequency statistical method, including the following steps:
[0008] S1. Initialize the ASIC chip, register and configure dedicated data acquisition channels for statistical data;
[0009] S2. Apply for shared memory including a circular buffer in the operating system and initialize the shared memory;
[0010] S3. In the operating system, obtain the first statistical data through a self-loop;
[0011] S4. Write the first statistical data into the memory partition corresponding to the first write index value;
[0012] S5. The upper-layer application enables a timer and sets a timing period;
[0013] S6. The upper-layer application obtains and organizes the statistical data within the timing period, and updates the initial read index value of the shared memory;
[0014] S7. After the timing period arrives, repeat steps S5 to S6.
[0015] Further, the statistical data consists of the number of received and transmitted packets and the number of received and transmitted bytes corresponding to the interface.
[0016] Further, initialize the shared memory, including:
[0017] Set both the initial write index value and the initial read index value of the shared memory to 0;
[0018] Divide the memory space of the shared memory into several memory partitions.
[0019] Further, when writing the first statistical data for the first time, the first write index value = initial write index value + 1; after each subsequent write of the first statistical data, the first write index value is automatically incremented by 1.
[0020] Further, in step S3, obtaining the first statistical data through a self-loop includes:
[0021] S3.1. Call the API to obtain the first statistical data;
[0022] S3.2. After a waiting time, jump back to step S3.1.
[0023] Further, the waiting time = cycle period - execution time of the API.
[0024] Further, organizing the statistical data within the timing period includes: extracting the second statistical data from the memory partition corresponding to the first read index value % number of memory partitions to the memory partition corresponding to the first write index value % number of memory partitions, and merging it with the previously obtained first statistical data.
[0025] Further, the first read index value = initial read index value + 1.
[0026] Further, updating the initial read index value of the shared memory includes: updating the initial read index value to the initial write index value.
[0027] The present invention also provides an application of a high-frequency statistical method in the training or inference of artificial intelligence large models.
[0028] Compared with the prior art, the high-frequency statistical method provided by the present invention can provide accurate high-frequency statistical data at the millisecond level or even sub-millisecond level, accurately depict the traffic characteristics of the large model network, such as the traffic characteristics of each network device within a period of time, provide accurate data support for the optimization of the artificial intelligence large model, and improve the training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 The flowchart of the high-frequency statistical method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The following describes the specific embodiments of the present invention in detail. It should be understood that the embodiments of the present invention are not limited to the embodiments shown in the drawings, and the protection scope of the present invention is not limited by the specific embodiments. The "first", "second" and similar terms used in the present invention do not represent any order, quantity or importance, but are only used to distinguish different components. Similarly, the terms such as "a", "one" or "the" do not represent a quantity limitation, but represent the existence of at least one. Unless otherwise clearly stated, throughout the specification and claims, the term "comprising" or its variations such as "including" or "having" etc. will be understood to include the stated elements or components, and do not exclude other elements or other components. "Connection" or "connected" and similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0031] Unless otherwise defined, all technical terms and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the art. In addition, the meanings of the technical terms and scientific terms used in the present invention should be interpreted as having the same meaning as the corresponding terms defined in the common technical manuals, and should not be interpreted as having an idealized or overly formal meaning, unless the present invention clearly defines so.
[0032] Figure 1 The flowchart of the high-frequency statistical method according to an embodiment of the present invention is shown. Refer to Figure 1 The high-frequency statistical method includes the following steps:
[0033] S1. Initialize the ASIC chip and register and configure a dedicated data acquisition channel for statistical data.
[0034] Specifically, the statistical data consists of the number of transmitted and received packets and the number of transmitted and received bytes corresponding to the interface.
[0035] The amount of statistical data inside the ASIC chip is extremely large. Usually, when processing statistical data, the whole of it is transmitted to the memory. However, due to the huge amount of statistical data, the overall time consumption is relatively long. After analysis and actual verification by the inventor, the specific content of the statistical data to be transmitted is simplified, such that the statistical data consists of the number of transmitted and received packets and the number of transmitted and received bytes corresponding to the interface, thereby reducing the amount of statistical data that needs to be acquired frequently while still being able to obtain accurate statistical results.
[0036] After simulation calculation and actual verification by the inventor, the amount of statistical data, that is, the number of transmitted and received packets and the number of transmitted and received bytes, for each interface is only 4 × 8 = 32 bytes. The specific calculation process is prior art and will not be elaborated here.
[0037] Furthermore, in this embodiment, taking 64 interfaces as an example, the amount of corresponding statistical data is only 64 × 4 × 8 = 2048 bytes, that is, 2K bytes. The time taken to transmit it to the memory is only dozens of microseconds.
[0038] Specifically, the dedicated data acquisition channel is a DMA (Direct Memory Access) channel.
[0039] S2. Apply for shared memory in the operating system and initialize the shared memory.
[0040] Specifically, the operating system is the operating system of network devices such as switches or routers.
[0041] Specifically, initializing the shared memory includes:
[0042] Setting both the initial write index value and the initial read index value of the shared memory to 0;
[0043] Dividing the memory space of the shared memory into several memory partitions.
[0044] Furthermore, in this embodiment, all the statistical data for 2 seconds is stored in the shared memory. For high-frequency statistical calculations at the millisecond level, the storage space of the shared memory needs to be 64 × 4 × 8 × 2000 + management data = 4Mb. Correspondingly, the specific number of several memory partitions is 2000 memory partitions, which is the same as the number of statistical times, and the storage space of each memory partition among several memory partitions is 64 × 4 × 8 = 2048 bytes, that is, 2K bytes.
[0045] Specifically, in order to improve the utilization rate of the shared memory, simplify the data synchronization mechanism, and achieve high-speed data transmission, the shared memory includes a circular buffer.
[0046] S3. In the operating system, obtain the first statistical data through a self-loop.
[0047] Specifically, obtaining the first statistical data through a self-loop further includes:
[0048] S3.1. Call the API (Application Programming Interface) to obtain the first statistical data;
[0049] S3.2. After a waiting time, jump back to step S3.1.
[0050] Furthermore, the waiting time = cycle period - execution time of the API.
[0051] Among them, because the statistical frequency is too high, if the statistical data is written to the database or sent through the socket every time, it will seriously affect the performance of the CPU. Therefore, in the operating system, the cycle period is 1 millisecond, and the API is called through a loop to obtain the statistical data, and the execution time of the API is calculated in the operating system.
[0052] S4. Write the first statistical data into the memory partition corresponding to the first write index value.
[0053] Among them, when writing the first statistical data for the first time, the first write index value = initial write index value + 1; after each subsequent write of the first statistical data, the first write index value is automatically incremented by 1.
[0054] Specifically, in the entire business logic of high-frequency statistics, only this thread has write operations, and the upper-layer application only needs to perform read operations when obtaining statistical data, which can avoid lock operations during read-write concurrency, thereby improving the execution efficiency. Every time a statistical data is written into the circular buffer in the shared memory, the corresponding write index value is recorded in the shared memory.
[0055] S5. The upper-layer application enables a timer and sets a timing period.
[0056] Specifically, to save costs, the timer is a low-precision timer, and the timing period is 1 second.
[0057] S6. The upper-layer application obtains and organizes the statistical data within the timing period, and updates the initial read index value of the shared memory.
[0058] Specifically, organizing the statistical data within the timing period includes: extracting the second statistical data in the memory partition corresponding to the first read index value % (representing the remainder operation) of the number of memory partitions (2000 in this embodiment) to the memory partition corresponding to the first write index value % of the number of memory partitions, and merging it with the previously obtained first statistical data.
[0059] Further, the first read index value = the initial read index value + 1.
[0060] Specifically, updating the initial read index value of the shared memory includes: updating the initial read index value to the initial write index value.
[0061] S7. After the timing period arrives, repeat steps S5 - S6.
[0062] In summary, the present invention provides a high - frequency statistics method. After initializing the ASIC chip and the shared memory, in the operating system, the first statistical data is obtained through a self - loop and written into the memory partition corresponding to the first write index value; in the upper - layer application, a timer is enabled to obtain the statistical data within the timing period, organize the statistical data within the timing period, and update the initial read index value. This high - frequency statistics method can provide accurate high - frequency statistical data at the millisecond level or even sub - millisecond level, accurately depict the traffic characteristics of the large - model network, such as the traffic characteristics of each network device within a period of time, provide accurate data support for the optimization of the artificial - intelligence large model, and improve the training efficiency.
[0063] The present invention also provides an application of the high - frequency statistics method in the training or inference of an artificial - intelligence large model.
[0064] The foregoing description of the specific exemplary embodiments of the present invention is for the purposes of illustration and exemplification. These descriptions are not intended to limit the invention to the precise forms disclosed, and obviously, many changes and variations are possible in light of the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the invention and its practical applications, so that those skilled in the art can implement and utilize the various different exemplary embodiments of the invention, as well as various different selections and changes. The protection scope of the present invention is defined by the claims and their equivalents.
Claims
1. A high-frequency statistical method, characterized in that, It includes the following steps: S1. Initialize the ASIC chip and register and configure a dedicated data acquisition channel for statistical data; S2. Apply for shared memory including a circular buffer in the operating system and initialize the shared memory; S3. In the operating system, obtain the first statistical data through a self-loop; S4. Write the first statistical data into the memory partition corresponding to the first write index value; S5. The upper-layer application enables a timer and sets a timing period; S6. The upper-layer application obtains and collates the statistical data within the timing period and updates the initial read index value of the shared memory; S7. After the timing period arrives, repeat steps S5 to S6.
2. The high-frequency statistical method according to claim 1, characterized in that The statistical data consists of the number of transmitted and received packets and the number of transmitted and received bytes corresponding to the interface.
3. The high-frequency statistical method according to claim 2, wherein The initialization of the shared memory includes: Set both the initial write index value and the initial read index value of the shared memory to 0; Divide the memory space of the shared memory into several memory partitions.
4. The high-frequency statistical method according to claim 3, characterized in that When writing the first statistical data for the first time, the first write index value = initial write index value + 1; after each subsequent write of the first statistical data, the first write index value is automatically incremented by 1.
5. The high-frequency statistical method according to claim 4, characterized in that In step S3, the obtaining of the first statistical data through a self-loop includes: S3.
1. Call the API to obtain the first statistical data; S3.
2. After a waiting time, jump back to step S3.
1.
6. The high-frequency statistics method according to claim 5, characterized in that The waiting time = the cycle period - the execution time of the API.
7. The high-frequency statistical method according to claim 6, characterized in that The collation of the statistical data within the timing period includes: extracting the second statistical data from the memory partition corresponding to the first read index value % the number of memory partitions to the memory partition corresponding to the first write index value % the number of memory partitions and merging it with the previously obtained first statistical data.
8. The high-frequency statistics method according to claim 7, wherein The first read index value = the initial read index value + 1.
9. The high-frequency statistical method according to claim 6, characterized in that The updating of the initial read index value of the shared memory includes: updating the initial read index value to the initial write index value.
10. Application of the high-frequency statistical method according to any one of claims 1 to 9 in the training or inference of artificial intelligence large models.
Citation Information
Cited By
Data real-time statistical method and system
CN121614351A
A data real-time statistics method and system
CN121614351B