HBF chip multichannel parallel access control method and device

By implementing a multi-channel parallel access control method for the HBF chip, dynamically analyzing access change patterns and load status, and constructing a bandwidth reallocation strategy, the complexity of multi-channel access control is solved, the system's response speed and stability are improved, resource allocation is optimized, and access latency and conflicts are reduced.

CN122019425APending Publication Date: 2026-05-12UNITED MEMORY TECHNOLOGY (JIANGSU) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNITED MEMORY TECHNOLOGY (JIANGSU) LTD
Filing Date
2025-12-25
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing HBF chip multi-channel parallel access control methods are difficult to effectively manage and coordinate multi-channel access requests, resulting in increased access latency and frequent data conflicts. This makes it difficult to meet the real-time and fine-grained management requirements in high-concurrency environments, thus limiting the performance of the system.

Method used

By collecting real-time access data streams from each storage segment of the HBF chip, analyzing access changes, marking high-conflict access request events, obtaining access load status parameters for each channel, constructing a bandwidth reallocation strategy, and based on this, conducting feasibility assessments and distribution of high-conflict access request events, thereby realizing the analysis of dynamic access patterns and intelligent control of bandwidth resources.

Benefits of technology

It enables dynamic monitoring and optimization of chip storage resources, reduces access conflicts, improves system response speed and stability, increases bandwidth utilization and overall data processing efficiency, avoids resource waste, and ensures the response speed and processing efficiency of critical tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019425A_ABST
    Figure CN122019425A_ABST
Patent Text Reader

Abstract

The invention relates to the field of chip access control, in particular to an HBF chip multichannel parallel access control method and device. The method comprises the following steps: collecting real-time access data streams of each storage segment of the HBF chip, and carrying out access change analysis to obtain a dynamic access change rule; performing access competition situation analysis based on a dynamic access change rule, and marking a high-conflict access request event; obtaining access load state parameters of each channel, and constructing a bandwidth reallocation strategy; and performing feasibility evaluation on the high-conflict access request event based on a bandwidth reallocation strategy, determining whether the delay cost of the candidate channel is lower than the conflict resolution cost, and if so, distributing the high-conflict access request event to the candidate channel. According to the method, the access conflict is relieved, the waiting time is reduced, the access concurrency is improved, and the chip parallel access efficiency and quality are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of chip access control, and more particularly to a method and apparatus for multi-channel parallel access control of HBF chips. Background Technology

[0002] With the continuous development of semiconductor technology and integrated circuit design, HBF (High Bandwidth Memory Friendly) chips, as a high-bandwidth, high-performance memory interface solution, have been widely used in high-performance computing, artificial intelligence accelerators, and graphics processors. With the rapid increase in data processing demands, multi-channel parallel access technology has become a key means to improve the overall throughput and response speed of chips. Through its multi-channel design, HBF chips can support multiple access requests simultaneously, thereby greatly improving data access efficiency and system performance.

[0003] However, as the scale of multi-channel parallel access continues to expand, the complexity of access control also increases significantly. Due to issues such as resource contention, access conflicts, and data consistency among channels, failure to effectively manage and coordinate access requests from multiple channels can lead to increased access latency, frequent data conflicts, and even system instability and performance bottlenecks. Especially in high-concurrency access environments, traditional access control mechanisms struggle to meet the demands of real-time performance and fine-grained management, limiting the performance of HBF chips. Existing multi-channel access control methods largely rely on fixed scheduling strategies and simple priority rules, lacking intelligent awareness of dynamic access patterns and system load, resulting in uneven resource allocation and low access efficiency. Furthermore, traditional methods often fail to achieve real-time adjustment and optimization when facing complex access conflicts and multi-task concurrency, making it difficult to guarantee the dual goals of high bandwidth utilization and low-latency access. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention proposes a multi-channel parallel access control method and apparatus for HBF chips, thereby resolving at least one of the aforementioned technical problems.

[0005] To achieve the above objectives, the present invention provides a multi-channel parallel access control method for an HBF chip, comprising the following steps: Real-time access data streams of each memory segment of the HBF chip are collected, access change analysis is performed, and dynamic access change patterns are obtained. Based on the dynamic changes in access patterns, analyze the access competition situation and mark high-conflict access request events. Obtain access load status parameters for each channel and construct a bandwidth reallocation strategy; A feasibility assessment is performed on high-conflict access request events based on the bandwidth reallocation strategy to determine whether the latency cost of the candidate channel is lower than the cost of conflict resolution. If so, the high-conflict access request events are distributed to the candidate channel.

[0006] This specification provides an HBF chip multi-channel parallel access control device for executing the HBF chip multi-channel parallel access control method described above, comprising: The access change analysis module is used to collect real-time access data streams from each memory segment of the HBF chip, perform access change analysis, and obtain dynamic access change patterns. The access situation analysis module is used to analyze the access competition situation based on dynamic access change patterns and mark high-conflict access request events. The bandwidth allocation module is used to obtain the access load status parameters of each channel and construct a bandwidth reallocation strategy. The channel distribution module is used to perform a feasibility assessment on high-conflict access request events based on the bandwidth reallocation strategy, determine whether the latency cost of the candidate channel is lower than the cost of conflict resolution, and if so, distribute the high-conflict access request events to the candidate channel. The specific benefits of this invention are as follows: Real-time monitoring of access to each memory segment of the chip allows for a comprehensive understanding of storage resource usage dynamics, preventing performance bottlenecks caused by concentrated access hotspots. Analysis of dynamic access patterns provides a data foundation for subsequent access conflict prediction and bandwidth regulation, enabling visualization and controllability of access behavior. It helps identify high-frequency storage access hotspots, facilitating targeted resource allocation optimization and improving overall data processing efficiency. Competitive situation analysis allows for early detection of conflict hotspots in storage access, avoiding latency and throughput reduction caused by resource contention. Marking high-conflict events helps in targeted scheduling and optimization of storage access, reducing the system burden caused by access conflicts. Support for intelligent early warning mechanisms provides a basis for decision-making regarding bandwidth allocation and access diversion strategies, improving system response speed and stability. Real-time acquisition of channel load status enables dynamic perception of bandwidth resources, preventing resource waste caused by overloaded channels and idle channels. The construction of bandwidth reallocation strategies effectively balances the access pressure of different channels, improving bandwidth utilization and enhancing system throughput. It ensures bandwidth for critical task channels, improving the response speed and processing efficiency of important access requests. By assessing costs, intelligent decision-making is achieved, avoiding additional latency caused by blind traffic distribution and ensuring that optimization measures bring substantial performance improvements. High-conflict requests are rationally allocated to less loaded candidate channels, alleviating access conflicts, reducing waiting time, and improving access concurrency. This mechanism promotes dynamic balancing and elastic adjustment of storage access paths, enhancing the system's adaptability and stability to complex access patterns. Attached Figure Description

[0007] Figure 1This is a flowchart illustrating the steps of a multi-channel parallel access control method for an HBF chip according to the present invention. Figure 2 This is a detailed flowchart illustrating the implementation steps of step S1. Figure 3 This is a flowchart illustrating the detailed implementation steps of step S2. Detailed Implementation

[0008] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0009] This application provides a method and apparatus for multi-channel parallel access control of an HBF chip. The execution entities of the HBF chip multi-channel parallel access control method and apparatus include, but are not limited to, mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc., which can be considered as general computing nodes in this application. The data processing platform includes, but is not limited to, at least one of an audio / image management system, an information management system, and a cloud data management system.

[0010] Please see Figures 1 to 3 This invention provides a multi-channel parallel access control method for HBF chips, comprising the following steps: Real-time access data streams of each memory segment of the HBF chip are collected, access change analysis is performed, and dynamic access change patterns are obtained. Based on the dynamic changes in access patterns, analyze the access competition situation and mark high-conflict access request events. Obtain access load status parameters for each channel and construct a bandwidth reallocation strategy; A feasibility assessment is performed on high-conflict access request events based on the bandwidth reallocation strategy to determine whether the latency cost of the candidate channel is lower than the cost of conflict resolution. If so, the high-conflict access request events are distributed to the candidate channel.

[0011] In the embodiments of the present invention, see Figure 1 The diagram below illustrates the steps of a multi-channel parallel access control method for an HBF chip according to the present invention. In this example, the steps of the multi-channel parallel access control method for an HBF chip include: Real-time access data streams of each memory segment of the HBF chip are collected, access change analysis is performed, and dynamic access change patterns are obtained. In this embodiment, real-time access data streams need to be collected from each memory segment of the chip. During the collection process, each memory segment is equipped with an access monitoring node to record key information such as the timestamp of the access request, channel number, target address, and access type. Data sampling is performed at a fixed period, typically set between 10 microseconds and 100 microseconds, to balance real-time performance and data stability.

[0012] After data acquisition, the access data stream is serialized, reconstructing access behaviors from different time periods into a continuous data sequence in chronological order. Next, a time-sliding window method is used to statistically calculate access frequency, request interval, and access density, extracting the fluctuation patterns of access patterns. When the access frequency exhibits periodic or sudden changes, it is identified as a dynamic access event. By analyzing the temporal distribution and duration of these events, the dynamic access variation patterns of the HBF chip under different channel loads can be identified. This pattern reveals which memory segments are in a high-access state and which are relatively idle within a specific time period, providing a temporal and statistical basis for subsequent competitive landscape analysis.

[0013] Based on the dynamic changes in access patterns, analyze the access competition situation and mark high-conflict access request events. In this embodiment, the access frequency, number of concurrent requests, and bandwidth utilization of all channels are summarized and statistically analyzed to establish an access contention model. This model reflects the degree of resource contention between different channels within the same time window.

[0014] When multiple requests converge on the same storage segment or adjacent address regions within a certain time period, the contention model will exhibit access conflict peaks. The severity of the conflict can be determined by calculating the access overlap metric between requests. If the overlap exceeds a set threshold (e.g., 70%), requests within that time window will be marked as high-conflict access request events.

[0015] To further improve identification accuracy, access priority and data dependency are also considered. For example, if a high-priority request is delayed due to blocking by a low-priority request, it is also recognized as a form of conflict event. Ultimately, all high-conflict access request events are categorized and recorded, including the time of occurrence, the channel to which they belong, the conflict depth, and the affected memory segment number. Through this process, the HBF chip can dynamically identify areas of access contention, providing accurate input for bandwidth allocation strategies.

[0016] Obtain access load status parameters for each channel and construct a bandwidth reallocation strategy; In this embodiment, after marking high-collision events, it is necessary to understand the current load status of each channel to provide a quantitative basis for subsequent bandwidth allocation. To this end, access load parameters for each channel are collected, including key indicators such as current bandwidth utilization, average transmission rate, queue length, latency, and transmission duty cycle. These parameters reflect the current workload and potential available bandwidth of the channel.

[0017] After acquiring load data, a load balancing assessment is performed on each channel. If a channel's bandwidth utilization exceeds 90%, it is identified as a high-load channel; if it is below 50%, it is identified as an idle candidate channel. Based on this classification result, a bandwidth reallocation strategy is constructed. This strategy aims to balance channel utilization and optimizes overall access efficiency through bandwidth reallocation and task migration.

[0018] The core of the strategy lies in determining the bandwidth adjustment range based on the differences in bandwidth capacity between channels and dynamic load trends. For example, among multiple high-load channels, some low-priority requests can be preferentially transferred to low-load channels, thereby reducing the overall probability of access conflicts. By continuously updating channel status data, the bandwidth reallocation strategy has adaptive characteristics, automatically adjusting bandwidth distribution according to real-time changes in access flow, and achieving dynamic resource optimization.

[0019] A feasibility assessment is performed on high-conflict access request events based on the bandwidth reallocation strategy to determine whether the latency cost of the candidate channel is lower than the cost of conflict resolution. If so, the high-conflict access request events are distributed to the candidate channel.

[0020] In this embodiment, after the bandwidth reallocation strategy is established, a feasibility assessment of high-conflict access requests is required to determine whether they are suitable for distribution to candidate channels. The key to this assessment is comparing the latency cost of switching candidate channels with the time overhead of resolving conflicts on the original channel. First, the bandwidth capacity and current load of the candidate channel are calculated to estimate the expected waiting time for executing the request on it. Then, the conflict waiting time of the original channel is predicted to obtain the latency difference between the two. If the expected latency of the candidate channel is less than the conflict resolution time of the original channel by more than a set threshold (e.g., 10%), then channel switching is deemed feasible.

[0021] Once the feasibility conditions are met, high-conflict access requests will be preferentially transferred to candidate channels for execution. To avoid additional interference from the switchover, a phased distribution mechanism is adopted, first performing small-batch migrations in idle channels, and then performing batch adjustments after confirming a decrease in latency.

[0022] If the latency cost of a candidate channel is not significantly lower than the conflict resolution cost, the original channel is retained for execution, and the execution order is optimized through a latency compensation mechanism. Ultimately, this strategy achieves efficient scheduling of multi-channel resources, enabling the HBF chip to maintain stable throughput and low latency characteristics in high-parallel access scenarios.

[0023] In this embodiment, see Figure 2 The specific steps for collecting real-time access data streams from each memory segment of the HBF chip, performing access change analysis, and obtaining dynamic access change patterns are as follows: By deploying real-time monitoring nodes, real-time access data streams of each storage segment of the HBF chip are collected; The real-time access data stream is adaptively segmented into monitoring time windows to obtain a multi-time-window data stream; Calculate the access frequency of the multi-time-window data stream, perform time distribution analysis, and generate an access hotspot distribution map; Based on the distribution of access hotspots, short-term clustering and long-term migration are identified to obtain short-term hotspot clustering areas and long-term hot-cold migration trends. Based on the aforementioned short-term hotspot clusters and long-term hotspot migration trends, an in-depth analysis of access change trends was conducted to obtain dynamic access change patterns.

[0024] In this embodiment, lightweight monitoring nodes are deployed at the interfaces of each physical channel and logical address space of the chip. Each monitoring node has access request identification, timestamp synchronization, signal separation, and buffering functions, enabling the capture of read and write operations with nanosecond-level time accuracy. The sampling frequency of the nodes is typically set at around 500MHz to ensure accurate tracking of the data flow even under high load. To avoid interference with the main access path, a mirror buffer structure is adopted, i.e., a replica channel is established outside the main path to synchronously output access events in the form of data frames. Each monitoring node is connected to the central acquisition and control module through a high-speed DMA channel for unified data packaging and timing correction. The acquired content includes key indicators such as channel number, physical address range, access type, latency, and hit status. To ensure data quality, noise reduction, time alignment, and anomaly filtering are performed during the acquisition phase, ultimately forming a continuously input access data stream. Through this distributed real-time acquisition mechanism, panoramic data of the access behavior of each channel can be obtained without affecting the normal read and write efficiency of the HBF chip, providing a foundation for subsequent dynamic analysis. After obtaining the continuous access data stream, the time series needs to be segmented to enable phased and scaled feature extraction and statistical analysis. The time window segmentation employs an adaptive adjustment mechanism, dynamically determining the time window length based on access load and volatility. This improves sampling accuracy during high-frequency access phases and reduces redundant data during low-frequency phases. System activity is assessed by calculating the sliding variance of the access event intervals. When the variance exceeds a set threshold, the time window is shortened (e.g., from 5 milliseconds to 1 millisecond); when the variance falls below a lower limit, the window is extended (e.g., from 5 milliseconds to 10 milliseconds), ensuring both real-time performance and stability in the analysis process. To avoid information loss due to boundary effects, a window overlap strategy is used, typically with an overlap ratio of 50%. Data from each time window is independently cached, forming a multi-time-window access data stream set. Temporal continuity is maintained between windows, and each window is accompanied by a window number, start and end times, and data packet identification information.

[0025] Within each time window, the number of accesses to each storage segment is counted, and the access intensity per unit time is obtained using the access frequency calculation formula f_i = N_i / Δt. To reduce statistical fluctuations caused by sudden accesses, a moving average or exponential smoothing method is used for data flattening. Then, the data from all time windows are constructed into a two-dimensional access distribution matrix according to time order and storage segment number, where the horizontal axis represents the time window number, the vertical axis represents the storage segment number, and the matrix elements represent the access frequency at the corresponding time. Fourier analysis of this matrix in the time direction reveals the periodic pattern of access behavior, while density estimation in the spatial direction (such as Gaussian kernel density function) yields the spatial distribution characteristics of access hotspots. The generated access hotspot distribution map uses color gradients to represent access intensity, thus forming a heat map that evolves over time. This map clearly reflects the degree of access concentration of different channels at different stages, as well as the appearance, disappearance, and migration patterns of hotspot areas. This process provides a basic data model for subsequent clustering analysis and migration detection. After obtaining the access hotspot distribution, it is necessary to further identify clustering phenomena with concentrated access frequencies in a short period of time, and migration trends of hot and cold changes in storage areas over a longer time scale. The identification of short-term clusters can be achieved based on time-correlation clustering methods. By calculating the similarity of hotspot locations and intensities between adjacent time windows, regions that appear consecutively and whose access frequency increases significantly are grouped into the same cluster. A time decay coefficient is introduced into the distance metric to emphasize the impact of changes in the most recent time window on the clustering results. Long-term migration identification is achieved through trend analysis of access frequency time series. A linear fitting method is used to calculate the slope of the change in access intensity over time to determine the direction of heat change in a certain storage segment. When the slope is positive and exceeds a set threshold, it is determined as "cold to hot" migration, and vice versa, "hot to cold" migration. Furthermore, the persistence of changes can be quantified by combining time correlation coefficients. Through this identification mechanism, the active change patterns of different storage segments can be discovered in multi-channel systems, distinguishing between short-term burst loads and long-term structural changes, thereby providing a basis for access scheduling and storage resource reallocation decisions. The entire process enables dynamic identification and trend tracking of access status, giving the system a certain degree of adaptive adjustment capability. Short-term aggregation results are fused with long-term migration trend data to form a comprehensive vector set containing multi-dimensional features such as local popularity, global trend, slope of change, and correlation coefficient. Subsequently, principal component analysis or feature reduction methods are used to extract the main directions of change to identify key factors influencing access patterns. In the trend prediction stage, an autoregressive sliding model is used to predict future changes in access popularity, thereby proactively identifying potential access concentration areas or channel overload risks. To assess the overall system stability, an information entropy index is introduced to calculate the complexity of the access distribution. A significant increase in entropy indicates a drastic change in access patterns, requiring dynamic adjustment of access control parameters.This analysis process organically combines short-term burst characteristics with long-term evolutionary trends, providing a macro-level understanding of the dynamic access behavior of the HBF chip. Through trend analysis results, adaptive optimization of access control strategies can be achieved, including channel scheduling weight adjustment, dynamic data block migration, and cache reconstruction. This improves access efficiency and system response performance in multi-channel parallel scenarios, enabling dynamic controllability and predictability of access patterns.

[0026] In this embodiment, the specific steps for analyzing access contention based on dynamic access change patterns and marking high-conflict access request events are as follows: Based on the dynamic access change pattern, the future access data volume is predicted and calculated to obtain the parallel access data volume in the future time period. The current access waiting queue is determined based on the amount of real-time access data. Identify the access target of the current access waiting queue; calculate the access overlap based on the access target to obtain the queue access overlap. Based on the parallel access data volume and queue access overlap diagram, a multi-channel access competition situation analysis is performed to obtain the multi-channel competition situation. Based on the multi-channel competition situation, high conflict risk is identified, the target address, timestamp and priority of the access are extracted, and high conflict access request events are marked.

[0027] In this embodiment, the dynamic access change patterns obtained from the previous stage analysis are used as the input feature set, including multi-dimensional indicators such as short-term hotspot clustering information, long-term hot-cold migration slope, and access frequency change rate. These features are then mapped onto a time axis to form continuous access sequence data. Using an autoregressive moving average model (ARIMA) or a time prediction method based on recursive feature networks (RNN), the access data volume within several future time slices (e.g., 10ms, 50ms, 100ms) is predicted. The model includes two parts: a trend term describing the overall growth direction of the access load, and a fluctuation term reflecting the uncertain fluctuations in short-term access. The prediction output is a sequence of parallel access data volume for each future time period, with units of "requests / millisecond" or "MB / s". To improve prediction accuracy, the system also introduces access periodic autocorrelation coefficients and inter-channel interference factors for correction. When the predicted access volume exceeds a set threshold (e.g., 85% of the single-channel throughput limit), a parallel access load early warning mechanism is triggered, providing a quantitative basis for subsequent access queue scheduling and conflict risk identification. The system extracts information such as the access request ID, access type (read / write), target address, data length, and initiation timestamp from the real-time request stream output by the access controller. Then, it writes these requests into a circular queue buffer in chronological order of arrival. To reflect the parallelism between different channels, multiple sub-queues are maintained simultaneously, each corresponding to a different physical channel or address range. The system calculates the current load factor for each channel based on access frequency thresholds and channel occupancy rates. When the load of a channel exceeds a set value (e.g., 0.8), a waiting mechanism is automatically triggered, temporarily suspending new access requests at the end of the queue. This waiting queue contains not only actual access instructions to be executed but also initiated but unresponded requests for subsequent overlap calculations and conflict analysis. To improve real-time performance, the queue update cycle is controlled at the microsecond level, and a circular caching mechanism is used to avoid memory overflow. This process enables dynamic capture of the HBF chip's access status, providing reliable real-time input for competitive situation modeling.

[0028] The system parses the target address range of each request in the queue to obtain its mapped segment in the physical storage space. Based on the HBF chip's storage mapping structure, the physical address is uniformly mapped to a two-level identifier: channel number and block number. Subsequently, the overlap ratio of any two access requests in the address space is calculated, i.e., the access address overlap. If the access intervals of two requests partially or completely overlap in the physical storage unit, they are marked as a potential conflict pair. Furthermore, to assess the time-level competition, access time overlap also needs to be considered, calculated by comparing the access start time and the expected completion time. Combining these two factors, the overall access overlap index is obtained, defined as: O = α·S + β·T, where S is the spatial overlap rate, T is the temporal overlap rate, and α and β are weighting coefficients (usually set to 0.6 and 0.4). A higher metric indicates a greater risk of conflict between accesses. To ensure real-time performance, the system employs a hierarchical parallel computing mechanism, performing overlap calculations for different channels in parallel. This ultimately forms a queue access overlap matrix, reflecting the correlation between any two access requests and providing data support for competitive situation analysis. The amount of future access prediction data is used as the access intensity input in the temporal dimension, and the queue access overlap matrix is ​​used as the competitive relationship input in the spatial dimension. A comprehensive competitive model is constructed as follows: C(i,t) = F(i,t) × O(i,j), where F(i,t) represents the access intensity of channel i at time t, and O(i,j) represents the access overlap between channels i and j. A multi-channel competition intensity distribution map can be obtained through matrix calculation. The system further calculates the competition index K_c = ΣC(i,t) / N to quantify the overall competition situation. When K_c exceeds a threshold (e.g., 0.7), it indicates that the system is in a high-competition state. The analysis results can be mapped to a visual competition heatmap, using color gradients to represent competition intensity. This map shows the mutual interference of access requests between channels over time and the areas of concentrated conflict. Based on this situational result, the system can identify which channels may experience bandwidth contention, response delays, or resource locking in future periods, providing a basis for decision-making in subsequent conflict risk identification.

[0029] The system filters regions with competition indices exceeding a threshold from the competition landscape map, designating the corresponding channels and time windows as high-risk areas. Then, it extracts key attributes of access requests within these regions, including the target address, access initiation timestamp, access priority level, and corresponding channel number. By analyzing the interactions between these access requests, the system further calculates the conflict contribution value C_r for each request, representing its impact on the overall competition intensity. If a request's C_r exceeds 1.5 times the average, it is identified as a high-conflict access request event. The system marks this event and outputs it to the access scheduling module as the basis for subsequent access arbitration and priority reordering. To avoid duplicate marking, the system implements a time decay mechanism, automatically removing the mark when a conflicting request does not generate high-conflict characteristics for several consecutive cycles. This mechanism enables real-time identification and dynamic management of high-conflict risks during multi-channel access of the HBF chip, effectively reducing performance fluctuations caused by channel contention and improving the overall parallel access efficiency and response consistency of the system.

[0030] In this embodiment, the specific steps for obtaining the access load status parameters of each channel and constructing the bandwidth reallocation strategy are as follows: Obtain the access load status parameters of each channel; calculate the request queue depth, data throughput rate, bandwidth utilization ratio and idle computing resources based on the access load status parameters to obtain the multi-dimensional load status characteristics of each channel; Based on the multidimensional load state characteristics, the inter-channel load imbalance is calculated to obtain the load imbalance index. The resource allocation difference coefficient between channels is quantified based on the load imbalance index; Based on the resource configuration difference coefficient, dynamic bandwidth resource reallocation is performed to obtain a bandwidth reallocation strategy.

[0031] In this embodiment, an embedded monitoring module continuously records access behavior within the channel, extracting core indicators reflecting the load level. These mainly include: number of access requests, average access latency, data throughput rate, bandwidth utilization, cache hit rate, and idle computing cycle percentage. The monitoring module is implemented as a hardware counting unit within the channel controller, with a sampling period typically set to 10 microseconds to ensure timely capture of load changes. All collected parameters are synchronized to the system control unit via a bus and uniformly aligned on the timeline to ensure comparability of parameters across different channels. To prevent the impact of instantaneous fluctuations on statistical results, the system uses moving average and exponential smoothing algorithms to dynamically correct the original data. The resulting access load status parameters not only reflect the current channel's operational intensity but also its data flow and resource utilization efficiency. These parameters serve as the foundational input for subsequent multi-dimensional load characteristic calculations, providing high-precision data support for inter-channel load balancing judgment and bandwidth allocation optimization. A structured analysis is performed on the access queue of each channel to calculate the current request queue depth. The queue depth is represented by the ratio of the number of incomplete requests to the maximum queue capacity, reflecting the current congestion level of the channel. Secondly, the actual data throughput rate within a fixed time window is calculated by dividing the amount of data accessed by the time window width. To reflect resource utilization, the bandwidth utilization ratio, i.e., the ratio of the current throughput rate to the theoretical bandwidth limit, also needs to be calculated. Simultaneously, to reflect the channel's idle capacity, an idle computing resource metric is introduced to measure the proportion of available cycles that can be used for additional processing per unit of time. These four metrics together constitute the channel's multi-dimensional load state feature vector: L = [Qd, Ts, Bw, Fr], where Qd is the request queue depth, Ts is the throughput rate, Bw is the bandwidth utilization rate, and Fr is the proportion of idle resources. To avoid bias from a single metric, the system performs comprehensive weighting through normalization and weight coefficient allocation (e.g., Qd: 0.4, Ts: 0.3, Bw: 0.2, Fr: 0.1). In this way, different performance metrics from multiple channels can be uniformly mapped to the same dimensional space, thus forming a channel load characteristic model that can be used for comparison and analysis, providing a basic data structure for subsequent load balancing decisions.

[0032] Using the multidimensional feature vector of each channel as input, the relative differences between channels are calculated using Euclidean distance or cosine similarity. First, the overall load value of all channels is normalized, and its average load level L_avg is calculated. Then, the deviation ΔL_i = |L_i - L_avg| is calculated for each channel, and the standard deviation σ_L characterizes the overall load dispersion of the system. To improve sensitivity to abnormal channels, a weighting factor k_i is introduced to amplify the weight of channels with high bandwidth occupancy or high queue depth, thereby highlighting potential congestion points. The overall calculated load imbalance index D_load can be defined as: Where N is the number of channels. The larger this index is, the more significant the load difference between channels, and the more unbalanced the system. By continuously monitoring the changing trend of D_load, periods of uneven load distribution can be dynamically identified, providing a basis for resource reallocation decisions. This method ensures that the system can detect performance bottlenecks and areas of concentrated load in real time, thereby achieving precise dynamic optimization.

[0033] After standardizing the available resource parameters (bandwidth, buffer capacity, computing unit utilization) of each channel, a resource configuration matrix R is formed. Then, the load imbalance index D_load obtained in the previous step is mapped to the resource space, and the resource matching degree M_i = 1 - |L_i - R_i| is calculated for each channel. Finally, by calculating the deviation between the global average matching degree and the matching degree of each channel, the resource configuration difference coefficient C_diff is obtained. C_diff = Σ|M_i - M_avg| / N.

[0034] This coefficient reflects the degree of deviation between the use of different channel resources and load demand. A large C_diff indicates that resource allocation has failed to effectively match actual load demand; a C_diff close to zero indicates that system resource utilization is approaching equilibrium. To facilitate real-time judgment, the system compares C_diff with a preset threshold (e.g., 0.2), and automatically triggers a bandwidth reallocation mechanism when it exceeds the threshold. This metric quantifies the system's resource distribution efficiency at different points in time, making the subsequent dynamic allocation process measurable and adaptive.

[0035] When the system detects that the resource configuration difference coefficient exceeds a set threshold, it enters the dynamic bandwidth resource reallocation phase. The goal of this step is to achieve adaptive reallocation of bandwidth resources among multiple channels while maintaining overall system stability. First, the system determines the target channel set requiring adjustment based on the calculation results of the previous phase, marking high-load channels as the "bandwidth compensation zone" and low-load channels as the "bandwidth relinquishment zone." Then, based on channel priority, task type, and access density, the bandwidth adjustment ratio for each channel is calculated. Where γ is the global adjustment coefficient used to control the adjustment intensity. The adjustment strategy follows the principle of "constant total bandwidth, dynamic change in allocation ratio," meaning that the total available bandwidth of the system remains unchanged, but the allocation weight between channels is dynamically updated according to the real-time load. To prevent scheduling jitter caused by frequent fluctuations, the system sets a minimum adjustment interval and an allocation smoothing coefficient to make the bandwidth change a continuous and gradual trend. After the bandwidth reallocation strategy is output, the system control module dynamically modifies the channel scheduling table and access arbitration priority according to the strategy to achieve real-time rebalancing of resources. The final result is that high-load channels obtain more data path bandwidth, and low-load channels correspondingly reduce resource consumption, thereby significantly reducing the load difference between channels and improving the overall parallel access efficiency of the HBF chip and the system throughput stability.

[0036] In this embodiment, the specific steps for conducting a feasibility assessment of high-conflict access request events based on the bandwidth reallocation strategy, determining whether the latency cost of the candidate channel is lower than the cost of conflict resolution, and if so, distributing the high-conflict access request event to the candidate channel; otherwise, suspending the high-conflict access request event for waiting for access are as follows: Based on the bandwidth reallocation strategy, idle channel planning is performed on high-conflict access request events to obtain candidate channels; Calculate the bandwidth capacity, expected waiting delay, and switching cost with the original channel of the candidate channel to obtain the channel switching matching coefficient; Feasibility assessment is performed based on the channel switching matching coefficient to determine whether the latency cost of the candidate channel is lower than the cost of conflict resolution. If so, high-conflict access request events are distributed to the candidate channel. If not, suspend the high-conflict access request event and wait for access; Perform latency compensation processing on real-time access data streams.

[0037] In this embodiment, when the HBF chip identifies a high-conflict access request event during multi-channel parallel access, it needs to mitigate the conflict through a bandwidth reallocation strategy. The core of the bandwidth reallocation strategy lies in dynamically allocating idle or low-load channels while maintaining data consistency. First, the real-time bandwidth utilization of all channels is sampled, with the sampling period typically set between 50 and 200 microseconds to capture bandwidth fluctuation characteristics. Subsequently, based on the sampled data, the current load percentage and bandwidth idle rate of each channel are calculated, and channels with a bandwidth utilization rate below 40% are selected as candidate channels.

[0038] During the selection process, the physical topology and data path latency between channels must also be considered. For example, if channels share the same bus branch, contention risks may arise; therefore, independent channels are prioritized during planning. Furthermore, to ensure the continuity of access for high-priority tasks, the planning strategy prioritizes matching channels that are topologically adjacent to the original access channel and have lower switching costs. Finally, by combining channel bandwidth utilization, topology location, and access direction consistency, a candidate channel set is generated, providing a decision-making basis for subsequent matching and switching evaluation. After obtaining the candidate channel set, the comprehensive availability index of each candidate channel needs to be further calculated to determine the most suitable channel for handling high-conflict access requests. The calculation process first evaluates the bandwidth capacity of the candidate channel by statistically analyzing the channel's average bandwidth utilization over the past 100 milliseconds and combining it with the maximum bandwidth value to calculate the remaining bandwidth capacity. Subsequently, based on the current task queue length and request execution time of the channel, its expected waiting latency is estimated. If the channel queue is long, the waiting latency increases proportionally to reflect the actual processing load.

[0039] The channel switching cost mainly consists of data cache remapping, address mapping latency, and channel context switching time. To quantify this cost, channel switching latency is statistically analyzed at the nanosecond level, and switching energy consumption is introduced as a supplementary parameter. When the switching cost is high, the matching coefficient decreases accordingly. Finally, the channel switching matching coefficient is calculated by combining the three parameters of bandwidth capacity, waiting latency, and switching cost. This coefficient reflects the balance between bandwidth utilization and switching efficiency of the candidate channel, providing a quantitative basis for subsequent feasibility assessment. After obtaining the switching matching coefficient for each candidate channel, its feasibility needs to be comprehensively evaluated to determine whether the channel switching has advantages in terms of time and energy consumption. The evaluation process uses latency cost as the core criterion. First, the expected waiting latency of the candidate channel is added to the switching cost to obtain the total switching latency cost; then, this cost is compared with the conflict waiting time required if the current high-conflict request is not switched. If the latency after switching is lower than the cost of waiting for conflict resolution, it is determined to be a feasible switch.

[0040] In practice, to avoid performance fluctuations caused by frequent switching, a latency cost threshold deviation range is set. If the candidate channel switching latency is only 5% lower than the original channel, migration is not performed; if it is more than 10%, channel redirection is immediately executed. During channel switching, the address of the access request is mapped to the cache space corresponding to the new channel, and the access index table is updated to ensure data consistency. After evaluation, high-conflict requests that meet the criteria are distributed to the new channel, achieving dynamic bandwidth load balancing and effectively reducing the peak pressure of multi-channel access competition. For high-conflict access request events that do not meet the switching criteria, a suspension mechanism is used to wait for access to prevent further competition for channel resources. Suspension is implemented by setting a latency flag for the access request, which records the initial timestamp, waiting duration, and priority status of the request. If the suspended request is a high-priority task, it is kept in the queue for monitoring and will participate in the evaluation again in the next time window; if it is a low-priority request, it enters the latency buffer to wait for the next round of bandwidth release.

[0041] During the waiting process, the load changes of the relevant channels are continuously monitored. When the bandwidth utilization drops to a preset threshold (e.g., below 60%), the suspended requests are automatically woken up and reallocated to available channels. To avoid excessively long waiting times leading to excessive response latency, a maximum waiting time limit parameter is set, typically not exceeding 500 microseconds. Requests that have not received an execution opportunity after this time will trigger an emergency scheduling mechanism, forcing them to participate in the next round of channel matching.

[0042] This suspension mechanism ensures the orderly execution of access tasks under high load, preventing performance degradation caused by bandwidth contention, and providing reference data for the next stage of latency compensation calculation. After high-conflict access requests are reallocated or suspended, latency compensation needs to be performed on the overall access process to correct the time deviation caused by channel switching and waiting. The core objective of latency compensation is to maintain the stability of the HBF chip's access timing, ensuring that the actual access latency is consistent with the theoretical value. First, the execution time of each access request is compared with the originally planned execution time to calculate the latency deviation value. Subsequently, high-latency requests are compensated for by dynamically adjusting the channel priority queue, cache refresh frequency, and access interval.

[0043] During compensation, requests with latency deviations exceeding twice the average are prioritized for optimization. The overall latency curve is balanced by shortening the access interval in the next stage. Simultaneously, to avoid triggering a new round of contention during compensation, a smoothing factor is introduced during the adjustment process, ensuring a gradual change in latency recovery. Ultimately, the compensated real-time access data stream is re-aligned with the bandwidth allocation strategy, maintaining the synchronization and stability of the multi-channel access process.

[0044] In this embodiment, the specific steps for performing latency compensation processing on the real-time access data stream are as follows: Extract the full-link access timestamp from the real-time access data stream; The timestamps of all access stages are determined based on the full-link access timestamps. The time delay for each step is calculated based on the timestamps of all access steps, including request enqueue delay, queue waiting delay, transmission delay, storage access delay, and result return delay. The time delay is identified by identifying the delay characteristics of each stage, and delay attribution analysis is performed to generate stage delay attribution.

[0045] Refined delay compensation processing is performed based on process delay attribution.

[0046] In this embodiment, during the multi-channel parallel access process of the HBF chip, in order to accurately assess the distribution of access latency, it is necessary to extract the full-link access timestamp from the real-time access data stream. The extraction scope of the full-link timestamp covers the entire process from access request generation, queuing, channel scheduling, data transmission to result return. First, in the access request generation stage, each request is assigned an initial timestamp to record its trigger time at the bus level. Subsequently, a high-precision clock synchronization module controls the timestamp accuracy to the nanosecond level to ensure the comparability of data sampling between different channels.

[0047] During the extraction process, a parallel sampling mechanism is employed to track access requests, simultaneously capturing the arrival and departure times of requests across multiple data channels. Each data stream is accompanied by an independent time sequence identifier, and cross-channel matching is achieved through access identifiers (such as channel IDs, address indexes, etc.). To prevent data loss or out-of-order issues, timestamp data is buffered and reassembled in an orderly manner to ensure the integrity and continuity of the entire timeline. The final end-to-end access timestamps record the precise time nodes of each request at different stages, providing a reliable timeline foundation for subsequent latency calculations and feature recognition.

[0048] After obtaining the complete end-to-end access timestamps, it is necessary to further decompose the time nodes of each stage of the access process to clarify the flow of the request at each stage. The access process of the HBF chip can be divided into several key stages, including request queuing, queue waiting, bus transmission, storage access, and result return. By analyzing the continuous relationship between timestamps, the start and end times of each stage can be accurately determined.

[0049] In the specific implementation, the data is first grouped according to the event type identifier in the timestamp sequence. For example, "REQ_START" indicates a request is initiated, "BUS_ENTRY" indicates entry into the transmission channel, "MEM_ACCESS" indicates access to the storage array, and "RESP_END" indicates the return of a result. By matching the time differences between adjacent identifiers, the time range of each stage can be obtained. To ensure the accuracy of segmentation, a time series consistency detection mechanism is introduced. When the time order is abnormal or an event is lost, time is filled in by interpolation between preceding and following frames.

[0050] Ultimately, the timestamps of all access stages were compiled into a structured data table, containing the start and end times of each access request at different stages, providing a hierarchical basis for subsequent delay calculations and attribution analysis.

[0051] After determining the timestamps for each stage, the time intervals for each stage need to be calculated to quantify the composition of access latency. Each access request has independent latency characteristics at different stages. For example, request enqueue latency reflects the time difference between the arrival of the request and its registration in the queue; queue waiting latency represents the time the request waits in line for channel allocation; transmission latency reflects the propagation time of data in the bus and crossbar network; storage access latency reflects the addressing and data retrieval speed of the storage array; and result return latency reflects the time for the return and acknowledgment of response data.

[0052] During the delay calculation process, the timestamp difference is processed at the nanosecond level and averaged using multiple sampling results to reduce the impact of random fluctuations. To improve accuracy, outlier values ​​are filtered out using an outlier detection method. When a sudden change occurs in the access delay, the delay calculation module reconfirms the logical relationship of the timestamps to avoid statistical bias caused by time overlap or sampling errors.

[0053] The final result is a set of latency distribution data that accurately reflects the time consumption structure of access requests at different stages, providing high-resolution basic data support for latency feature identification.

[0054] After calculating the time delay for each stage, it is necessary to further identify delay characteristics and perform attribution analysis. Delay characteristic identification involves comparing the delay distribution of each stage to determine the main performance bottlenecks. First, statistical analysis is performed on each type of delay data, calculating its mean, variance, and maximum value to identify stages with large fluctuations. If queue waiting delay accounts for more than 50% of the total delay, it indicates insufficient scheduling efficiency; if storage access delay is significantly higher than the average, it may be related to storage array conflicts or a decrease in cache hit rate.

[0055] To attribute latency, it's necessary to analyze the contextual data. For example, during periods of frequent access conflicts, transmission latency often accompanies increased queue latency, indicating that the latency stems from channel contention. If latency is concentrated in the result return phase, it may be related to output bandwidth limitations. Through multidimensional data correlation analysis, each type of latency can be mapped to specific influencing factors, such as bandwidth usage, data remapping overhead, or cache refresh frequency.

[0056] The final generated process delay attribution results are output in a categorized format, clearly indicating the main sources of delay in each process and their contribution ratio, providing precise target guidance for subsequent delay compensation.

[0057] After identifying the sources of latency at each stage, refined latency compensation is needed to optimize the overall access response performance of the HBF chip. The compensation process employs differentiated strategies for different types of latency. For latency in the request queuing and waiting phases, compensation is achieved by dynamically adjusting request priorities and queue allocation weights. When excessive transmission latency is detected, network congestion is reduced through bandwidth fine-tuning and channel reallocation. If latency is mainly concentrated in the storage access phase, performance can be corrected by increasing prefetch depth and implementing cache replacement strategies.

[0058] The compensation strategy employs a gradual adjustment mechanism, with each adjustment controlled within 10% to avoid timing oscillations. After compensation, real-time monitoring of latency changes at each stage continues, and the adjustment results are solidified into operating parameters once the latency decline trend stabilizes. To ensure the overall balance of compensation, latency correction is performed in time slices, typically with a period of 1 to 5 milliseconds. Through refined latency compensation, the HBF chip can adaptively adjust access timing in multi-channel parallel access, significantly reducing average response latency and improving access stability, resulting in more efficient multi-channel access control performance.

[0059] In this embodiment, the specific steps for refined delay compensation processing based on link delay attribution are as follows: Perform end-to-end delay mapping on the delay attribution of the aforementioned links to generate an end-to-end delay mapping diagram; Based on the end-to-end latency mapping, latency deviation trend analysis is performed to identify high-latency requests and low-latency requests. High-latency requests are processed by sending access requests in advance, while low-latency requests are processed by adjusting the access timing.

[0060] In this embodiment, after obtaining the latency attribution results for each stage, the scattered latency data needs to be integrated to construct a full-link latency mapping map reflecting the overall access process characteristics. This mapping map forms a visualized latency distribution model by linearly superimposing latency data from multiple access stages over time and matching it with spatial structure. Specifically, the stage latency data is first sorted according to the access order, corresponding to stages such as request initiation, queue waiting, transmission, storage access, and response return. The latency of each stage is marked on a unified timeline and displayed as nodes. The connections between nodes represent the access flow path, and their color intensity or marker density reflects the latency strength.

[0061] To ensure mapping accuracy, a time synchronization and channel mapping mechanism is introduced, enabling the access latency of different channels to be aligned under the same time base. On channels with high accumulated latency, the mapping map displays clearly highlighted areas, indicating potential access bottlenecks. This mapping visually presents the latency distribution characteristics of the HBF chip during multi-channel parallel access, revealing latency differences and dependencies between channels and stages. The generated end-to-end latency mapping map not only provides a data foundation for latency trend analysis but also offers a visual reference for the formulation of subsequent access scheduling strategies.

[0062] After the end-to-end latency mapping is generated, trend analysis of the data is required to identify fluctuation patterns and abnormal distribution characteristics of access latency. The core objective of the analysis is to distinguish between high-latency and low-latency types of access requests, thereby providing a basis for subsequent access scheduling. Specific methods include latency deviation calculation and trend modeling. First, the access latency curves of different channels are normalized so that the latency characteristics of each channel can be compared on the same scale. Then, the mean and deviation of latency are calculated using a sliding time window to identify requests that exceed the average level by a certain threshold; these requests are judged as high-latency requests. Conversely, requests that are below the average by a certain proportion are considered low-latency requests.

[0063] To further distinguish between short-term anomalies and long-term trends, time series regression analysis was used to extract the slope of latency data changes over continuous periods. When the latency of a certain channel continues to rise, the corresponding deviation trend in the mapping graph is marked as a red high-risk area; while areas with low-latency requests are displayed as stable areas or blue areas. Ultimately, through trend analysis, the types of high-latency requests that cause performance fluctuations in multi-channel parallel access can be accurately identified, providing a basis for decision-making in subsequent access priority reconstruction and scheduling optimization.

[0064] After identifying latency deviations, targeted access optimization strategies need to be implemented based on different latency types. For high-latency requests, to reduce their impact on overall access efficiency, a pre-sending approach is adopted. Specifically, when an access request is detected to exhibit a high latency trend in historical accesses, it is inserted into the early stage of the scheduling queue, allowing bandwidth to be allocated in advance when the channel is idle or about to release resources. This effectively reduces the waiting time of high-latency requests and alleviates access bottlenecks.

[0065] For low-latency requests, an access timing adjustment strategy is adopted to avoid these requests excessively preempting resources and causing overall performance imbalance. By introducing a small access timing offset for low-latency requests, requests of different priorities can form a reasonable interleaved execution relationship in time, thereby achieving a balanced improvement in resource utilization.

[0066] To ensure the stability of timing adjustments, an adaptive latency compensation mechanism is introduced to dynamically adjust the timing offset based on real-time access traffic and channel load. Ultimately, by pre-sending high-latency requests and timing-tuning low-latency requests, multi-channel parallel access can maintain high throughput and low access fluctuations under high load conditions, achieving dynamic optimization of overall access latency.

[0067] In this embodiment, an HBF chip multi-channel parallel access control device is provided for executing the HBF chip multi-channel parallel access control method described above, including: The access change analysis module is used to collect real-time access data streams from each memory segment of the HBF chip, perform access change analysis, and obtain dynamic access change patterns. The access situation analysis module is used to analyze the access competition situation based on dynamic access change patterns and mark high-conflict access request events. The bandwidth allocation module is used to obtain the access load status parameters of each channel and construct a bandwidth reallocation strategy. The channel distribution module is used to perform a feasibility assessment on high-conflict access request events based on the bandwidth reallocation strategy, determine whether the latency cost of the candidate channel is lower than the cost of conflict resolution, and if so, distribute the high-conflict access request events to the candidate channel. Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0068] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein are implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for multi-channel parallel access control of an HBF chip, characterized in that, Includes the following steps: Real-time access data streams of each memory segment of the HBF chip are collected, access change analysis is performed, and dynamic access change patterns are obtained. Based on the dynamic changes in access patterns, analyze the access competition situation and mark high-conflict access request events. Obtain access load status parameters for each channel and construct a bandwidth reallocation strategy; A feasibility assessment is performed on high-conflict access request events based on the bandwidth reallocation strategy to determine whether the latency cost of the candidate channel is lower than the cost of conflict resolution. If so, the high-conflict access request events are distributed to the candidate channel.

2. The HBF chip multi-channel parallel access control method according to claim 1, characterized in that, The specific steps for collecting real-time access data streams from each memory segment of the HBF chip, performing access change analysis, and obtaining dynamic access change patterns are as follows: By deploying real-time monitoring nodes, real-time access data streams of each storage segment of the HBF chip are collected; The real-time access data stream is adaptively segmented into monitoring time windows to obtain a multi-time-window data stream; Calculate the access frequency of the multi-time-window data stream, perform time distribution analysis, and generate an access hotspot distribution map; Based on the distribution of access hotspots, short-term clustering and long-term migration are identified to obtain short-term hotspot clustering areas and long-term hot-cold migration trends. Based on the aforementioned short-term hotspot clusters and long-term hotspot migration trends, an in-depth analysis of access change trends was conducted to obtain dynamic access change patterns.

3. The HBF chip multi-channel parallel access control method according to claim 1, characterized in that, The specific steps for analyzing access contention based on dynamic access change patterns and marking high-conflict access request events are as follows: Based on the dynamic access change pattern, the future access data volume is predicted and calculated to obtain the parallel access data volume in the future time period. The current access waiting queue is determined based on the amount of real-time access data. Identify the access target of the current access waiting queue; The access overlap is calculated based on the access target to obtain the queue access overlap. Based on the parallel access data volume and queue access overlap diagram, a multi-channel access competition situation analysis is performed to obtain the multi-channel competition situation. Based on the multi-channel competition situation, high conflict risk is identified, the target address, timestamp and priority of the access are extracted, and high conflict access request events are marked.

4. The HBF chip multi-channel parallel access control method according to claim 1, characterized in that, The specific steps for obtaining the access load status parameters of each channel and constructing the bandwidth reallocation strategy are as follows: Obtain the access load status parameters of each channel; calculate the request queue depth, data throughput rate, bandwidth utilization ratio and idle computing resources based on the access load status parameters to obtain the multi-dimensional load status characteristics of each channel; Based on the multidimensional load state characteristics, the inter-channel load imbalance is calculated to obtain the load imbalance index. The resource allocation difference coefficient between channels is quantified based on the load imbalance index; Based on the resource configuration difference coefficient, dynamic bandwidth resource reallocation is performed to obtain a bandwidth reallocation strategy.

5. The HBF chip multi-channel parallel access control method according to claim 1, characterized in that, The specific steps for conducting a feasibility assessment of high-conflict access request events based on the bandwidth reallocation strategy, determining whether the latency cost of candidate channels is lower than the cost of conflict resolution, and if so, distributing the high-conflict access request events to candidate channels; otherwise, suspending the high-conflict access request events for waiting for access, are as follows: Based on the bandwidth reallocation strategy, idle channel planning is performed on high-conflict access request events to obtain candidate channels; Calculate the bandwidth capacity, expected waiting delay, and switching cost with the original channel of the candidate channel to obtain the channel switching matching coefficient; Feasibility assessment is performed based on the channel switching matching coefficient to determine whether the latency cost of the candidate channel is lower than the cost of conflict resolution. If so, high-conflict access request events are distributed to the candidate channel. If not, suspend the high-conflict access request event and wait for access; Perform latency compensation processing on real-time access data streams.

6. The HBF chip multi-channel parallel access control method according to claim 5, characterized in that, The specific steps for performing latency compensation processing on the real-time access data stream are as follows: Extract the full-link access timestamp from the real-time access data stream; The timestamps of all access stages are determined based on the full-link access timestamps. The time delay for each step is calculated based on the timestamps of all access steps, including request enqueue delay, queue waiting delay, transmission delay, storage access delay, and result return delay. The time delay is identified by identifying the delay characteristics of each stage, and delay attribution analysis is performed to generate stage delay attribution. Refined delay compensation processing is performed based on process delay attribution.

7. The HBF chip multi-channel parallel access control method according to claim 6, characterized in that, The specific steps for refined delay compensation based on process delay attribution are as follows: Perform end-to-end delay mapping on the delay attribution of the aforementioned links to generate an end-to-end delay mapping diagram; Based on the end-to-end latency mapping, latency deviation trend analysis is performed to identify high-latency requests and low-latency requests. High-latency requests are processed by sending access requests in advance, while low-latency requests are processed by adjusting the access timing.

8. A multi-channel parallel access control device for an HBF chip, characterized in that, The method for executing the HBF chip multi-channel parallel access control method as described in claim 1 includes: The access change analysis module is used to collect real-time access data streams from each memory segment of the HBF chip, perform access change analysis, and obtain dynamic access change patterns. The access situation analysis module is used to analyze the access competition situation based on dynamic access change patterns and mark high-conflict access request events. The bandwidth allocation module is used to obtain the access load status parameters of each channel and construct a bandwidth reallocation strategy. The channel distribution module is used to perform a feasibility assessment on high-conflict access request events based on the bandwidth reallocation strategy, determine whether the latency cost of the candidate channel is lower than the cost of conflict resolution, and if so, distribute the high-conflict access request events to the candidate channel.