Chip data processing method and system
By extracting data packet processing attributes and constructing efficient processing channels, scheduling conflict events are proactively generated, solving the problems of low resource utilization and unbalanced load in traditional chip processing, thereby optimizing chip processing performance and improving stability.
Patent Information
- Application Number
- CN202511751763.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-03-06
AI Technical Summary
Traditional chip processing architectures struggle to adapt to dynamically changing processing demands, resulting in low resource utilization, low processing efficiency, and a lack of proactive prediction and intelligent adjustment capabilities. They are also unable to effectively address performance inconsistencies and load imbalances among hardware components.
By extracting packet processing attributes and constructing efficient processing channels, scheduling conflict events are proactively generated. A dynamic load field is formed using reverse scheduling gradients. Combined with adaptive scheduling parameters, bottleneck bypass networks, and fault tolerance enhancement technologies, comprehensive optimization of chip processing performance is achieved.
It improves chip resource utilization efficiency and processing stability, ensures reliability and consistency in complex working environments, and achieves load balancing and elimination of hardware differences.
Smart Images

Figure CN121614259A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of integrated circuit optimization technology, and in particular to a chip data processing method and system. Background Technology
[0002] With the rapid development of artificial intelligence and high-performance computing applications, modern chips face increasingly complex data processing needs and diverse workload patterns. Traditional chip processing architectures mainly rely on static resource allocation and fixed scheduling strategies, which are difficult to adapt to dynamically changing processing requirements, resulting in low resource utilization, low processing efficiency, and failure to fully realize the computing potential of chips.
[0003] Most existing chip optimization techniques rely on passive responses to address performance bottlenecks and resource conflicts, lacking proactive prediction and intelligent adjustment capabilities, and struggling to effectively handle performance inconsistencies between hardware components. When faced with complex multi-task parallel processing scenarios, these methods often exhibit problems such as unbalanced loads, increased processing latency, and elevated error rates, severely limiting the overall performance of the chip system. Therefore, a method is urgently needed to address at least one of these issues. Summary of the Invention
[0004] This invention provides a chip data processing method and system. By extracting data packet processing attributes and constructing efficient processing channels, it actively generates scheduling conflict events and uses reverse scheduling gradients to form a dynamic load field. Combined with adaptive scheduling parameters, bottleneck bypass networks, and fault tolerance enhancement technologies, it achieves comprehensive optimization of chip processing performance, solves key technical problems such as low resource utilization, unbalanced load, and the impact of hardware differences in traditional chip processing, and provides an intelligent processing solution for high-performance chip computing.
[0005] The first aspect of this invention provides a chip data processing method, comprising the following steps:
[0006] The system acquires chip input data streams and hardware configuration information, performs structured preprocessing on the input data streams and hardware configuration information to extract data packet processing attributes;
[0007] Based on packet processing attributes, hidden dependencies are removed to form an efficient processing channel, and resource scheduling strategies are determined along the efficient processing channel.
[0008] According to the resource scheduling strategy, scheduling conflict events are actively generated. Based on the scheduling conflict events, a reverse scheduling gradient is constructed to form a dynamic load field. The load imbalance in the dynamic load field is used to create a processing window. The processing window is coupled with the standard scheduling algorithm to generate adaptive scheduling parameters.
[0009] Based on adaptive scheduling parameters, the internal state changes of the chip are triggered to generate a state monitoring map. Bottleneck computing units are identified from the state monitoring map to form a bottleneck bypass network. The bottleneck bypass network is then remapped to the processing window to form a load-balanced processing data flow.
[0010] Response characteristic data is obtained based on the load-balanced processing data stream. Hardware inconsistency is transformed based on the response characteristic data to form a fault tolerance enhancement factor. The fault tolerance enhancement factor is integrated into the load-balanced processing data stream to generate a pre-compensation processing sequence. The pre-compensation processing sequence eliminates the impact of hardware differences and generates a standardized processing signal.
[0011] A second aspect of the present invention provides a chip data processing system, comprising:
[0012] The data acquisition module is used to acquire the chip input data stream and hardware configuration information, and to perform structured preprocessing on the input data stream and hardware configuration information to extract data packet processing attributes.
[0013] The dependency resolution module is used to remove hidden dependencies based on packet processing attributes to form an efficient processing channel, and to determine resource scheduling strategies along the efficient processing channel.
[0014] The conflict scheduling module is used to actively generate scheduling conflict events according to the resource scheduling strategy, construct a reverse scheduling gradient based on the scheduling conflict events to form a dynamic load field, create a processing window by utilizing the load imbalance in the dynamic load field, and couple the processing window with the standard scheduling algorithm to generate adaptive scheduling parameters.
[0015] The bypass control module is used to trigger changes in the internal state of the chip based on adaptive scheduling parameters to generate a state monitoring map. The bottleneck computing unit is identified from the state monitoring map to form a bottleneck bypass network. The bottleneck bypass network is then remapped to the processing window to form a load-balanced processing data flow.
[0016] The fault tolerance enhancement module is used to obtain response characteristic data based on the load-balanced processing data stream, transform hardware inconsistencies based on the response characteristic data to form a fault tolerance enhancement factor, integrate the fault tolerance enhancement factor into the load-balanced processing data stream to generate a pre-compensation processing sequence, and eliminate the impact of hardware differences through the pre-compensation processing sequence to generate a standardized processing signal.
[0017] The beneficial effects of this invention are reflected in the following points: First, by using packet processing attribute extraction technology, a physical layer processing baseline configuration and a virtual processing layer mapping are constructed, forming cross-layer data features. Combined with serial path region identification and parallel propagation chain reconstruction, the hidden dependencies between packets are effectively eliminated, realizing the construction of efficient parallel processing channels and improving the parallelism and throughput of data processing. Second, by adopting a strategy of actively generating scheduling conflict events, a dynamic load field is constructed through reverse scheduling gradients. Combined with bottleneck bypass networks and adaptive scheduling parameter generation technology, traditional passive performance optimization is transformed into active intelligent scheduling, realizing dynamic balanced allocation of processing resources and preventive resolution of bottleneck problems, significantly improving the chip's resource utilization efficiency and processing stability. Finally, a hardware inconsistency conversion and fault tolerance enhancement mechanism is established. Through response characteristic data analysis, error signal source capture, and correction signal generation technology, combined with pre-compensation processing sequences and residual reallocation algorithms, the impact of performance differences between hardware components is effectively eliminated, generating standardized processing signal outputs and ensuring the reliability and consistency of the chip system in complex working environments.
[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0019] The accompanying drawings illustrate specific examples of the technical solutions of the present invention and, together with the detailed embodiments, form part of the specification, serving to explain the technical solutions, principles, and effects of the present invention.
[0020] Unless otherwise specified or otherwise, the same reference numerals in different figures represent the same or similar technical features, and different reference numerals may be used to represent the same or similar technical features.
[0021] Figure 1 This is a schematic flowchart of a chip data processing method according to the present invention.
[0022] Figure 2 This is a structural block diagram of a chip data processing system according to the present invention. Detailed Implementation
[0023] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0024] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0025] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0026] The technical solutions of the embodiments of this application will be described below.
[0027] like Figure 1 As shown, this embodiment of the invention provides a chip data processing method, including the following steps S110-S150:
[0028] Step S110: Obtain chip input data stream and hardware configuration information, and perform structured preprocessing on the input data stream and hardware configuration information to extract data packet processing attributes.
[0029] Specifically, the process involves acquiring the chip's input data stream and hardware configuration information. Real-time data stream information is collected through the chip's input interface, including key parameters such as packet size, transmission rate, data type, and timestamps. The bandwidth utilization of the input data stream is monitored, and the amount of data transmitted per second is calculated. The burst characteristics of the data stream are analyzed, and the ratio of peak to average data transmission rate is calculated. Priority identifiers for the data stream are extracted to distinguish between high-priority and low-priority data transmission needs. The chip's hardware configuration information is acquired, including specifications such as the number of processor cores, cache capacity, memory bandwidth, and clock frequency. The chip's architecture configuration is read, identifying the number of pipeline stages, number of execution units, and instruction set support. The chip's power management status is detected, obtaining the current operating voltage, power consumption level, and temperature status. The chip's bus configuration information is collected, including data bus width, address space allocation, and peripheral interface types. For example, a certain AI inference chip might have an input data stream of 32MB / s image data, a hardware configuration of an 8-core processor, 16MB L3 cache, DDR4-3200 memory interface, and an operating frequency of 2.4GHz. The acquired input data stream and hardware configuration information are formatted to unify the data representation format and units of measurement.
[0030] In some embodiments, the structured preprocessing of the input data stream and hardware configuration information to extract packet processing attributes includes: collecting physical layer processing baseline configurations using the input data stream and hardware configuration information; creating virtual processing layer mapping operations using the physical layer processing baseline configurations to form cross-layer data features; performing hierarchical matching parsing on the cross-layer data features to obtain processing attribute parameters; and extracting packet processing attributes based on the processing attribute parameters.
[0031] The system collects physical layer processing baseline configurations using input data streams and hardware configuration information. Based on data type characteristics, it selects appropriate execution unit configurations, prioritizing vector processing units for image data and scalar processing units for text data. It determines caching strategies based on data stream timestamp information, using larger cache lines for sequential access data and smaller cache lines for random access data. It calculates memory bandwidth requirements using data stream bandwidth utilization and transmission rate to determine the necessary DDR4-3200 memory interface configuration. It analyzes the latency requirements of high-priority data and sets corresponding processing priorities and scheduling strategies. It adjusts processing frequency and voltage settings based on hardware power consumption limitations to find a balance between performance and power consumption. It collects key configuration parameters for the physical layer, such as baseline clock frequency, core voltage, cache allocation, and bus bandwidth. It performs feasibility checks on the baseline configurations to ensure they are within hardware capabilities. Finally, it compiles the collected physical layer processing baseline configurations into a configuration parameter table, containing complete information on processor configuration, storage configuration, and interface configuration. For example, for an image data stream of 32MB / s, the collected physical layer baseline configuration includes: 8 processing cores operating at 2.4GHz, 16MB L3 cache allocated in a 4:8:4 ratio (instruction cache: data cache: shared cache), DDR4-3200 memory bandwidth of 25.6GB / s, core voltage of 1.2V, and vector processing unit priority set to high.
[0032] A virtual processing layer mapping operation is created using the physical layer processing baseline configuration to form cross-layer data characteristics. On-chip cache access patterns are detected during data transmission intervals, cache access monitoring points are set, and access counters are deployed in L1, L2, and L3 caches. Cache hit rate data is collected, and the ratio of the number of hits to the total number of accesses for each cache level is calculated. Cache access patterns for high-priority and low-priority data are statistically analyzed based on data flow priority identifiers. Cache access latency is monitored, and hardware timestamps are used to record the response time of access requests. Based on the detected cache access patterns, virtual layers are divided: a high-performance virtual layer is formed with a cache hit rate >95% and access latency <50 nanoseconds; a standard virtual layer is formed with a hit rate of 70-95% and latency of 50-100 nanoseconds; and a low-power virtual layer is formed with a hit rate <70% and latency >100 nanoseconds. Combined with the physical layer's baseline frequency configuration, the high-performance virtual layer operates at a maximum frequency of 2.4GHz, the standard virtual layer operates at a medium frequency of 1.8GHz, and the low-power virtual layer operates at a reduced frequency of 1.2GHz. Data channels are partitioned based on cache bandwidth utilization. Regions with bandwidth utilization B > 80% are created as fast data channels, while regions with B < 50% are created as slow channels. This virtual mapping, based on physical layer configuration and dynamic detection, forms cross-layer data characteristics, including characteristic parameters such as the processing capacity distribution of each virtual layer, cache performance indicators, and frequency configuration.
[0033] Layered matching parsing is performed on cross-layer data features to obtain processing attribute parameters. Based on cross-layer data features, layered matching parsing is performed to extract cache performance indicators and load distribution characteristics for analysis. The matching degree of each virtual layer is detected using cache hit rate differences; when the hit rate of a high-performance layer drops below 90%, it is identified as a matching anomaly. Inter-layer coordination is evaluated using access latency indicators; matching adjustments are triggered when latency exceeds a preset range by 50%. The communication overhead matching between virtual layers is detected based on cache bandwidth utilization; communication bottlenecks are identified when the bandwidth utilization difference is >30%. The load distribution characteristics of each virtual layer are analyzed, and the number of data packets processed and the proportion of time occupied by each layer are statistically analyzed. Performance imbalance areas in virtual layer matching parsing are identified, and these areas require attribute adjustments. Key processing attribute parameters for each virtual layer are extracted, including processing capacity coefficients, latency sensitivity, bandwidth requirements, and load tolerance. For overloaded virtual layers, the processing capacity coefficient is set to a higher value, indicating that the layer needs more resource support. For lightly loaded virtual layers, the processing capacity coefficient is set to a lower value, indicating that the layer can release resources. The obtained processing attribute parameters are associated with and stored with the corresponding virtual layer identifiers to form a complete attribute parameter mapping table.
[0034] Data packet processing attributes are extracted based on processing attribute parameters (including processing capacity coefficient, latency sensitivity, bandwidth requirement, and load tolerance). Based on the processing capacity coefficient and load tolerance of each virtual layer, the data packet allocation strategy and processing requirements are determined. For virtual layers with high latency sensitivity, data packets with high real-time requirements are allocated; for virtual layers with high bandwidth requirements, large data packets are allocated. Arriving data packets are matched with the processing attribute parameters of the virtual layers according to their basic attributes (data packet size, data type, priority) to determine the optimal processing layer. The processing layer attributes of the data packets are extracted, identifying whether the data packets should be allocated to the high-performance layer, standard layer, or low-power layer for processing. The processing requirements of the data packets in the corresponding virtual layer are analyzed, and the required computing resources, storage space, and bandwidth requirements are calculated based on the frequency configuration of the virtual layer. Based on the target layer of data packet allocation, the application characteristics of the data packets are identified, determining whether the data packets belong to application types such as image processing, audio processing, or general computing. The content characteristics of the data packets are analyzed, and the degree of data structure and complexity are identified through pattern matching. The temporal characteristics of the data packets are analyzed to extract the arrival time interval and estimated processing time. Historical processing information of statistical data packets is collected to analyze the processing time and resource consumption patterns of similar data packets. For example, after attribute parameter matching, the extracted data packet processing attributes of a 1MB image data packet include: target processing level is high-performance virtual layer, data type is RGB image, resolution is 1920×1080, compression format is JPEG, processing priority is high, estimated processing time is 12 milliseconds, required processing cores are 2, required cache is 2MB, and bandwidth requirement is 8MB / s.
[0035] Step S120: Based on the packet processing attributes, remove hidden dependencies to form an efficient processing channel, and determine the resource scheduling strategy along the efficient processing channel.
[0036] In some embodiments, the efficient processing channel is formed by removing hidden dependencies based on packet processing attributes, including: identifying serial path regions using packet processing attributes; performing dependency reconstruction tracing through the serial path regions to form a parallelized propagation chain; performing path reduction positioning on the parallelized propagation chain to obtain reduced control points; and forming an efficient processing channel based on the reduced control points.
[0037] Identify serial path regions using packet processing attributes. Analyze timestamp information in packet attributes to identify packet sequences with strict timing requirements. Examine packet dependency fields to find combinations of packets explicitly marked as dependent. Identify upstream and downstream relationships in the data flow based on packet data source and destination attributes. Use attribute matching algorithms to categorize packets with similar processing needs into potential serial processing paths. Analyze packet resource requirement attributes to identify packet sequences that require exclusive use of certain resources. Determine the preemption relationship between high-priority and low-priority packets through packet priority attribute analysis. Construct the topology of the serial path and represent the serial dependencies between packets using a directed graph. Statistically assess the impact of serial processing by calculating the length and complexity, path depth, and number of nodes of the serial path. Analyze the distribution density of serial path regions to identify regions with concentrated and dispersed serial dependencies. For example, in a video encoding task, frame sequence packets must be processed serially according to frame number order, forming a serial path region of length 30. Classify and label the identified serial path regions to distinguish different types of serial dependencies.
[0038] Dependency reconstruction tracing is performed within the serial path region to form a parallel propagation chain. The true dependency degree of each data packet in the serial path is detected, and strong and weak dependencies are distinguished through data flow monitoring and timing constraint detection. For data packets with weak dependencies, dependency decoupling techniques are used to separate them from the serial path. Data replication and buffering techniques are employed to create independent processing paths for data packets that can be processed in parallel. The reconstructed dependencies form a branching propagation chain structure, decomposing the original serial path into multiple parallel sub-paths. Necessary synchronization points are preserved in the parallel propagation chain to ensure data consistency and processing correctness. The transformation path of each data packet from the serial path to the parallel chain is recorded. The branching factor of the reconstructed propagation chain is analyzed, and the average number of parallel branches that can be decomposed from each serial path is calculated. For example, 10 image filtering data packets that originally needed to be processed serially are decomposed into 3 parallel propagation chains after detection: the first chain contains 4 data packets, the second chain contains 3 data packets, and the third chain contains 3 data packets. Each parallel chain handles data exchange and state synchronization between chains through a data exchange interface.
[0039] For example, the path reduction positioning of the parallel propagation chain to obtain the reduction control point includes: using the parallel propagation chain to identify disconnected dependent chain segments; determining the recycling direction by disconnecting dependent chain segments; identifying the reconstructed path network based on the recycling direction; and converting the reconstructed path network into parallel processing path accumulation to form the reduction control point.
[0040] The parallel propagation chain is used to identify disconnected dependency segments. In the parallel propagation chain, dependency analysis identifies weak dependency segments that can be disconnected, further improving the degree of parallelization. Data traffic between data packets is monitored; when the amount of data transmitted from packet A to packet B exceeds 1MB, it is recorded as a strong data dependency. Timing constraints are identified through timestamp comparison; when packet B must begin processing within 10ms after packet A completes, it is recorded as a strong timing dependency. Conflicts in processing core and cache usage are detected; when two data packets require exclusive access to the same processing core, it is recorded as a strong resource dependency. A dependency strength threshold of 0.3 is set; when the overall dependency coefficient is below this threshold, it is determined to be a disconnectable weak dependency segment. The impact of disconnecting weak dependency segments is analyzed, assessing the degree of impact on processing correctness and operational effectiveness. Dependency substitution techniques are employed, replacing disconnected dependencies through data duplication or delayed processing. For example, in the parallel propagation chain of image processing, the dependency strength between the filtering operation and the subsequent compression operation is found to be only 0.25; this dependency segment can be disconnected through an intermediate cache.
[0041] The recycling direction is determined by disconnecting dependent chain segments. Monitoring and analysis are performed on the processing resources released by disconnected chain segments, including processing cores, cache space, and bandwidth resources. Time window analysis is used to calculate the available time windows for various resources, determining the time periods and durations of resource idleness. Demand analysis is used to identify peak resource demand in other propagation chains, pinpointing time periods and chains with insufficient resources. Resource demand time analysis determines the directionality of recycling; when chain segment A releases resources at time T1 and chain segment B needs the same resources at time T2, the recycling direction from A to B is determined. Prioritized recycling directions are determined based on resource type matching, with processing core resources preferentially flowing to compute-intensive chains and cache resources preferentially flowing to data-intensive chains. For example, a disconnected image filtering chain releases 2 processing cores and 4MB of cache, which are allocated to the audio processing chain within a 0.5ms time window. A resource flow direction map is constructed, with each direction representing a feasible recycling path. The determined recycling directions are organized into a direction configuration table, containing information such as resource type, release time, target chain segment, and transmission path.
[0042] The reorganized path network is identified based on the direction of resource recycling. The execution sequence and path connections of the propagation chain are redesigned according to the time schedule of resource recycling. Propagation chain segments that can share resources are staggered in time to avoid resource conflicts. New path connection relationships are formed, connecting the originally independent propagation chains through resource-sharing nodes. A network topology optimization algorithm is used to adjust the structure of the reorganized path network and reduce data transmission latency. Critical paths and bottleneck nodes in the reorganized path network are identified, and limiting factors affecting the overall processing time are identified. Path merging and splitting points in the reorganized network are identified, and key connection positions in the network topology are determined. The connectivity of each node in the reorganized network is statistically analyzed, and the number of input and output connections for each node is calculated through topology scanning. Path redundancy in the network is detected, and the number of independent paths between any two nodes is calculated. Connectivity testing tools are used to verify the reachability of the network, ensuring that any source node can reach any target node. Load analysis is performed to analyze the load distribution characteristics of the reorganized network and evaluate the load balancing degree of each path. Complexity analysis is performed on the reorganized path network, and the complexity of the network is evaluated through the number of nodes and connection density. When a shared node fails, the system can quickly switch to a backup path to maintain stable operation.
[0043] The reconstructed path network is transformed into a parallel processing path accumulation to form reduction control points. The distribution of parallel processing capabilities in the reconstructed network is identified, and the network node with the highest parallelism is identified. The number of processing paths connected to each node and the corresponding number of processing cores are counted, and the node's processing capability index is calculated using a weighted formula: P = α×N_path + β×N_core, where P is the processing capability index, N_path is the number of processing paths connected to the node, N_core is the number of processing cores equipped in the node, α is the path weight coefficient, and β is the core weight coefficient. For example, if α=2 and β=1, then the processing capability index of node A (3 paths, 4 cores) is 10; the processing capability index of node B (2 paths, 6 cores) is also 10. The contribution of each node to the overall processing capability is determined by the accumulation and proportion of the processing capability index. The node with the highest cumulative value is selected as the main reduction control point; these nodes have the strongest processing capability concentration effect. Auxiliary control points are set between the main reduction control points to ensure coordination and smooth data flow between control points. The accumulation pattern of parallel processing paths is identified, and different accumulation types such as linear accumulation and non-linear accumulation are identified. A mapping relationship between parallel processing paths and control points is established, and the control point number and weight coefficient corresponding to each path are determined. For example, in the reorganized path network, three main reduction control points and two auxiliary control points are identified, with the cumulative processing capacities of the main control points being 45%, 30%, and 25%, respectively.
[0044] Efficient processing channels are formed based on reduced control points. The topology and data flow of the processing channels are designed with the reduced control points as core nodes. Efficient data transmission paths and processing node connections are constructed based on the cumulative processing capacity distribution of the reduced control points. Corresponding processing resources, including processing cores, cache space, and I / O bandwidth, are configured at each reduced control point. The main reduced control point is configured with 4-6 processing cores, and the auxiliary control points with 2-3 cores, with dedicated cache space allocated according to processing capacity ratios. Each reduced control point ensures synchronized processing progress through a coordination interface, and a status signal exchange mechanism maintains coordination and consistency between channels. A load balancing strategy for the processing channels is designed, and the processing load of each channel is monitored and adjusted in real time through the reduced control points. Different channels can borrow each other's processing resources through resource interfaces when needed, achieving dynamic resource sharing and optimized configuration. A hierarchical structure of processing channels is formed, with higher-priority channels receiving more control point support and resource allocation to ensure the processing priority of critical tasks. The throughput capacity of the processing channels is analyzed, and the data processing rate and response time performance of each channel are evaluated.
[0045] Resource scheduling strategies are determined along efficient processing channels. Based on efficient processing channels formed by reducing control points, the number of data packets and total processing time for each channel are statistically analyzed, and the average and peak loads of the channels are calculated. CPU utilization, memory usage, and task queue lengths for each channel are collected in real time. The standard deviation σ and average load μ of each channel's load are calculated, and load imbalance is quantified using the coefficient of variation L = σ / μ. For example, if the loads of three channels are 90%, 60%, and 30% respectively, the average load μ = 60%, the standard deviation σ = 24.5%, and the imbalance L = 0.41. When L > 0.3, it is considered a significant load imbalance. Bottleneck detection is used to identify the bottleneck locations of processing channels, pinpointing the key links limiting overall processing capacity. Temporal characteristic analysis identifies the periodic and bursty characteristics of load changes over time, and spatial characteristic analysis determines the distribution patterns of high-load and low-load regions. For example, analysis reveals that the image processing channel accounts for 60% of the total load, while the audio processing channel accounts for only 20%, showing a significant load imbalance. Based on the analyzed load distribution characteristics, a resource scheduling strategy is determined, allocating processing resources according to the load level of each processing channel, with higher-load channels receiving more processing cores and cache space. A load balancing strategy is adopted to migrate some data packets from overloaded channels to less-loaded channels for processing. An adaptive scheduling algorithm is designed to automatically adjust the resource allocation ratio based on real-time load changes. For uneven load distribution, a resource borrowing strategy is adopted, allowing busy channels to temporarily borrow resources from idle channels. A priority scheduling strategy is constructed to arrange the processing order according to the priority of data packets and processing deadlines. A load threshold trigger condition is set to activate the load splitting strategy when the channel load exceeds 80%. A time window strategy for resource allocation is formulated, using different resource allocation ratios for different time periods. For example, when the image processing channel is overloaded, some image data packets are diverted to the general computing channel, and the number of processing cores in the image processing channel is increased from 4 to 6. The scheduling effect is evaluated through processing latency monitoring and throughput measurement to ensure the effectiveness of the resource scheduling strategy.
[0046] Step S130: Actively generate scheduling conflict events according to the resource scheduling strategy, construct a reverse scheduling gradient based on the scheduling conflict events to form a dynamic load field, create a processing window using the load imbalance in the dynamic load field, and couple the processing window with the standard scheduling algorithm to generate adaptive scheduling parameters.
[0047] Specifically, scheduling conflicts are proactively created according to resource scheduling strategies. Based on load balancing strategies, additional data packets are intentionally added to the medium-load channel in a processing channel with an original load distribution of 90%, 60%, and 30%, creating load redistribution conflicts. Based on resource borrowing strategies, multiple channels simultaneously request to borrow the same processing core resources, creating resource contention scenarios. Targeting the 80% load threshold trigger condition, the load on a certain channel is intentionally pushed to 85% to test the response of the load balancing strategy. Using adaptive scheduling algorithms, packet priorities are suddenly changed when the algorithm adjusts resource allocation, creating scheduling decision conflicts. Priority scheduling strategies are adopted, allowing high-priority tasks to compete with low-priority tasks for the same processing time window, triggering timing conflicts. Load injection controls the intensity and duration of conflicts to ensure they remain within the system's tolerance. For example, based on a 4-core configuration for an image processing channel, a load requiring 6 cores is intentionally added, creating a 2-core processing gap conflict. Conflict events are categorized into resource conflicts, timing conflicts, and priority conflicts.
[0048] In some embodiments, constructing a reverse scheduling gradient based on scheduling conflict events to form a dynamic load field includes: performing conflict intensity analysis on scheduling conflict events to obtain conflict characteristic parameters; using the conflict characteristic parameters to perform system stability testing to form a resilience assessment configuration; constructing a conflict response mapping through the resilience assessment configuration to obtain response switching nodes; and constructing a reverse scheduling gradient based on the response switching nodes to form a dynamic load field.
[0049] Perform conflict intensity analysis on scheduling conflict events to obtain conflict characteristic parameters. Statistically count the number of processing channels and data packets affected by the conflict; record a channel as affected when its load exceeds a preset threshold of 15%. Measure the time from the start of the conflict to the system's recovery. Monitor the percentage decrease in throughput and the increase in response latency in milliseconds caused by the conflict. Analyze the propagation time of the conflict from its source to other components. Track and identify the main data packet types and resource types causing the conflict. Statistically count the number of conflicts per unit time and the average interval. Calculate the conflict intensity coefficient I = A × T × F, where A is the number of affected channels, T is the duration, and F is the frequency. Mark the distribution location of the conflict in the processing area. For example, a resource contention conflict affecting 3 channels, lasting 15 milliseconds, and occurring 2 times / second has an intensity coefficient I = 90. Organize the characteristic parameters into a parameter vector containing intensity, range, time, and frequency.
[0050] Use the conflict characteristic parameters to conduct system stability tests to form a resilience evaluation configuration. Set an increasing test sequence according to the conflict intensity. Gradually increase the conflict intensity and observe the changes in system throughput and response latency. Test the stability of multiple conflict superpositions, triggering resource conflicts and timing conflicts simultaneously. Measure the number of processing cycles required to recover from the conflict state to the normal state. Calculate the system resilience coefficient R = Pnormal / (Pconflict × Trecover), where Pnormal is the system processing capacity in the normal state, Pconflict is the system processing capacity during the conflict, and Trecover is the recovery time coefficient. Measure the performance degradation caused by different conflict types respectively. Test the time for the system to automatically recover to 90% of the normal performance without manual intervention. Determine the maximum conflict intensity and frequency that the system can withstand. Configure the resilience parameters according to the test results, including buffer capacity, recovery speed, and self-healing threshold. Organize the evaluation results into a resilience evaluation configuration file containing the response parameters for each conflict scenario.
[0051] Obtain the response switching nodes by constructing a conflict response mapping through the resilience evaluation configuration. Use the resilience coefficient R, the recovery time coefficient Trecover, and the self-healing threshold in the resilience evaluation configuration to construct the system conflict response mapping relationship. Analyze the key parameters in the resilience evaluation configuration and use the resilience coefficient R as the main basis for judging the response intensity. Create a response strategy matrix according to the results of the resilience evaluation configuration: when R > 0.8, adopt a mild response strategy; when 0.5 < R ≤ 0.8, adopt a moderate response strategy; when R ≤ 0.5, adopt a severe response strategy. Configure corresponding response strategies for each resilience level: the mild response adopts a buffering strategy, the moderate response adopts a reallocation strategy, and the severe response adopts a transfer strategy. Establish a response decision tree based on the resilience evaluation configuration parameters. Identify the response switching nodes according to the threshold changes of the resilience coefficient R: set the switching node A when the system resilience crosses from the high-resilience state to the medium-resilience state, and set the switching node B when it crosses from the medium-resilience state to the low-resilience state. Set the node trigger conditions based on the resilience evaluation configuration: node A is triggered when R drops below 0.8, and node B is triggered when R drops below 0.5. Measure the time interval for the response strategy to switch. For example, when the resilience coefficient drops from R = 0.85 to R = 0.75, node A is triggered, and the system response switches from the mild buffering strategy to the moderate reallocation strategy. Set the processing priority when multiple nodes are triggered simultaneously.
[0052] A dynamic load field is formed by constructing a reverse scheduling gradient based on response switching nodes. The physical location of the response switching nodes is used as the starting point for gradient construction. Starting from the node, a reverse scheduling gradient vector is constructed along the direction of decreasing load pressure. The gradient magnitude is calculated as (node load pressure - target area pressure) / distance. A gradient field covering the processing area is formed by synthesizing gradient vectors from multiple nodes. Real-time load density is superimposed on the gradient field to construct a dynamic load field. The load field density changes in real time with conflict events, with an update frequency of every 10 milliseconds. The optimal path for load transfer is determined, with streamlines pointing in the direction of gradient descent. The potential energy distribution of the load field is analyzed; high potential energy indicates high load pressure. An evolutionary model of the temporal variation of the load field is established. For example, a reverse gradient field centered on two switching nodes has an average magnitude of 0.8, with the main trend shifting from the 90% load area to the 30% load area. The load field is divided into three regions: high load, medium load, and low load.
[0053] Processing windows are created based on load imbalances in a dynamic load field. By monitoring the distribution of the dynamic load field in real time and calculating the load variance of each region, a sliding window method is used to analyze load fluctuations. Regions with a standard deviation σ > 0.2 are identified as having load imbalances. Buffer processing windows are created where load gradients change drastically. When the load difference between adjacent regions exceeds 40%, a gradient boundary is established, and buffer windows are deployed in the boundary region to smooth load transitions. Dense processing windows are created in high-load regions, configured with 1.5 times the standard number of processing cores to enhance the data processing capacity of these regions. Sparse processing windows are created in low-load regions to receive and process overflow tasks diverted from high-load regions. The size of the processing window is determined based on the magnitude of the load variance, using the empirical formula: Window size = Load variance × Adjustment coefficient. The adjustment coefficient is dynamically set according to the overall system load level to match the window size with the degree of load imbalance. The time-based operating characteristics of the processing windows are designed: the opening time is synchronized with load imbalance detection, the duration is adjusted according to the load balancing speed, and the closing condition is set when the load difference drops below a threshold. Resource contention and functional duplication issues between different processing windows are analyzed and avoided to ensure coordinated operation of each window.
[0054] Adaptive scheduling parameters are generated by coupling the processing window with standard scheduling algorithms. The parameter interfaces of standard scheduling algorithms (round-robin, shortest job first, priority scheduling) are identified, including adjustable parameters such as time slice, priority weight, and queue length. The processing capacity of the processing window is converted into scheduling weight parameters, and the weight coefficient is obtained by the ratio of the window's processing capacity to the total system processing capacity. The time slice allocation is adjusted according to the activation state of the processing window; when a dense processing window is active, the time slice is extended from 10ms to 15ms, and shortened to 5ms when a sparse processing window is active. The load prediction parameters of the scheduling algorithm are corrected using real-time load information of the processing window; when the window load deviation exceeds 20%, the prediction value is corrected accordingly. A mapping function between processing window parameters and scheduling parameters is established, defining a linear mapping relationship f(x) = ax + b, where x is the window parameter, a is the mapping coefficient, and b is the baseline offset. The influence of the window on the scheduling algorithm is evaluated, and the degree of influence is quantified by the proportion of the window weight in the total scheduling weight. The core scheduling logic of the scheduling algorithm remains unchanged; only the values and weight allocation are adjusted through the parameter interfaces. Generate an adaptive parameter set including time slice weights, priority adjustment coefficients, and queue capacity coefficients. For example, when a dense processing window is activated, the resource allocation weight for round-robin scheduling increases from 0.6 to 0.8, and the high-priority weight for priority scheduling is adjusted from 1.5 to 1.8. Set limits on the parameter variation range, limiting the weight coefficients to between 0.3 and 1.2, and the time slices to between 5 and 30 ms, ensuring that parameter adjustments are within the algorithm's tolerance range.
[0055] Step S140: Based on the adaptive scheduling parameters, the internal state changes of the chip are triggered to generate a state monitoring map. Bottleneck computing units are identified from the state monitoring map to form a bottleneck bypass network. The bottleneck bypass network is remapped to the processing window to form a load-balanced processing data flow.
[0056] Specifically, a state monitoring map is generated based on adaptive scheduling parameters that trigger internal chip state changes. The time slice weight is adjusted from 0.6 to 0.8 using these parameters, triggering changes in the time allocation of each processing core. This weight increase directly affects the core's workload and power consumption. The priority order of the task queue is reassigned by increasing the priority adjustment coefficient from 1.5 to 1.8, affecting the task execution order and response latency. The capacity allocation of L1, L2, and L3 cache queues is altered by adjusting the queue capacity coefficient, impacting data access hit rate and latency. The resource allocation ratio of the memory interface and bus is adjusted based on changes in resource allocation weights, affecting data transfer speed and bandwidth utilization. State change information for each computing unit after parameter adjustments is collected, including workload, resource utilization, and response time. Performance counters are used to collect real-time state data from key units such as the processor core, cache controller, memory controller, and I / O controller. Time series analysis is used to identify the temporal characteristics of state changes, revealing trends and periodic patterns. A visual design is used to construct a state monitoring map, with color intensity representing state strength and arrows representing the direction and magnitude of state changes.
[0057] In some embodiments, identifying bottleneck computing units from a state monitoring map to form a bottleneck bypass network includes: identifying bottleneck computing units using the state monitoring map; converting bottleneck computing units into filter trigger thresholds; performing data quality screening based on the filter trigger thresholds to obtain filtered data features; and performing bottleneck reverse utilization on the filtered data features to form a bottleneck bypass network.
[0058] Identify bottleneck computing units using a condition monitoring map. Scan all computing units in the condition monitoring map for key indicators such as load saturation, response latency, and resource utilization. Calculate the performance deviation of each computing unit: Deviation = (Actual Performance - Standard Performance) / Standard Performance × 100%. Set bottleneck identification criteria: a computing unit is identified as a bottleneck unit when its load saturation exceeds 80%, its response latency exceeds twice the baseline latency, and its resource utilization exceeds 90%. Analyze the state change trends of computing units to identify units with continuously declining or drastically fluctuating performance. Statistically analyze the frequency of bottleneck computing units; units with frequent bottlenecks are marked as critical bottleneck units. Detect the cascading effects of bottlenecks on downstream computing units. The bottleneck severity index S is calculated as (L-80)×(D / baseline latency-1)×(U-90), where S is the severity index (dimensionless), L is the load saturation (expressed as a decimal, e.g., 0.92 represents 92%), D is the actual response latency, the baseline latency is the standard response latency, and U is the resource utilization (expressed as a decimal, e.g., 0.95 represents 95%). A bottleneck unit is identified when S>0. The spatial distribution characteristics of the bottleneck computing units are analyzed. Identified bottleneck units are categorized into processing bottlenecks, storage bottlenecks, and communication bottlenecks. For example, the status monitoring map identifies processing core 2 (92% load, 120ms latency) and memory controller 1 (95% utilization, 80ms latency) as the main bottleneck computing units. The characteristic information of the bottleneck computing units is recorded, including bottleneck type, severity, impact range, and duration.
[0059] The bottleneck computing unit is converted into a filtering trigger threshold. Based on the response latency characteristics of the bottleneck unit, a time-triggered threshold for data processing, T_threshold, is set to = bottleneck latency × 0.8, where T_threshold is the time-triggered threshold and bottleneck latency is the current response latency of the bottleneck unit. Using the resource utilization information of the bottleneck unit, a resource allocation trigger threshold, R_threshold, is determined to be = bottleneck utilization rate × 0.9, where R_threshold is the resource trigger threshold and bottleneck utilization rate is the current resource utilization rate of the bottleneck unit. The numerical range of the filtering trigger threshold is calculated to ensure that the threshold is within a reasonable operating range. A multi-level threshold system is established, with the early warning threshold set at 70% of the bottleneck indicator, the trigger threshold at 85%, and the emergency threshold at 95%. Different conversion methods are applied based on the bottleneck type: For processing bottlenecks, a linear conversion is used: T_new = T_old × k, where T_new is the new threshold, T_old is the original threshold, and k is the linear conversion coefficient; for storage bottlenecks, a logarithmic conversion is used: T_new = T_old × log(utilization), where utilization is the storage unit utilization rate; and for communication bottlenecks, an exponential conversion is used: T_new = T_old × e^(load coefficient), where the load coefficient is the communication load coefficient. The effective and ineffective times of the thresholds are set through time characteristic analysis. Spatial impact analysis is used to determine the computing unit coverage area corresponding to each threshold. For example, 92% load saturation of processing core 2 is converted into a data traffic filtering threshold of 1200 packets / second, and 95% utilization of memory controller 1 is converted into a memory access trigger threshold of 800 times / millisecond.
[0060] Data quality screening is performed based on filtering trigger thresholds to obtain the characteristics of the filtered data. Using the set filtering trigger thresholds, the input data stream is screened for quality, retaining high-quality data and filtering out low-quality data. A data quality assessment system is established, checking data packet integrity (packet loss rate <1%), timeliness (latency < threshold), accuracy (verification passed), and consistency (format matching) through integrity checks, accuracy checks, and consistency checks. The processing requirements of data packets are compared with the filtering trigger thresholds. When the processing requirements of a data packet exceed the trigger threshold, it is marked as high-load data. High-load data is degraded by reducing processing precision or delaying processing to reduce resource requirements. Timely data packets are filtered, prioritizing the retention of latency-sensitive data and caching latency-tolerant data. The distribution characteristics of the filtered data are analyzed, and the ratio of retained to filtered data is statistically analyzed. Quality indicators are calculated for the filtered data, including average processing time, resource requirements, and priority distribution. The pattern characteristics of the filtered data are identified, and the data type, size distribution, and processing complexity are analyzed. Cluster analysis is performed on the features of the filtered data to group data with similar characteristics into data categories.
[0061] Bottleneck reversal is implemented to form a bottleneck bypass network based on the characteristics of the filtered data. The processing requirements of the filtered data are analyzed, and lightweight data types that can utilize the idle time of bottleneck units are identified based on packet size, processing time, and resource requirements. Appropriate lightweight data processing tasks are scheduled during periods when the bottleneck unit load is below 60%, achieving bottleneck resource reversal. Bottleneck reversal refers to reusing computing units that were originally bottlenecks due to high load as auxiliary processing nodes during their idle periods when the load decreases, turning waste into treasure and improving the overall system processing capacity. The idle time periods of bottleneck units are determined through real-time load monitoring, and a bottleneck unit is marked as an available processing window when its utilization rate is below 60% for 5 consecutive seconds. The filtered data is divided into three levels according to processing complexity: low complexity data (processing time <10ms), medium complexity data (10-50ms), and high complexity data (>50ms), with low complexity data prioritized for processing during the idle time of bottleneck units. A time scheduling table for bottleneck reversal is constructed to record the available time periods and corresponding data processing arrangements for each bottleneck unit. The design incorporates a network topology with bypass paths, transforming the original bottleneck units into auxiliary processing nodes, thus forming a bottleneck bypass network. This network is a dynamic data transmission structure where, when the main processing path becomes overloaded, data can bypass these former bottleneck nodes for processing, thereby avoiding the current processing bottleneck and achieving intelligent load balancing.
[0062] The bottleneck bypass network is remapped to processing windows to form a load-balanced data flow. The processing window configuration is analyzed: dense processing windows are configured with 1.5 times the processing cores for high-load areas, sparse processing windows are used to receive overflow tasks, and buffer processing windows are placed in boundary areas with drastic load gradient changes. The three main paths of the bottleneck bypass network are mapped to processing windows: high-throughput bypass paths are mapped to dense processing windows, low-throughput paths to sparse processing windows, and intermediate transition paths to buffer processing windows. Allocation is based on the load characteristics of the bypass paths: high-load data paths are assigned to dense windows, overflow data paths to sparse windows, and data paths with fluctuating loads to buffer windows. A round-robin allocation strategy ensures a balanced distribution of bypass paths among windows; when the number of paths connected to a window exceeds its processing capacity, new paths are automatically assigned to the lightest-loaded window. The load balancing effect of each processing window is evaluated using the load balancing degree calculation L = (σ / μ) × 100%, where L is the load balancing degree (percentage), σ is the standard deviation of the load of each window, and μ is the average load. A load balance is considered achieved when L < 20%. The data flow execution order is constructed as follows: Data packets first activate the corresponding processing window based on the trigger condition of load imbalance > 0.3, and then enter the corresponding window for processing according to the detour path allocation rules. Data flow priority scheduling is set: high-load data packets enter the dense processing window first, overflow data packets enter the sparse processing window, and fluctuating data packets are buffered through the buffer window.
[0063] Step S150: Obtain response characteristic data based on the load-balanced processing data stream, perform hardware inconsistency transformation based on the response characteristic data to form a fault tolerance enhancement factor, integrate the fault tolerance enhancement factor into the load-balanced processing data stream to generate a pre-compensation processing sequence, and eliminate the impact of hardware differences through the pre-compensation processing sequence to generate a standardized processing signal.
[0064] Specifically, response characteristic data is obtained based on load-balanced data stream processing. The response time of the data stream in different processing windows is measured, recording the time span from data input to processing completion. Throughput data for each processing window is collected, and the number of data packets processed and the total data volume per unit time are statistically analyzed. The latency distribution characteristics of the data stream are analyzed, and statistical indicators such as average latency, maximum latency, and latency variance are calculated. The error rate and packet loss rate of the data stream during processing are detected to identify the reliability level of data processing. Resource consumption in each processing window is measured, including CPU utilization, memory usage, and power consumption data. Load fluctuation information of the data stream is collected, and the amplitude and frequency characteristics of load changes over time are analyzed. Performance differences of the data stream on different processing paths are analyzed, and changes in response characteristics between paths are identified. The concurrent processing capability of the data stream is measured, and the number of data packets processed simultaneously and the concurrency level are statistically analyzed. For example, the response time for a dense processing window processing image data is 12ms, while the sparse processing window requires 18ms to process the same data, showing a significant difference in processing capability. State transition information during data stream processing is recorded, including queue state, cache state, and processor state changes. The collected response characteristic data is organized into a multi-dimensional feature vector, which contains feature parameters in multiple dimensions such as time, throughput, latency, and error rate.
[0065] In some embodiments, the process of transforming hardware inconsistencies based on response characteristic data to form a fault tolerance enhancement factor includes: identifying hardware inconsistencies using response characteristic data; capturing error signal sources based on hardware inconsistencies; performing signal reversal processing through the error signal sources to generate a correction signal; and performing fault tolerance strength analysis based on the correction signal to generate a fault tolerance enhancement factor.
[0066] Identify hardware inconsistencies using response characteristic data. Analyze the response time distribution of each hardware component to identify the degree of performance inconsistency between components through distribution differences. Analyze the throughput variation characteristics of hardware components to identify inconsistency patterns in throughput capabilities. Detect the error rate distribution of hardware components to identify components with significant differences in error frequency. Statistically analyze the resource utilization differences of hardware components, calculating the coefficient of variation of resource utilization such as CPU, memory, and bandwidth. Identify the spatial distribution characteristics of hardware inconsistencies and analyze the distribution patterns of inconsistencies on the chip. For example, if the average response time of the four cores on the left side of the processing core array is 10ms, while that of the four cores on the right side is 15ms, it indicates a significant spatial inconsistency. Analyze the temporal characteristics of hardware inconsistencies to identify the trends and periodicity of inconsistencies over time. Calculate quantitative indicators of hardware inconsistency, using indicators such as coefficient of variation and relative standard deviation to quantify the degree of inconsistency. Classify the types of hardware inconsistencies, distinguishing inconsistencies caused by different factors such as manufacturing process differences, aging degree differences, and environmental impact differences.
[0067] Error signal sources are captured based on hardware inconsistencies. The correlation between hardware inconsistencies and system errors is analyzed to identify which inconsistencies directly lead to processing errors. Signal spectrum analysis is used to identify the characteristic frequencies and amplitude distribution of error signals in the frequency domain. The temporal characteristics of error signals are detected, and their occurrence time, duration, and repetition patterns are analyzed. Trigger conditions for error signal capture are set, initiating capture when inconsistency indicators exceed a set threshold. Multi-channel signal capture technology is employed to simultaneously collect error signal outputs from multiple hardware components. The propagation path of error signals is analyzed, tracing their propagation from source to final impact. The intensity distribution of error signals is quantified, measuring their intensity at different locations and times. Pattern characteristics of error signals are identified, categorizing similar signals into error signal patterns. For example, a periodic temperature shift signal caused by manufacturing differences is captured in a temperature sensor array, with an amplitude shift of 0.5 degrees every 10 ms. Contextual information of the captured error signals is recorded, including the system state and operating conditions at the time of the error. The captured error signals are filtered to separate useful error characteristic signals from irrelevant noise signals.
[0068] A correction signal is generated by performing signal reversal processing on the source of the error signal. Based on the captured error signal source information, signal reversal technology is used to generate a corresponding correction signal to cancel the error's effect. The waveform characteristics of the error signal are analyzed to determine the type of signal reversal processing method and parameter settings. A mathematical model for signal reversal is designed, generating a correction signal with the opposite phase and equal amplitude to the error signal through inverse operation. The time delay of signal reversal is calculated to ensure that the correction signal and the error signal are precisely aligned in time. Frequency domain signal reversal processing is implemented, converting the time-domain error signal to the frequency domain using Fourier transform for reversal. The amplitude and phase parameters of the correction signal are generated, and effective error cancellation is achieved by precisely controlling the characteristic parameters of the correction signal. A method for synthesizing the correction signal is designed to combine multiple single-frequency correction signals into a composite correction signal. For example, for a 0.5-degree periodic offset error in a temperature sensor, a -0.5-degree correction signal with the opposite phase is generated, and the temperature offset is reduced to within 0.05 degrees through signal superposition. The stability requirements of the correction signal are analyzed to ensure that the correction signal maintains a stable cancellation effect during long-term use.
[0069] Fault tolerance enhancement factors are generated based on the fault tolerance strength analysis of the correction signal. The cancellation capability of the correction signal is calculated as follows: Cancellation efficiency = Error signal reduction / Original error signal strength × 100%. The types and strength ranges of errors that the correction signal can effectively cancel are determined. The effective duration of the correction signal under different operating conditions is analyzed. The stability and reliability of the correction signal in long-term use are analyzed. The fault tolerance strength value is calculated as follows: Fault tolerance strength = Cancellation efficiency × Duration coefficient × Stability coefficient. The fault tolerance enhancement factor F is calculated as follows: Fault tolerance strength / Baseline fault tolerance capability, where F is the fault tolerance enhancement factor, with a value range of 0.5-2.0, and F>1 indicates improved fault tolerance. The weight allocation of the enhancement factor is adjusted according to the fault tolerance requirements of different hardware components. The applicability of the factor under different operating scenarios is determined. The fault tolerance enhancement factor is standardized, mapping the factor value to a standard range of 0.5-2.0.
[0070] In some embodiments, incorporating a fault tolerance enhancement factor into a load-balanced processing data stream to generate a pre-compensation processing sequence includes: monitoring the current compensation intensity using the fault tolerance enhancement factor; performing threshold dynamic adjustment on the current compensation intensity to obtain a variable threshold configuration; implementing self-adjusting allocation to the load-balanced processing data stream through the variable threshold configuration to obtain a modulated data stream; and forming a pre-compensation processing sequence based on the modulated data stream.
[0071] The current compensation intensity is monitored using a fault tolerance enhancement factor. The real-time compensation intensity is calculated as C = F × L, where F is the fault tolerance enhancement factor and L is the current load level. The changing trend of the compensation intensity is analyzed to identify the pattern of its variation with time and load. The distribution characteristics of the compensation intensity are statistically analyzed, and the differences in compensation intensity across different processing windows are statistically analyzed. Abnormal fluctuations in the compensation intensity are detected, and anomalies such as sudden increases or decreases in compensation intensity are identified. The stability indices of the compensation intensity are evaluated, and the stability of the compensation intensity is assessed using variance and coefficient of variation. The correlation between the compensation intensity and system performance is analyzed to identify the impact of the compensation intensity on processing efficiency. The monitoring cycle for the compensation intensity is set, and an appropriate monitoring frequency is determined based on the system's response characteristics. The compensation intensity is classified into three levels: low (C < 1.0), medium (1.0 ≤ C < 1.5), and high (C ≥ 1.5).
[0072] A variable threshold configuration is obtained by dynamically adjusting the current compensation intensity. Based on the monitored current compensation intensity, the system's operating threshold configuration is adjusted using a threshold adjustment method. The difference between the current compensation intensity and the standard threshold is analyzed to assess the demand and direction of threshold adjustment. A strategy for dynamic threshold adjustment is designed, employing a combination of proportional, integral, and derivative adjustments. The range of threshold adjustment amplitude is determined to ensure that the adjusted threshold remains within the system's acceptable range. Different threshold adjustment strategies are set according to the compensation intensity level: conservative adjustment is used for high compensation intensity, and aggressive adjustment is used for low compensation intensity. The time constant of threshold adjustment is analyzed to determine the time required for the threshold to adjust from the current value to the target value. Parameter constraints for the variable threshold configuration are set to limit the maximum amplitude and rate of change of the threshold. For example, when the system detects that the compensation intensity increases from 1.2 to 1.6, the error tolerance threshold is automatically adjusted from 0.5% to 0.3%, and the processing timeout threshold is extended from 20ms to 25ms. The stability requirements for the variable threshold configuration are determined to ensure that threshold adjustment does not cause system oscillation. A parameter table for the variable threshold configuration is generated, containing threshold setting parameters under various operating conditions.
[0073] For example, obtaining a modulated data stream by implementing self-adjusting allocation to a load-balanced processing data stream through a variable threshold configuration includes: parsing the compensation residual weight using the variable threshold configuration; determining the residual recovery intensity based on the compensation residual weight; using the range where the residual recovery intensity exceeds the threshold as the recovery window; and redistributing the residual to the load-balanced processing data stream according to the recovery amount corresponding to the recovery window to obtain the modulated data stream.
[0074] The compensation residual weights are analyzed using a variable threshold configuration. Based on the generated variable threshold configuration information, the distribution of residual weights generated during the compensation process is extracted using analytical techniques. The compensation residuals for each processing window are calculated; the residual equals the difference between the actual compensation value and the ideal compensation value. The distribution characteristics of the compensation residuals are analyzed, and statistical parameters such as the mean, variance, and distribution range of the residuals are statistically analyzed. The weight allocation pattern of the compensation residuals is identified to determine which processing windows produce larger compensation residuals. For example, in the temperature compensation process, core A's target compensation is +2 degrees, but the actual compensation is +2.3 degrees, resulting in a positive residual of +0.3 degrees, while core B's target compensation is -1 degree, but the actual compensation is -0.7 degrees, resulting in a negative residual of +0.3 degrees. The compensation residual weight W = |R| / ∑|R| is calculated, where W is the residual weight, R is the compensation residual, and ∑|R| is the sum of the absolute values of all residuals. The temporal variation characteristics of the compensation residual weights are analyzed to identify the evolution law of the residual weights over time. The spatial distribution of the compensation residual weights is detected to determine the degree of concentration of residuals in different processing regions. Evaluate the relevance of the compensation residual weights and analyze the correlation between residual weights in different processing windows. Identify outliers in the compensation residual weights and find processing windows where the weights significantly deviate from the normal range.
[0075] The residual recovery intensity is determined based on the residual weight. The mapping relationship between residual weight and recovery intensity is analyzed; windows with larger residual weights require higher recovery intensities. The residual recovery intensity S = k × W^α is calculated, where k is the proportionality coefficient, W is the residual weight, and α is the exponential parameter. The allocation strategy of recovery intensity is adjusted according to the processing capacity of the processing window, with windows with stronger processing capacity undertaking higher recovery intensities. The constraints on residual recovery intensity are analyzed to ensure that the recovery intensity does not exceed the carrying capacity of each processing window. The overall balance of residual recovery intensity is evaluated to ensure that all residuals are effectively recovered. The priority of residual recovery intensity is set, and the priority order of recovery is determined according to the importance of the residuals. The temporal characteristics of residual recovery intensity are analyzed to determine the execution time and duration of the recovery operation. For example, for a processing window with a temperature residual weight of 0.4, the recovery intensity is calculated to be 1.6, indicating that this window needs to perform residual recovery operations at 1.6 times the standard intensity.
[0076] The range where residual recovery intensity exceeds a threshold is designated as the recovery window. Based on the analyzed residual recovery intensity distribution, regions with intensity exceeding the set threshold are identified and defined as residual recovery windows. A threshold standard for residual recovery intensity is set, which is 1.2 times the average recovery intensity. The recovery intensity distribution of all processing areas is scanned, and continuous regions with intensity exceeding the threshold are marked. The geometric parameters of the recovery windows are determined, including features such as window size, shape, and boundary location. The distribution of the number of recovery windows is analyzed, and the total number and distribution density of recovery windows in the system are statistically analyzed. The coverage of the recovery windows is evaluated, and the number and type of residuals that each window can recover are determined. The priority ranking of recovery windows is set, and the priority is determined based on the recovery intensity and residual importance of the windows. The temporal characteristics of the recovery windows are analyzed, and the activation time, duration, and failure time of the windows are determined. The recovery capacity of the recovery windows is evaluated to understand the maximum residual recovery capability of each window. For example, when the recovery intensity threshold is set to 1.3, the system identifies two processing areas with recovery intensities of 1.6 and 1.8 as recovery windows, with window 1 specifically recovering temperature residuals and window 2 specifically recovering voltage residuals.
[0077] The modulated data stream is obtained by redistributing residuals to the load-balanced processing data stream based on the recovery amount corresponding to the recovery window. Using the determined recovery window and recovery amount information, the residual redistribution operation is performed on the load-balanced processing data stream. The actual recovery amount for each recovery window is determined; the recovery amount equals the smaller value between the window capacity and the residual requirement. A strategy for residual redistribution is designed, employing an optimal matching method to allocate residuals to appropriate processing paths. The constraints of residual redistribution are analyzed to ensure that the allocation process satisfies load balancing and resource constraints. The residual redistribution operation is executed, processing the residuals of each recovery window sequentially according to priority. The load distribution after residual redistribution is analyzed to understand the impact of the allocation operation on overall load balancing. The traffic parameters of the modulated data stream are adjusted according to the results of residual redistribution. A characteristic description of the modulated data stream is generated, including key characteristic parameters such as traffic, latency, and priority. The quality improvement effect of the modulated data stream is analyzed by comparing the data stream quality before and after modulation to understand the degree of improvement. For example, by redistributing the residuals, the temperature residuals are transferred from the high-load core A to the low-load core C, and the voltage residuals are transferred from the core B to the core D, thus achieving a rebalancing of the residual loads.
[0078] A pre-compensation processing sequence is formed based on the modulated data stream. Using the generated modulated data stream, a processing sequence with pre-compensation capabilities is constructed to offset the impact of hardware differences. The processing requirements of the modulated data stream are analyzed to determine the design parameters of the pre-compensation processing sequence. The execution flow of the pre-compensation processing sequence is designed, including three stages: preprocessing, compensation processing, and post-processing. The timing arrangement of the pre-compensation processing sequence is configured, determining the execution time and duration of each compensation operation. The resource requirements of the pre-compensation processing sequence are evaluated to understand the computational and storage resources required for the compensation operations. A parameter adjustment interface for the pre-compensation processing sequence is designed, allowing adjustment of compensation parameters according to real-time conditions. The execution efficiency of the pre-compensation processing sequence is analyzed to ensure that the compensation operations do not significantly degrade the system's processing performance. A state management system for the pre-compensation processing sequence is constructed to track the sequence's execution status and completion progress. An exception handling strategy for the pre-compensation processing sequence is set up to address situations where anomalies occur during compensation operations. Functional testing of the pre-compensation processing sequence is performed to confirm its compensation effect under various operating conditions.
[0079] A standardized processing signal is generated by eliminating the impact of hardware differences through a pre-compensation processing sequence. The constructed pre-compensation processing sequence is used to pre-compensate the data stream, eliminating the impact of performance differences between hardware components. Compensation operations in the pre-compensation processing sequence are executed to standardize the processing results of different hardware components. The compensation amount for hardware differences is determined, and corresponding compensation parameters are determined based on the performance deviation of each component. Error correction is performed on the processing results, reducing processing errors caused by hardware inconsistencies through compensation methods. Signal standardization processing is implemented to convert the processed signals from different hardware components into a unified standard format. The quality indicators of the standardized processed signals are analyzed, including characteristics such as signal stability, consistency, and accuracy. The output format of the standardized processed signals is designed to ensure that the signal format meets the input requirements of subsequent processing modules. A completeness check is performed on the standardized processed signals to ensure that the signals contain all necessary information. The timing characteristics of the standardized processed signals are analyzed to ensure that the timing relationships of the signals meet system requirements. For example, the original signal amplitudes output by the three processing cores were 3.2V, 2.8V, and 3.6V, respectively; after pre-compensation processing, they are uniformly adjusted to a standard output signal of 3.0V. The generated standardized processing signal is used as the system output to finally complete the chip data processing.
[0080] To implement the chip data processing method corresponding to the above method embodiments, in order to achieve the corresponding functions and technical effects. See also Figure 2 , Figure 2 This diagram illustrates a structural block diagram of a chip data processing system 200 according to an embodiment of this application. For ease of explanation, only the parts relevant to this embodiment are shown. The chip data processing system 200 provided in this embodiment includes:
[0081] The data acquisition module 201 is used to acquire the chip input data stream and hardware configuration information, and to perform structured preprocessing on the input data stream and hardware configuration information to extract data packet processing attributes.
[0082] Dependency resolution module 202 is used to remove hidden dependencies based on packet processing attributes to form an efficient processing channel, and to determine resource scheduling strategies along the efficient processing channel;
[0083] The conflict scheduling module 203 is used to actively generate scheduling conflict events according to the resource scheduling strategy, construct a reverse scheduling gradient based on the scheduling conflict events to form a dynamic load field, create a processing window by utilizing the load imbalance in the dynamic load field, and couple the processing window with the standard scheduling algorithm to generate adaptive scheduling parameters.
[0084] The bypass control module 204 is used to trigger changes in the internal state of the chip based on adaptive scheduling parameters to generate a state monitoring map, identify bottleneck computing units from the state monitoring map to form a bottleneck bypass network, and remap the bottleneck bypass network to the processing window to form a load-balanced processing data flow.
[0085] The fault tolerance enhancement module 205 is used to obtain response characteristic data based on the load-balanced processing data stream, perform hardware inconsistency transformation based on the response characteristic data to form a fault tolerance enhancement factor, integrate the fault tolerance enhancement factor into the load-balanced processing data stream to generate a pre-compensation processing sequence, and eliminate the impact of hardware differences through the pre-compensation processing sequence to generate a standardized processing signal.
[0086] The chip data processing system 200 described above can implement a chip data processing method according to the above method embodiments. The options in the above method embodiments are also applicable to this embodiment, and will not be detailed here. The remaining content of this application's embodiments can be referred to the content of the above method embodiments, and will not be repeated in this embodiment.
[0087] The purpose of the above embodiments is to reproduce and derive the technical solution of the present invention by way of example, and to fully describe the technical solution, purpose and effect of the present invention. The purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosure of the present invention, and not to limit the scope of protection of the present invention.
[0088] The above embodiments are not an exhaustive list based on the present invention, and there may be many other embodiments not listed. Any substitutions and improvements made without departing from the concept of the present invention are within the protection scope of the present invention.
Claims
1. A chip data processing method, characterized by, The method comprises the following steps: acquiring chip input data stream and hardware configuration information, and performing structural preprocessing on the input data stream and the hardware configuration information to extract data packet processing attributes; based on the data packet processing attributes, hidden dependency relationships are resolved to form an efficient processing channel, and a resource scheduling strategy is determined along the efficient processing channel; according to the resource scheduling strategy, a scheduling conflict event is actively generated, a reverse scheduling gradient is constructed based on the scheduling conflict event to form a dynamic load field, a processing window is created by using load imbalance in the dynamic load field, the processing window is coupled with a standard scheduling algorithm to generate adaptive scheduling parameters; based on the adaptive scheduling parameters, a state monitoring atlas is generated by triggering chip internal state changes, bottleneck computing units are identified from the state monitoring atlas to form a bottleneck bypass network, and the bottleneck bypass network is remapped to the processing window to form a load-balanced processing data stream; based on the load-balanced processing data stream, response characteristic data is acquired, hardware inconsistency is converted based on the response characteristic data to form a fault tolerance capability enhancement factor, the fault tolerance capability enhancement factor is integrated into the load-balanced processing data stream to generate a pre-compensation processing sequence, and the hardware difference is eliminated through the pre-compensation processing sequence to generate a standardized processing signal.
2. The method according to claim 1, characterized in that, The structural preprocessing on the input data stream and the hardware configuration information to extract the data packet processing attributes comprises the following steps: collecting physical layer processing benchmark configurations by using the input data stream and the hardware configuration information; creating a virtual processing layer mapping operation by using the physical layer processing benchmark configurations to form a cross-layer data feature; obtaining processing attribute parameters by performing hierarchical matching analysis on the cross-layer data feature; extracting the data packet processing attributes based on the processing attribute parameters.
3. The method according to claim 1, characterized in that, Based on the data packet processing attributes, hidden dependency relationships are resolved to form an efficient processing channel, which comprises the following steps: identifying serial path regions by using the data packet processing attributes; forming parallel propagation chains by performing dependency relationship reconstruction tracking through the serial path regions; obtaining reduction control points by performing path reduction positioning on the parallel propagation chains; forming the efficient processing channel based on the reduction control points.
4. The method according to claim 1, characterized by, Based on the scheduling conflict event, a reverse scheduling gradient is constructed to form a dynamic load field, which comprises the following steps: obtaining conflict characteristic parameters by performing conflict intensity analysis on the scheduling conflict event; forming a toughness evaluation configuration by performing system stability testing using the conflict characteristic parameters; obtaining response switching nodes by constructing a conflict response mapping through the toughness evaluation configuration; constructing a reverse scheduling gradient based on the response switching nodes to form a dynamic load field.
5. The method of claim 1, characterized in that, From the state monitoring atlas, bottleneck computing units are identified to form a bottleneck bypass network, which comprises the following steps: identifying bottleneck computing units by using the state monitoring atlas; converting the bottleneck computing units into filter trigger thresholds; obtaining filtered data features by performing data quality screening based on the filter trigger thresholds; forming a bottleneck bypass network by performing bottleneck reverse utilization on the filtered data features.
6. The method of claim 1, characterized in that, Based on the response characteristic data, hardware inconsistency is converted to form a fault tolerance capability enhancement factor, which comprises the following steps: identifying hardware inconsistency by using the response characteristic data; capturing error signal sources based on the hardware inconsistency; generating correction signals by performing signal inversion processing through the error signal sources; generating a fault tolerance capability enhancement factor by performing fault tolerance intensity analysis based on the correction signals.
7. The method of claim 1, characterized in that, The fault tolerance enhancement factor is integrated into the load-balanced processing data stream to generate a pre-compensation processing sequence, including: The current compensation strength is monitored by using the fault tolerance enhancement factor; The variable threshold configuration is obtained by performing threshold dynamic adjustment on the current compensation strength; The modulated data stream is obtained by implementing self-adjusting distribution to the load-balanced processing data stream through the variable threshold configuration; The pre-compensation processing sequence is formed based on the modulated data stream.
8. The method according to claim 3, characterized in that, The path reduction positioning is implemented on the parallel propagation chain to obtain the reduced control point, including: The broken dependency chain segment is identified by using the parallel propagation chain; The recycling direction is determined through the broken dependency chain segment; The reorganized path network is identified based on the recycling direction; The reduced control point is accumulated by converting the reorganized path network into a parallel processing path.
9. The method according to claim 7, characterized in that, The modulated data stream is obtained by implementing self-adjusting distribution to the load-balanced processing data stream through the variable threshold configuration, including: The compensation residual weight is parsed by using the variable threshold configuration; The residual recovery strength is determined based on the compensation residual weight; The range where the residual recovery strength exceeds the threshold is taken as the recovery window; The modulated data stream is obtained by performing residual re-distribution to the load-balanced processing data stream through the recovery amount corresponding to the recovery window.
10. A chip data processing system, characterized by comprising: It includes: A data acquisition module is configured to acquire chip input data stream and hardware configuration information, and to perform structured preprocessing on the input data stream and hardware configuration information to extract data packet processing attributes; A dependency analysis module is configured to resolve hidden dependency relationships based on the data packet processing attributes to form an efficient processing channel, and to determine a resource scheduling strategy along the efficient processing channel; A conflict scheduling module is configured to actively create scheduling conflict events according to the resource scheduling strategy, to build a reverse scheduling gradient based on the scheduling conflict events to form a dynamic load field, to create a processing window using load imbalance in the dynamic load field, and to couple the processing window with a standard scheduling algorithm to generate adaptive scheduling parameters; A bypass control module is configured to trigger a state change in the chip based on the adaptive scheduling parameters to generate a state monitoring atlas, to identify a bottleneck calculation unit from the state monitoring atlas to form a bottleneck bypass network, and to remap the bottleneck bypass network to the processing window to form a load-balanced processing data stream; A fault tolerance enhancement module is configured to obtain response characteristic data based on the load-balanced processing data stream, to perform hardware inconsistency conversion based on the response characteristic data to form a fault tolerance enhancement factor, to integrate the fault tolerance enhancement factor into the load-balanced processing data stream to generate a pre-compensation processing sequence, and to eliminate the influence of hardware differences through the pre-compensation processing sequence to generate a standardized processing signal.