Systolic array scheduling processing method, device, equipment and medium

By dynamically generating scheduling strategies on a pulsating array processing device and combining real-time operating status and data characteristics, the problem of collaborative optimization between the compilation module and the scheduling decision module is solved, achieving efficient and adaptive task scheduling and execution, and improving computing efficiency and energy efficiency ratio.

CN120929220APending Publication Date: 2025-11-11PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511185271.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In the existing technology, there is a lack of a collaborative optimization mechanism between the compilation module and the scheduling decision module based on real-time running status and data characteristics, which makes it impossible to achieve dynamic, efficient and adaptive scheduling and execution of processing tasks on the pulse array processing device.

Method used

The system acquires a data processing model and compiles it into an initial processing task based on the hardware architecture characteristics of the systolic array. It then acquires the data to be processed and analyzes its data characteristics, monitors the real-time operating status of the systolic array processing device, generates a scheduling strategy that includes task allocation and data transmission paths, and updates the parameters of the compilation module and the scheduling decision module based on performance data.

Benefits of technology

It achieves real-time adaptation of task allocation and data transmission path, improves computing efficiency and energy efficiency ratio, solves the performance bottleneck and resource waste problem of traditional chip architecture when facing complex computing tasks, and improves processing efficiency and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929220A_ABST
    Figure CN120929220A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a systolic array scheduling processing method, device, equipment and medium, and the method comprises the steps: obtaining a data processing model, and compiling the data processing model into an initial processing task based on hardware architecture features of a systolic array, acquiring to-be-processed data and analyzing data characteristics of the to-be-processed data, monitoring a real-time operation state of the systolic array processing device, inputting the real-time operation state, the data characteristics and an initial processing task into a scheduling decision module to generate a scheduling strategy, executing the scheduling strategy to complete data processing, and collecting performance data; and updating optimization parameters of the compiling module and strategy generation parameters of the scheduling decision module based on the performance data. According to the method, the initial processing task is generated through compiling, the scheduling strategy is dynamically generated in combination with the data features and the running state, parameter updating is achieved through performance data feedback, self-adaptive closed-loop optimization of compiling and scheduling is formed, and the calculation efficiency and the energy efficiency ratio are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a pulse array scheduling processing method, apparatus, device, and storage medium. Background Technology

[0002] In the field of chip computing, especially in scenarios dealing with artificial intelligence tasks, traditional computing architectures have exposed multiple performance bottlenecks. Faced with the ever-increasing computational complexity and data volume in neural network models, general-purpose processors (CPUs) and graphics processing units (GPUs) often suffer from low energy efficiency, fixed instruction scheduling, and wasted computing unit resources, making it difficult to fully utilize the hardware's parallel processing capabilities. Furthermore, in terms of chip resource utilization and computational load allocation, traditional compilation mechanisms lack adaptive capabilities, leading to unstable performance of the same computational model under different data scales or load conditions, and a lack of dynamic tuning mechanisms to optimize overall processing efficiency.

[0003] In the fintech sector, neural networks are frequently used for large-scale risk control modeling, customer profiling, and transaction behavior analysis. These tasks place extremely high demands on model response speed, data throughput, and the accuracy of resource scheduling. However, traditional chip architectures suffer from static scheduling strategies and a lack of targeted task allocation when facing highly concurrent and computationally intensive financial neural network inference tasks. This makes it difficult to cope with the real-time processing needs under peak business loads, thus affecting model response latency and the timeliness of business decisions.

[0004] In the healthcare sector, applications such as intelligent assisted diagnosis, image recognition, and patient status monitoring rely on complex deep neural network models for rapid inference. Traditional chips often face challenges such as uneven performance, slow data loading, and excessive power consumption during operation, especially in mobile or wearable medical devices where achieving low-power, high-efficiency computing is difficult, limiting the promotion and application depth of intelligent diagnostic and treatment systems. Summary of the Invention

[0005] The main objective of this invention is to provide a pulsating array scheduling processing method, apparatus, device, and storage medium, aiming to solve the technical problem in the prior art where the lack of a collaborative optimization mechanism between the compilation module and the scheduling decision module based on real-time running status and data characteristics leads to the inability to achieve dynamic, efficient, and adaptive scheduling and execution of processing tasks on the pulsating array processing device.

[0006] To achieve the above objectives, the present invention provides a pulsating array scheduling processing method, comprising:

[0007] A data processing model is obtained, and the data processing model is compiled into an initial processing task according to the hardware architecture characteristics of the pulsating array through a compilation module.

[0008] Acquire the data to be processed, and analyze the data features of the data to be processed through the feature extraction module;

[0009] Monitor the real-time operating status of the pulse array processing device, and input the real-time operating status, the data characteristics, and the initial processing task into the scheduling decision module to generate a scheduling strategy that includes task allocation and data transmission path;

[0010] The scheduling strategy is executed to drive the pulsating array to process the data to be processed, generate processing results, and collect performance data during the processing.

[0011] Based on the performance data, update the optimization parameters of the compilation module and the strategy generation parameters of the scheduling decision module.

[0012] Furthermore, to achieve the above objectives, the present invention provides a pulse array scheduling processing device, comprising:

[0013] The compilation module is used to obtain the data processing model and compile the data processing model into an initial processing task according to the hardware architecture characteristics of the systolic array.

[0014] The feature extraction module is used to acquire the data to be processed and analyze the data features of the data to be processed.

[0015] The scheduling decision module is used to monitor the real-time operating status of the pulse array processing device, and input the real-time operating status, the data characteristics and the initial processing task into the scheduling decision module to generate a scheduling strategy that includes task allocation and data transmission path;

[0016] The execution module is used to execute the scheduling strategy to drive the pulsating array to complete the processing of the data to be processed, generate the processing result, and collect the performance data during the processing.

[0017] The feedback optimization module is used to update the optimization parameters of the compilation module and the strategy generation parameters of the scheduling decision module based on the performance data.

[0018] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a systolic array scheduling process stored in the memory and executable on the processor, wherein when the systolic array scheduling process is executed by the processor, it implements the steps of the systolic array scheduling method as described above.

[0019] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a systolic array scheduling process, wherein the systolic array scheduling process, when executed by a processor, implements the steps of the systolic array scheduling method as described above.

[0020] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a pulsating array scheduling processing method, apparatus, device, and medium, comprising: acquiring a data processing model and compiling it into an initial processing task based on the hardware architecture characteristics of the pulsating array; acquiring data to be processed and analyzing its data characteristics; monitoring the real-time operating status of the pulsating array processing device; inputting the real-time operating status, data characteristics, and initial processing task into a scheduling decision module to generate a scheduling strategy; executing the scheduling strategy to complete data processing and collecting performance data; and updating the optimization parameters of the compilation module and the strategy generation parameters of the scheduling decision module based on the performance data. This invention compiles the data processing model into an initial processing task adapted to the pulsating array structure and dynamically generates a scheduling strategy by combining the characteristics of the data to be processed and the operating status of the processing device. This enables task allocation and data transmission paths to adapt to hardware resource conditions in real time. Furthermore, by driving parameter updates through performance data feedback, an adaptive closed-loop optimization of compilation and scheduling is formed, improving computational efficiency and energy efficiency ratio. Attached Figure Description

[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0022] Figure 1 This is a schematic diagram of an application environment for the pulsating array scheduling processing method in one embodiment of the present invention;

[0023] Figure 2 This is a flowchart illustrating an embodiment of the pulsating array scheduling processing method of the present invention;

[0024] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the pulse array scheduling processing device of the present invention;

[0025] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0026] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0027] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0028] The pulsating array scheduling processing method provided in this embodiment of the invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can obtain a data processing model from the user terminal and compile it into an initial processing task based on the hardware architecture characteristics of the systolic array. It acquires the data to be processed and analyzes its data characteristics, monitors the real-time operating status of the systolic array processing device, inputs the real-time operating status, data characteristics, and initial processing task into the scheduling decision module to generate a scheduling strategy, executes the scheduling strategy to complete data processing and collect performance data, and updates the optimization parameters of the compilation module and the strategy generation parameters of the scheduling decision module based on the performance data. This invention compiles the data processing model into an initial processing task adapted to the systolic array structure and dynamically generates a scheduling strategy based on the characteristics of the data to be processed and the operating status of the processing device. This allows task allocation and data transmission paths to adapt to hardware resource conditions in real time. Furthermore, performance data feedback drives parameter updates, forming an adaptive closed-loop optimization of compilation and scheduling, improving computational efficiency and energy efficiency. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster composed of multiple servers. The invention will be described in detail below through specific embodiments.

[0029] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the pulsating array scheduling processing method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0030] like Figure 2 As shown, the pulse array scheduling processing method proposed in this invention includes the following steps:

[0031] S10, Obtain the data processing model, and compile the data processing model into an initial processing task according to the hardware architecture characteristics of the pulsating array through the compilation module;

[0032] In this embodiment, acquiring the data processing model typically involves loading a neural network model file from an external model storage repository. This model file contains structure definitions, parameter configurations, training weights, etc., and commonly used formats include ONNX, PB, and H5. The model acquisition process needs to support the parsing capability of heterogeneous model formats to achieve compatible loading of different types of network structures. In practical applications, the model graph structure and tensor flow can be automatically parsed using a model description language to extract the types of computation nodes, connection methods, and the dimensions and types of input and output tensors, and mapped into a structured descriptive representation.

[0033] The compilation module needs to possess model semantic parsing and hardware instruction conversion capabilities. It must be able to receive the aforementioned structured model representation and convert it according to the computational resource characteristics of the systolic array. A systolic array is a two-dimensional array-structured hardware acceleration device characterized by using multiply-accumulate (MAC) units as basic computational components, performing periodic synchronous processing along the data path. The compilation module analyzes the operator types and data flow connections in the model, identifying features such as parallel computation paths, serially dependent nodes, and repeated operator calls. Combined with the hardware layout information of the systolic array, it determines the spatial mapping scheme of the operators on the array.

[0034] During the mapping process, the compilation module needs to evaluate the computational resources required by each computing node, including the tensor dimension, number of channels, and activation function complexity, and dynamically optimize the resource allocation strategy of the MAC units in the systolic array. This process also includes evaluating memory usage patterns to prevent bandwidth bottlenecks or access conflicts in the input and output tensors after mapping. Subsequently, the compilation module converts the above spatial mapping information into a set of instructions recognizable by the systolic array, including but not limited to: tensor loading instructions, computation scheduling instructions, synchronization and data transfer instructions, and output write-back instructions. This ultimately forms the initial processing task, which has a complete scheduling sequence, execution path, and resource usage graph, and can be directly deployed to the target hardware for execution.

[0035] In the implementation process, the data processing model can be obtained by connecting to the local file system, cloud model service, or enterprise private model repository. The model file is parsed by calling the corresponding model identifier code based on the task configuration file. The compilation module can be built based on a tensor compilation framework and a customized backend instruction generation module. During the model graph structure parsing stage, an optimizer based on graph isomorphism algorithms is applied to identify subgraph structures that can be merged in the computation graph, reducing computational path redundancy.

[0036] Support for systolic array hardware architecture in the compilation module can be manifested in a dynamic parameter configuration interface, such as automatically identifying the array's row and column size, supported operator types for MAC units, and supported data format precision (e.g., INT8, BF16, FP32), and dynamically updating these parameters through configuration files or hardware interfaces. The instruction generation stage employs a scheduling graph expansion technique, layering and mapping tasks to time-slice execution tables to support pipelined execution and concurrent resource reuse.

[0037] Data layout rearrangement strategies can also be employed. During the model compilation phase, tensor channel remapping and block-level tensor pruning can improve the on-chip cache utilization of the systolic array during the initial processing tasks, reducing DRAM access latency. Furthermore, for neural network models with branching structures, a conditional execution instruction generation mechanism can be used to enable the systolic array to activate only a subset of computational units based on the classification path of the actual input data, thus reducing power consumption.

[0038] Example Description: In the healthcare business domain, structured models can be used for hardware mapping of medical image segmentation models such as UNet. The compilation module automatically identifies repeated convolutional structures in UNet and merges nodes. Then, based on the number of channels that can be processed in parallel in the systolic array, it allocates resources to generate a scheduling graph, enabling the model to be quickly deployed on edge devices for real-time lesion detection tasks.

[0039] In the fintech business, complex transaction anti-fraud models can undergo structural compression and scheduling optimization. For example, when processing graph neural network models containing multi-layer fully connected layers and attention mechanisms, the compilation module identifies the sparse matrix multiplication part and improves the efficiency of multiplication accumulation through block-level scheduling of systolic arrays, enabling rapid modeling and inference of large-scale user transaction data.

[0040] This embodiment introduces a structure-aware mechanism for pulsating arrays during the compilation process of the data processing model. This makes the initial processing tasks generated by compilation more aligned with hardware execution characteristics in terms of spatial mapping and resource scheduling, reducing computational path conflicts and memory transfer overhead. By integrating structure optimization and instruction generation into a unified compilation process, the deployment efficiency and runtime performance of the model on heterogeneous computing platforms are improved.

[0041] S20, acquire the data to be processed, and analyze the data features of the data to be processed through the feature extraction module;

[0042] In this embodiment, acquiring the data to be processed typically involves pulling input data from a data interface or data source system. This data may exist in the form of images, sequences, tensor matrices, or structured entries, requiring the ability to parse different data types and formats. The source of the data to be processed can be external sensing devices, data acquisition links, business databases, message queues, etc. Data arrival times, formats, and completeness vary from data source to data source, thus requiring a preprocessing component to ensure data format uniformity, boundary alignment, and dimension validity. During data preparation, noise reduction, deduplication, and format parsing operations are also required to ensure the integrity and stability of the input data.

[0043] The feature extraction module is designed around modeling the computational features of the data to be processed. This module comprises three main functions: a feature mapping unit, an index quantization unit, and a structure analysis unit. The feature mapping unit identifies information such as tensor shape, number of channels, and sequence length in the structural dimension of the data to be processed, providing a static input basis for subsequent computational resource scheduling. The index quantization unit analyzes the complexity and sparsity of the data in the content dimension, such as extracting edge density and texture information distribution from image data, and calculating field sparsity ratios, coefficients of variation, or entropy values ​​for structured data. The structure analysis unit identifies nested structures, cross-dimensional correlations, or hierarchical organizational structures in the data to support model adaptability for downstream modules.

[0044] The extraction of data features is not limited to a single static attribute, but rather involves constructing multi-dimensional feature vectors that cover multiple indicator dimensions such as structural complexity, sparse distribution, content diversity, and dynamic variability. These indicators include both static structural attributes and distribution indicators that are dynamically calculated based on the input data content. To ensure computational consistency, the feature extraction module needs to perform normalization and scale alignment operations to ensure that multi-dimensional features can be uniformly evaluated and compared.

[0045] The final output data features are encapsulated into a structured feature vector. This feature vector has clearly defined fields, quantifiable value ranges, and supports similarity analysis or outlier detection with historical feature data. This feature vector will serve as an important input to the scheduling module, supporting dynamic resource allocation and execution path selection.

[0046] In the actual implementation, the data to be processed can be acquired through a multi-threaded acquisition module that connects to a Kafka queue, HTTP interface, or NFS shared directory to achieve real-time data retrieval. Format detection logic is integrated into the acquisition process to mark data records that do not conform to the agreed-upon format as anomalies and send them to a data quality buffer pool.

[0047] The feature extraction module can be deployed as a lightweight front-end computing node, supporting edge deployment and acceleration on heterogeneous platforms. In image processing scenarios, convolution operators are used to statistically analyze image edge complexity and color entropy. In structured data processing, sparse matrix construction and non-zero value ratio estimators are used to obtain sparsity indices. Hash bucketing can also be used to approximate the distribution of field categories.

[0048] Furthermore, a structure analysis module based on a self-attention mechanism can be used to subdivide time-series data and perform dynamic pattern recognition to extract stability and abrupt change indicators of the input data over time, assisting downstream scheduling strategies in selecting steady-state or fast-adaptive paths. Feature vectors can be compressed into standardized input using sparse vector encoding, supporting direct input into the scheduling model for forward inference.

[0049] Example description: In the field of healthcare business, the input data can be a CT slice sequence. The feature extraction module extracts the texture gradient intensity and density distribution changes of each slice to determine whether the image data contains abnormal areas or redundant slices, thereby selecting whether to perform compressed storage, frame skipping, or precision mode adjustment.

[0050] In the fintech business, input data may include user behavior logs and transaction sequences. The feature extraction module extracts indicators such as transaction frequency, amount volatility, and operation type distribution to construct multi-dimensional feature vectors, which support the priority scheduling or precise analysis of abnormal user behavior in subsequent model inference tasks.

[0051] This embodiment introduces a multi-dimensional feature extraction mechanism, including structural complexity, sparsity, and distribution, into the data processing path to achieve a comprehensive characterization of the input data at both the structural and semantic levels. This allows downstream task scheduling and precision control to dynamically adjust strategies based on the actual processing difficulty of the data, thereby improving overall execution efficiency and hardware resource utilization.

[0052] S30, monitor the real-time operating status of the pulse array processing device, and input the real-time operating status, the data characteristics and the initial processing task into the scheduling decision module to generate a scheduling strategy that includes task allocation and data transmission path;

[0053] In this embodiment, monitoring the operating status of the processing device is a prerequisite for the dynamic scheduling mechanism. The processing device is typically a pulse array processing device, and its operating status can be collected in real time from multiple dimensions, including the task queue length of each multiplication-accumulation unit, the remaining execution time of tasks, the current energy consumption level, cache hit rate, and the communication blocking frequency between processing units. All of these indicators need to be detected and quantified in real time by an embedded status acquisition device, and simultaneously subjected to moving averages and anomaly smoothing within a time window to prevent short-term fluctuations from affecting the stability of scheduling decisions. The status acquisition process should also be accompanied by a heartbeat synchronization mechanism and a master-slave clock alignment mechanism to ensure the comparability and timing consistency of cross-module operating statuses.

[0054] Real-time operational status and data features are combined to form the input for scheduling decisions, requiring a unified representation format and structural encapsulation. Tensor structure encoding is typically used to map operational status indicators to a feature vector space, which is then concatenated with previously extracted data features to form a joint representation, creating a unified input interface. This interface features structured field identifiers, dimensional encoding, and position masks, supporting heterogeneous indicator alignment and multi-source input fusion.

[0055] The initial processing task serves as another input, described as a scheduling instruction set, data channel identifier, and task execution graph. This needs to be converted into a graph structure embedding representation by a parsing module, which also identifies data dependencies between tasks and labels the execution order. Each node in the task graph representation carries attribute information such as operator type, data requirement, and processing cost estimate, and is dynamically matched with a scheduling resource matching model.

[0056] The scheduling decision module is typically constructed using neural network architectures, graph optimization engines, or multi-objective reinforcement learning strategies. Its main function is to generate scheduling strategies based on these three types of inputs. The scheduling strategy comprises two key parts: a task allocation scheme and a data transmission path. The task allocation scheme searches for the optimal task mapping position on the processing unit set using a resource estimation function, combining load balancing factors and dependency proximity to achieve the goal of minimizing communication latency. The data transmission path design, on the other hand, generates a low-congestion, high-throughput hierarchical path configuration scheme by analyzing cache levels and data affinity.

[0057] The scheduling strategy outputs a set of structured control instructions, including a task binding table, data flow graph paths, and resource lock acquisition and release sequences, which can be directly called by subsequent processing units during the execution phase.

[0058] In practical implementation, the runtime status collector can encapsulate the execution time, storage usage, and input / output bandwidth of each processing unit into a status matrix by configuring hardware performance counters and low-overhead sampling logic, and report it to the scheduling control module at a periodic frame rate. Data features are provided by the feature extraction module and transmitted using a shared memory area to reduce interface call latency.

[0059] The initial processing tasks are given in the form of a computational graph output after model compilation. Each task unit has a static estimation index, and the task block structure can be obtained through topological sorting and strongly connected subgraph partitioning. The scheduling decision module adopts a scheduling model based on graph neural networks, which learns the priority and scheduling order between tasks through node embedding and edge relationships, and generates a scheduling graph with the minimum overall execution time.

[0060] The data transfer path can be dynamically generated by constructing a two-level transfer matrix, combining DMA path occupancy, shared cache hit probability, and data reuse rate estimation, and the scheduling results can be distributed to each execution unit through a unified scheduling control bus.

[0061] Example: In the field of healthcare, image reconstruction tasks involve a large number of matrix transformations and convolution calculations. The scheduling module can prioritize allocating tasks with high memory requirements to low-load nodes based on the cache utilization and energy consumption level of the processing unit, thereby avoiding power overload that could lead to heat dissipation bottlenecks and ensuring the stability and continuity of the medical image reconstruction process.

[0062] In the fintech business, processing high-frequency trading data involves concurrent inference of multiple sub-models. The scheduling module dynamically divides the task set based on the dependency strength between tasks and data sparsity, assigns latency-sensitive real-time risk assessment tasks to nodes with high response capabilities, and calculates and plans low-priority but high-throughput paths for the back-end risk control model to achieve a balance between business response speed and resource utilization.

[0063] This embodiment integrates the running status, data characteristics, and task information into the scheduling module, enabling task scheduling to consider not only the dependencies of the tasks themselves but also the current resource status of the processing device. This achieves state-aware resource optimization allocation, thereby significantly reducing task waiting time and data transmission conflicts, and improving overall computing throughput and response stability.

[0064] S40, execute the scheduling strategy to drive the pulsating array to complete the processing of the data to be processed, generate the processing result, and collect the performance data during the processing;

[0065] In this embodiment, the scheduling strategy includes a task allocation scheme and data transmission path information. When driving the systolic array to perform processing actions, the scheduling strategy must first be parsed to extract the task allocation mapping relationship and data path configuration. The task allocation scheme binds each computation task to a specific multiplication-accumulation processing unit, and establishes a mapping table by combining the task identifier and the processing unit address. The mapping process must follow resource constraints, including the current occupancy status of the processing unit, available cache space, and operation instruction support capabilities. The data transmission path describes the data dependency channels between tasks using a topology graph structure, clarifying the transport relationship between raw data and intermediate results between different processing units.

[0066] According to the task allocation scheme, the initial processing tasks are distributed to multiple multiply-accumulate units within the systolic array processing unit. During task allocation, instruction decoding, register preloading, and data path initialization are required to ensure that each unit completes its computational load within a preset period. Data path initialization involves the configuration of multi-level memory access, including the loading of input data from external memory to on-chip cache, and the sharing and transfer of intermediate results between processing units. This process requires coordination of data transfer between multiple units using synchronization control signals, based on the flow order defined in the data transmission path.

[0067] The computation execution of the pulsating array is initiated by the scheduling and control module. A pipelined scheduling mechanism ensures tasks progress cyclically within the computation array. Data is delivered between processing units via local communication, avoiding global bus congestion. During execution, the system synchronously collects multiple performance metrics, including total task execution time, unit task processing time, cache hit rate, average load per processing unit, energy consumption, and accuracy deviation. This data is recorded in a structured format by the performance acquisition module. The performance acquisition process is implemented using hardware counters, on-chip current and voltage sensing modules, and a computational output error evaluation module, ensuring data accuracy and timeliness.

[0068] Performance data collection must be performed in parallel during task execution, employing a non-blocking collection strategy to avoid interfering with the main computation process. All performance data will be centrally reported once after task completion and used as feedback input for subsequent optimization modules.

[0069] In practical deployments, an intermediate control layer can be used to parse the scheduling strategy. The DMA controller preloads task instructions and data into the corresponding on-chip storage areas and sets execution status flags for each computing unit within the systolic array. The execution phase employs a hierarchical synchronization mechanism, divided into local synchronization control and global execution coordination. The task processing flow is based on a fixed-cycle pipeline mechanism, combined with input data inflow control and computational state transitions, to achieve parallel computation and communication.

[0070] The performance data acquisition section can integrate dedicated monitoring circuitry to capture computation cycle count, access count, cache call frequency, and power consumption estimates via embedded registers. This data is then combined with the main control module's soft interrupt mechanism to trigger data packet uploading, forming a high-frequency, low-latency data sampling link. If the task requires floating-point precision, an output verification module must also be integrated to calculate the precision deviation by comparing the expected output with the actual output error.

[0071] Example: In the medical and health field, such as in ultrasound image enhancement processing tasks, a large number of image frames need to be filtered and edge-enhanced in real time. This execution mechanism evenly maps the computational tasks to each computing unit and collects runtime and energy consumption indicators, which can achieve high frame rate processing while avoiding thermal runaway problems, ensuring the clarity of medical images and processing response speed.

[0072] In the field of fintech, such as batch credit scoring model inference, multiple linear and nonlinear transformation operations are involved. By precisely scheduling tasks and data paths, multiple models can be executed in parallel. After performance bottlenecks are identified by combining performance data analysis, the model inference order and task combination can be adjusted to improve model throughput and scoring accuracy.

[0073] This embodiment achieves dynamic observability and controllability of computational behavior by precisely distributing task allocation and data path execution to each processing unit of the systolic array and combining it with real-time acquisition of multi-dimensional performance indicators. This ensures that the scheduling strategy is not only theoretically feasible, but also efficient, stable, and energy-controllable in actual hardware operation, providing closed-loop feedback support for subsequent scheduling and compilation optimization.

[0074] S50, based on the performance data, update the optimization parameters of the compilation module and the strategy generation parameters of the scheduling decision module.

[0075] In this embodiment, the update process is based on performance data, involving the construction of a mapping relationship between performance feedback and internal module parameters. It also improves the overall performance of the processing task by dynamically adjusting key control factors during the optimization process. Performance data includes multiple dimensions, such as processing latency, energy consumption, load balancing, cache hit rate, and computational accuracy. Each type of data has different dimensions and optimization objectives, requiring normalization to ensure comparability and composability. Normalization employs a sliding window-based statistical analysis method, constructing a dynamic distribution model for historical performance indicators. The performance level of the current execution data is determined by its position relative to the historical distribution.

[0076] Optimization parameters refer to a series of numerical settings that control the compilation module during task transformation and structure mapping, including task granularity partitioning thresholds, operation fusion strategy weights, and resource utilization density target values. These parameters directly affect the initial task partitioning pattern and operation-level scheduling strategy, and their adjustment can significantly impact the efficiency of computing resource utilization and concurrency. Strategy generation parameters control the path scoring function, resource contention penalty factor, and node priority strategy used by the scheduling decision module during task allocation and path generation.

[0077] To establish the correlation between performance data and the aforementioned parameters, the system introduces an online learning module to construct a bidirectional mapping structure. The first mapping structure takes performance metrics as input and outputs suggestions for adjusting the optimization parameters of the compiler module; the second mapping structure uses the same input and output to provide fine-tuning suggestions for the policy parameters of the scheduling decision module. This mapping relationship is dynamically updated during task execution, and data features from new execution instances are introduced through an incremental learning mechanism, thereby gradually optimizing the parameter adjustment strategy.

[0078] Before each parameter update, the parameters to be updated must be verified on a local test dataset. The verification process includes replay simulation, performance evaluation, and error tolerance testing to ensure that the modification does not introduce systemic performance degradation or resource anomalies. Update actions are completed in atomic transactions to ensure that parameter consistency is not compromised in concurrent processing scenarios. After a successful update, relevant parameter version information must be recorded in the version management module, including parameter hash, update time, verification dataset identifier, and performance improvement metrics, to ensure complete traceability for subsequent rollbacks, audits, or model reconstructions.

[0079] In one implementation, the results of each round of execution can be stored in a structured performance database, and a reward feedback function can be established for parameter adjustment actions using a reinforcement learning policy. The policy learning process employs a proximal policy optimization (PPO) algorithm to optimize the policy in each round. In other implementations, a gradient estimation-based method can be used, where the parameters are differentiated using a performance loss function and updated via backpropagation.

[0080] For different types of systolic array architectures, the weight of metrics in performance feedback can be adjusted. For example, energy-sensitive chips focus more on energy consumption per unit task and temperature rise rate, while high-throughput chips focus more on task execution concurrency and cache access efficiency. Parameter update strategies can be set with different trigger frequencies. For example, updating once after each round of execution under low resource utilization conditions, while a delayed update strategy can be set during stable operation to avoid frequent disturbances.

[0081] Example: In the healthcare field, in a chip processing system used for real-time ECG anomaly detection, by continuously collecting processing latency and energy consumption data of the detection task and adjusting compilation granularity parameters and channel reuse configuration, delay drift and power consumption increase during high-frequency anomaly detection can be effectively avoided, ensuring the real-time performance and stability of the patient monitoring system.

[0082] In the fintech field, during the inference process of high-frequency trading risk assessment models, performance data feedback reveals scheduling conflicts among multiple models. This allows for further fine-tuning of scheduling priority weights and resource allocation density parameters, enabling high-value trading models to obtain resources first and improving the accuracy and response speed of trading decisions at critical moments.

[0083] This embodiment constructs a learning mapping structure between performance feedback and optimization parameters, enabling the compilation and scheduling modules to continuously and adaptively adjust during operation. This improves the stability, energy efficiency, and execution efficiency of the processing model under various operating conditions, solves the problem of performance degradation of static optimization in dynamic scenarios, and forms a closed-loop system for software and hardware collaborative optimization.

[0084] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a pulsating array scheduling processing method, apparatus, device, and medium, comprising: acquiring a data processing model and compiling it into an initial processing task based on the hardware architecture characteristics of the pulsating array; acquiring data to be processed and analyzing its data characteristics; monitoring the real-time operating status of the pulsating array processing device; inputting the real-time operating status, data characteristics, and initial processing task into a scheduling decision module to generate a scheduling strategy; executing the scheduling strategy to complete data processing and collecting performance data; and updating the optimization parameters of the compilation module and the strategy generation parameters of the scheduling decision module based on the performance data. This invention compiles the data processing model into an initial processing task adapted to the pulsating array structure and dynamically generates a scheduling strategy by combining the characteristics of the data to be processed and the operating status of the processing device. This allows task allocation and data transmission paths to adapt to hardware resource conditions in real time. Furthermore, by driving parameter updates through performance data feedback, an adaptive closed-loop optimization of compilation and scheduling is formed, improving computational efficiency and energy efficiency.

[0085] In one embodiment, step S10 above includes:

[0086] S101, Obtain the topology of the data processing model;

[0087] S102, Identify redundant processing nodes in the topology;

[0088] S103, merge consecutive redundant processing nodes of the same type in the topology and update the topological connection relationship of the topology to generate an optimized topology;

[0089] S104, Based on the cell layout of the pulsating array, the optimized topology is divided into multiple processing substructures;

[0090] S105, convert each processing substructure into a pulse array hardware instruction set;

[0091] S106, using a genetic algorithm task mapping strategy, the hardware instruction set is allocated to the multiplication and accumulation processing unit of the systolic array during the compilation stage to generate the initial processing task.

[0092] In this embodiment, the data processing model is typically obtained from a network model specified by an upper-layer AI application or programming task. Its structure is generally a directed graph, where nodes represent operators or computational units, and edges represent data flow relationships. After the compilation module receives such a model, it first needs to extract its complete topology information. The topology is a prerequisite for constructing the mapping relationship between the model and the underlying hardware, and must include the types of all computational nodes, parameter dimensions, dependency order, input-output relationships, and their contextual structure identifiers. This structure can be abstracted and constructed using a static graph parser, which expands and merges complex structures such as cross-layer connections and skip connections to form an unambiguous execution graph.

[0093] Identifying redundant processing nodes in the topology is a critical path to reducing hardware execution load. Redundant nodes include those that perform duplicate computations, whose results are not used downstream, and equivalent computational units that produce the same output. Redundant units can be identified through static data flow analysis, result similarity matching, and operator ablation verification. Among the identified redundant nodes, those with consistent types and continuous connection directions are considered suitable for merging. The merging operation requires synchronously updating the topology connections, removing redundant edges, reconstructing edge weights, and recalculating data flow paths and tensor dimension mappings. The updated structure is the optimized topology, which features more compact data dependency paths and higher resource utilization efficiency.

[0094] The cell layout of a systolic array represents the physical distribution of multiplication-accumulation processing units in a two-dimensional space, typically including row and column dimensions, array alignment, data access step size, and control clock delay structure. Based on this layout, the optimized topology needs to be divided into multiple adaptable substructures. The goal is to group processing nodes with strong data dependencies into the same substructure to reduce cross-cell communication bandwidth consumption. During the partitioning process, the frequency of data interaction between nodes, computational density, and temporal concurrency need to be considered. Graph partitioning algorithms (such as METIS, Louvain, or Kernighan-Lin) are used to divide the topology graph into multiple processing subgraphs with minimal boundaries according to the hardware array's constraints. Each processing substructure must be fully mappable to a physical sub-region within the array and must meet resource constraints such as bandwidth, latency, and power consumption.

[0095] After the processing substructure is divided, it needs to be converted one by one into a hardware instruction set executable by the systolic array. This instruction set includes operation instructions such as data loading, register assignment, multiply-accumulate operations, memory transfer, and synchronization control. Each processing node needs to be mapped to a set of hardware instructions, and the dependency chain relationship needs to be preserved. This process needs to consider factors such as instruction density, hardware execution stack length, and local memory capacity limitations. After instruction generation, structural expansion and resource reuse optimization will be performed to generate an unambiguous low-level instruction set description that can be directly executed on the array.

[0096] The genetic algorithm task mapping strategy is used to efficiently allocate the generated hardware instruction set to the multiply-accumulate processing units of the systolic array. The mapping objective function needs to minimize data transmission paths, balance processing load, reduce waiting cycles, and improve overall throughput. The initial population consists of multiple mapping combinations. Each individual's encoding method combines the task number, unit number, and timestamp. The fitness function is weighted based on actual simulation latency and energy consumption. Crossover is performed by locally exchanging mapping groups, mutation introduces random perturbations to the positions of some task units, and selection preserves the optimal mapping combination. The task mapping scheme output by the algorithm after convergence is used to generate the final hardware resource configuration and scheduling strategy, forming the initial processing tasks.

[0097] Throughout the process, there are close logical relationships between operations. Optimizing the topology determines the compactness of the instruction set, dividing the processing substructure determines the granularity of resource allocation, and the task mapping strategy directly affects execution efficiency. By transforming the software model into structured, hardware-aware initial processing tasks, a mapping path from model abstraction to chip execution is established, laying the foundation for subsequent task scheduling and execution.

[0098] This embodiment effectively reduces redundant computing nodes, improves task mapping efficiency on the systolic array, enhances the locality of computationally intensive regions, reduces data transfer latency, and improves overall resource utilization and operational efficiency by introducing topology optimization, substructure partitioning, and a genetic algorithm-based mapping mechanism during the compilation process. This approach establishes a hardware-software collaborative link from model input to task execution, enhancing the scalability and execution stability of the systolic array when handling complex data processing tasks.

[0099] In one embodiment, step S104 includes:

[0100] S1041, Analyze the row and column dimension parameters of the cell layout of the pulsating array;

[0101] S1042, Analyze the data dependency strength between processing nodes in the optimized topology;

[0102] S1043, Based on the row and column dimension parameters and the processing capability parameters of the multiplication and accumulation units in the pulsating array, a node clustering algorithm is used to group the processing nodes in the optimized topology into multiple node groups;

[0103] S1044, Map each node group to a processing substructure that adapts to the physical structure of the pulsating array according to the row and column dimension parameters;

[0104] S1045, Identify high-dependency processing nodes in the optimized topology whose data dependency strength exceeds a preset dependency threshold;

[0105] S1046, Optimize the boundary partitioning of the processing substructure to maintain the locality of highly dependent processing nodes;

[0106] S1047, Verify the balance of processing load distribution in each processing substructure;

[0107] S1048, When the load distribution balance is lower than the preset balance threshold, adjust the node group allocation and re-verify;

[0108] S1049, When the load distribution balance reaches the preset balance threshold, output the final set of processing substructures.

[0109] In this embodiment, the cell layout of the systolic array refers to the structural arrangement of the multiply-accumulate (MAC) units defined during the chip design phase in a two-dimensional physical space, typically including a matrix structure with rows and columns. This structure determines the flow of data in space and the selection strategy for scheduling paths. Parsing the row and column dimension parameters of this cell layout is a prerequisite for mapping the abstract topology to a specific hardware structure. These dimension parameters involve not only the number of cells but also physical constraints such as their connectivity, control domain boundaries, and local cache access ranges. This process can be obtained by consulting the chip configuration description file or the system-level hardware abstraction interface, and initialization is completed by combining the hardware structure metadata from the specific chip documentation.

[0110] Based on the analysis of the hardware structure, a deep analysis of the optimized topology to be partitioned is required to identify the strength of data dependencies between processing nodes. Data dependency strength is an indicator that measures the density of data interaction and the degree of real-time coupling between different computing nodes, and is typically expressed as a function of dependency frequency, latency sensitivity, data volume, and update rate. In the graph model, dependency strength can be represented by edge weights, with higher-weighted edges indicating strong dependency paths. This analysis can be improved by combining tensor flow tracing, static analysis of the computation graph, and runtime log statistics to enhance the accuracy of dependency identification.

[0111] Next, to map the topology onto a systolic array with limited resources while maintaining the locality of dependencies, node clustering is performed based on the resolved row and column dimension parameters and the processing capacity parameters of each multiplication-accumulation unit. Processing capacity parameters include the unit's maximum throughput, the number of concurrent tasks, the supported data bit width range, and the single-cycle execution capacity. The clustering algorithm should aim to minimize cross-group communication overhead and maximize the task correlation within a group. Algorithms such as hierarchical clustering, spectral clustering, or graph-embedded K-means can be used to group the processing node set. Each node group constitutes a logical partitioning unit, which will subsequently be mapped to a hardware processing substructure.

[0112] The mapping of node groups needs to conform to the physical structural characteristics of the systolic array, especially when there are significant restrictions on the arrangement order of cells and the direction of communication paths in a two-dimensional plane. Improper mapping can lead to increased cross-row and cross-column communication, creating bandwidth bottlenecks. Therefore, during the mapping phase, each node group needs to be assigned to a MAC cell submatrix that conforms to the row and column dimensions through nested mapping functions. At the same time, the adjacency relationship with control logic units and data cache units should be considered to construct a processing substructure with high locality.

[0113] This process also requires identifying high-dependency processing nodes in the optimized topology whose data dependency strength exceeds a preset dependency threshold. These nodes often act as hubs in the data flow, and their output affects multiple downstream nodes. If they are divided into different substructures, it will cause data transmission costs to skyrocket. Therefore, it is necessary to extract these nodes separately and prioritize merging them at the partition boundaries to keep them in the same or adjacent processing substructures, thereby maintaining their locality characteristics during execution. This identification process can be achieved by depth-first traversing the computation graph and recording the node output degree and the sum of edge weights.

[0114] Boundary partitioning optimization aims to further enhance locality and reduce boundary communication overhead. Highly dependent processing nodes and their adjacent nodes should maintain a relatively compact structural layout during physical mapping. Boundary optimization can employ edge shrinkage algorithms or graph convolutional migration mechanisms to fine-tune the boundary partitioning path without disrupting the original node groups, and to reallocate some boundary nodes to improve the internal coupling of substructures.

[0115] The balanced distribution of processing load is one of the core indicators for evaluating the quality of substructure partitioning. It is necessary to verify whether the load of each processing substructure is close to balanced in terms of task quantity, data volume, execution cycle, and other dimensions. The load deviation of each substructure can be calculated using the load entropy calculation method or the mean square error function. If there is a significant imbalance, that is, if the load of any substructure exceeds the set threshold of the load of other structures, then the node group allocation adjustment is performed.

[0116] When the load distribution balance is lower than the preset balance threshold, it indicates that some substructures are overloaded or underloaded, which may lead to resource waste or processing bottlenecks during the scheduling phase. In this case, node groups need to be reallocated. Graph partitioning fine-tuning algorithms, such as local neighborhood rearrangement and weighted boundary sliding, are typically used to migrate some boundary nodes in high-load substructures to structures with lower loads, ensuring balanced resource allocation.

[0117] After adjusting the load distribution balance, a re-verification is required. If the set balance threshold is met, the current partitioning result is deemed acceptable, and the final processing substructure set is output. This set forms the basis for subsequent instruction set generation and physical mapping scheduling.

[0118] This embodiment introduces a hardware-structure-aware processing substructure partitioning mechanism, which effectively identifies critical computing paths in data-dependent topologies, improves the physical deployment locality of processing nodes, and thus reduces cross-unit communication overhead. Furthermore, by dynamically adjusting the load distribution of the processing substructures, the load on each computing sub-region is more balanced, effectively preventing bottlenecks caused by uneven resource utilization and achieving more efficient parallel task processing. This partitioning mechanism enhances the adaptability of the processing structure to heterogeneous tasks, improving overall throughput and energy efficiency.

[0119] In one embodiment, step S20 above includes:

[0120] S201, Load the pre-trained multimodal feature extraction model;

[0121] S202, the data to be processed is input into the multimodal feature extraction model, the complexity features and sparsity features of the data to be processed are extracted by the multimodal feature extraction model, and the data distribution features of the data to be processed are quantified;

[0122] S203, assign a first influence factor, a second influence factor, and a third influence factor to the complexity feature, the sparsity feature, and the data distribution feature, respectively;

[0123] S204, Adjust the complexity feature using the first influence factor to generate a weighted complexity feature;

[0124] S205, use the second influence factor to adjust the sparsity features to generate weighted sparsity features;

[0125] S206, Use the third influencing factor to adjust the data distribution characteristics and generate a weighted distribution characteristic;

[0126] S207, Integrate the weighted complexity feature, weighted sparsity feature and weighted distribution feature to generate a weighted feature set;

[0127] S208, Normalize the weighted feature set to generate normalized feature values;

[0128] S209, verify the validity of the normalized feature values ​​within the confidence interval, filter feature values ​​that exceed the confidence interval, and generate a set of valid feature values;

[0129] S210, use the set of effective feature values ​​to generate a data feature description.

[0130] In this embodiment, the data to be processed may include structured data, images, audio, text, or a multimodal combination of the above types. To understand and measure these data at the feature level, a multimodal feature extraction model with cross-modal understanding capabilities needs to be loaded. This model is designed to integrate an ensemble learning structure of image convolutional networks, text encoders, graph structure awareness modules, or other modal feature encoders, and through a pre-training process, enables the model to extract high-quality features from complex multi-source inputs. Model loading can be accomplished by accessing the underlying model library or an integrated model parameter loading mechanism, and the mapping rules between input and output channels are set based on the current application scenario.

[0131] When inputting the data to be processed into this feature extraction model, each modality of data needs to be preprocessed to match the model's input requirements. For example, images need to have their size and number of channels standardized, text needs to be segmented and vector embedded, and structured data needs to be standardized or flattened from nested structures. During model execution, complexity features, sparsity features, and data distribution features are extracted respectively. Among them, complexity features are quantitative indicators that measure the internal correlation, hierarchical dependence, and redundancy of data, and can be derived from the model's internal attention distribution, convolutional kernel activation map, or recursive deep path information. Sparsity features reflect the non-zero distribution density or the proportion of information present in different dimensions of the data, and can be calculated based on the L0 norm of the feature vector or the proportion of non-zero elements in the sparse matrix. Data distribution features reflect the distribution structure of data values ​​in the value space, and can be obtained through histogram distribution analysis, skewness and kurtosis calculation, or kernel density estimation, and converted into a quantifiable statistical description.

[0132] To express the relative importance of the three types of features in the final scheduling strategy generation process, a first, second, and third influencing factor need to be assigned to the complexity feature, sparsity feature, and data distribution feature, respectively. These influencing factors can be preset by system parameters or adaptively adjusted based on historical performance data. These factors are essentially weighting parameters that control the magnitude of the impact of each feature on subsequent scheduling trade-off decisions.

[0133] Next, each type of feature value is weighted and calculated with its corresponding influencing factor to obtain weighted complexity features, weighted sparsity features, and weighted distribution features. This weighting process not only improves the model's responsiveness to important feature dimensions but also enables differentiated management of the contributions between features. The weighting process can be implemented based on linear multiplication, logarithmic scaling, or exponential smoothing, depending on the task's expectation of feature amplification or suppression.

[0134] Subsequently, the three weighted features are integrated into a unified weighted feature set to facilitate subsequent normalization and modeling. Feature integration methods can include vector concatenation, feature-level fusion, or channel stacking; different methods will affect the degree of interaction between features. The integrated feature set may have issues with inconsistent scales and dimensions, therefore, it needs to be normalized. Normalization can employ methods such as max-min scaling, Z-score standardization, or sigmoid compression to maintain all features within a uniform scale, enhancing the stability and generalization ability of the subsequent model.

[0135] After normalization, the generated feature values ​​need to be validated using confidence intervals. Confidence intervals can be constructed based on the statistical distribution of historical training samples, determining the reliable range of feature values ​​by setting a confidence level. If a feature value falls outside a confidence interval, it indicates a risk of deviating from the training distribution, potentially affecting the system's robustness, and therefore should be filtered out. After filtering, all feature values ​​falling within the confidence intervals are retained, forming a set of valid feature values. This set reflects the high-confidence structural properties of the data in the current processing context.

[0136] Finally, the set of valid feature values ​​is transformed into a structured data feature description, serving as the data-driven foundation for the scheduling module and subsequent task mapping mechanism. The data feature description must include the quantified value, confidence score, influence factor weight, and dimension number for each feature type, and be encapsulated in a unified vector or key-value structure to ensure that subsequent modules can directly access, parse, and use this description.

[0137] This embodiment introduces a multimodal pre-trained model, a weighting mechanism, and a confidence interval verification strategy during the feature extraction stage. This enables the system to achieve a high-dimensional, multi-faceted structural understanding of the input data, thereby generating a data feature description with discriminative capabilities. Weighting operations make the impact of features on scheduling behavior more flexible and controllable, while normalization and confidence verification ensure the numerical stability and semantic credibility of the features. Guided by data features, subsequent task scheduling decisions will better align with the structural complexity, sparsity, and distribution of the input data, effectively improving the accuracy and adaptability of the scheduling strategy, reducing the risk of deviation in the pulsating array resource scheduling process, and achieving hardware and software co-optimization.

[0138] In one embodiment, step S30 above includes:

[0139] S301, monitors the task queue length of each processing unit in the pulse array processing device;

[0140] S302, monitors the real-time bandwidth status of the data transmission channel in the pulse array processing device;

[0141] S303, integrate the task queue length and real-time bandwidth status to generate a real-time running status profile;

[0142] S304, input the real-time running status profile, the data features and the initial processing task into the scheduling decision module, and generate a task allocation scheme and data transmission path through the scheduling decision model;

[0143] S305, Verify whether the task allocation scheme meets the resource constraints of the pulse array processing device;

[0144] S306, Integrate the verified task allocation scheme and the data transmission path to generate a scheduling strategy.

[0145] In this embodiment, the pulsating array processing device is a specialized hardware architecture designed for high-throughput parallel computing scenarios. Its core feature lies in the use of regularly arranged computing units, which achieve pipelined data flow-driven computation through fixed control logic and data paths. This device typically consists of multiple multiply-accumulate (MAC) units arranged in a two-dimensional grid structure, capable of processing continuously arriving data streams at a uniform clock rhythm. The processing units are directionally connected, and data is transmitted systematically within the array, forming a "pulsating" data propagation process, thereby significantly reducing the complexity of control logic and the latency overhead caused by intermediate storage.

[0146] In typical applications, systolic array processing devices include not only core computing units, but also local data buffers, scheduling and control modules, on-chip interconnects, memory interfaces, and logic interfaces for task configuration and loading of compilation results. These modules work together to form a complete, deployable chip-level computing platform. Its design goal is to improve computational efficiency and energy efficiency by supporting the intensive tensor multiplication and addition operations during neural network computation through highly customized hardware logic.

[0147] For example, the TPU (Tensor Processing Unit) is a typical example of a systolic array processing device. It contains a two-dimensional array of 256×256 MAC units, suitable for inference and training tasks in large-scale neural networks. In a TPU, data is propelled along a fixed path in a periodic manner, completing one computation and transfer per clock cycle. This ensures efficient data flow within the array, avoiding the frequent storage read / write bottlenecks found in traditional CPUs or GPUs. In tasks such as medical image analysis and financial transaction prediction, processing devices based on this structure can quickly perform operations such as convolution, matrix multiplication, and activation transformations, significantly improving model response speed and system throughput.

[0148] Based on a comprehensive understanding of the operational dynamics of the systolic array processing unit, a closed-loop input system for real-time scheduling is constructed to ensure context-adaptive scheduling behavior. First, regarding the internal structure of the systolic array processing unit, the task queue length of each processing unit needs continuous monitoring. This queue length is an instantaneous quantitative indicator of the number of tasks pending processing in each multiply-accumulate unit (MAC) or processing channel, accurately sampled by accessing the local instruction buffer, task delivery pointer, and execution completion status bits. These task queue lengths are summarized in a two-dimensional tensor and mapped to the spatial dimension of the physical array layout to describe the current load distribution along the computational path.

[0149] Simultaneously, the real-time bandwidth status of the data transmission channel needs to be collected. This bandwidth status includes not only the remaining available bandwidth of the physical link, but also the contention relationship, transmission delay, and arbitration status on the logical data path. The data channel can cover path selection overhead, bus usage frequency, communication delay statistics, and bandwidth utilization under the NoC architecture. Key status variables can be extracted through the throughput counter, buffer pool status register, and flow control protocol feedback in the transmission arbitrator. This dimension is used to measure the spatial distribution pattern of communication bottleneck locations and data exchange capabilities.

[0150] After acquiring the raw state information from the above two dimensions, the system integrates them into a real-time operational status profile. This profile is organized using a multi-channel feature tensor structure. The channel dimension indicates the resource type (e.g., compute load channel, communication bandwidth channel, etc.), and the spatial dimension indicates the location and adjacency structure of each resource node in the physical array. This supports the scheduling model in making topology-sensitive policy inferences under spatial constraints. The operational status profile can also incorporate short-term state evolution trends from the time axis dimension, adding operational trend perception capabilities through a sliding time window mechanism, thereby improving the scheduling system's tolerance to volatility.

[0151] Subsequently, the runtime status profile, data features from the feature extraction module (including normalized complexity, sparsity, and data distribution attributes), and the initial processing task structure generated from the compilation module are input into the scheduling decision module. This module typically consists of a scheduling model constructed using a fusion graph structure embedding and a deep policy optimization algorithm. Its function is to predict the optimal task deployment graph under resource-constrained conditions by combining the current system load and task content. The output includes a task allocation scheme, specifying which processing unit each processing subtask should be assigned to, as well as the data transmission path, defining the routing, hop count, and communication priority of data flows between tasks.

[0152] This is followed by the resource constraint verification phase. Based on the hardware configuration constraints of the systolic array processing device, including the maximum number of concurrent tasks supported by a single unit, bandwidth bottleneck thresholds, data reuse cache capacity, and control instruction window size, the system performs feasibility reasoning using a mapping graph structure. If issues such as allocation conflicts, link overload, or time window conflicts exist, the system either rolls back to adjust input parameters or enters a retraining process. The verified task allocation and data path combinations are integrated into the final scheduling strategy, serving as the set of execution scheduling instructions for that scheduling cycle, driving the systolic array to execute cycle by cycle according to the clock rhythm.

[0153] Throughout the process, each link transmits the scheduling context state in a lossless, closed-loop, and structure-preserving manner. The scheduling strategy is no longer based on a static orchestration template, but rather dynamically learns and maps the state context jointly shaped by the system and data at each moment, thereby achieving the immediacy of strategy selection, the structural rationality of task allocation, and the maximization of system resource utilization.

[0154] This embodiment constructs a state profile that integrates computing load and communication bandwidth status by dynamically sensing the operating status of the systolic array processing device from multiple dimensions. This enables the scheduling decision-making process to adapt to changes in internal system resources in real time. By combining data characteristics and initial processing tasks as scheduling inputs, task allocation and data transmission paths are generated under global resource constraints, effectively avoiding task backlog and communication conflicts. The generation process of the scheduling strategy undergoes rigorous resource feasibility verification to ensure the executability and optimization of the strategy in terms of spatial structure, task load, and communication topology. This achieves a state-driven closed-loop control mechanism for scheduling, significantly improving the processing efficiency and system throughput of the systolic array processing device.

[0155] In one embodiment, step S40 above includes:

[0156] S401, In the near-memory storage unit, a sparse coding module is used to compress the high sparse data in the data to be processed to generate compressed data;

[0157] S402, the compressed data is stored in a multi-level storage structure;

[0158] S403, according to the data transmission path in the scheduling strategy, transmit the compressed data in the multi-level storage structure;

[0159] S404, The compressed data is decompressed in the near-memory storage unit to generate decompressed data;

[0160] S405, according to the task allocation scheme in the scheduling strategy, the initial processing task is distributed to the specific processing unit of the pulsating array during the execution phase;

[0161] S406, input the decompressed data into the specific processing unit;

[0162] S407, The data features are received in the specific processing unit;

[0163] S408, in the specific processing unit, the processing accuracy mode is adjusted according to the data characteristics;

[0164] S409, the multiplication and accumulation unit in the specific processing unit is started in the adjusted processing precision mode, the initial processing task is executed using the decompressed data, and the processing result is generated;

[0165] S410 collects and processes task execution time, energy consumption, and processing result accuracy as performance data.

[0166] In this embodiment, the scheduling strategy is executed to drive the pulsating array to process the data to be processed. First, the data needs to be preprocessed in the nearest-memory storage unit. The nearest-memory storage unit has localized data processing capabilities and offers advantages such as low latency and low power consumption. A sparse coding module is introduced into this storage unit to compress the highly sparse portion of the data to be processed. Sparsity refers to the low proportion of valid values ​​relative to the total number of values ​​in the data. Sparse coding, such as CSR (Compressed Sparse Row), COO (Coordinate), or bitmap-based coding methods, can significantly reduce the storage and transmission load. After the compression process is completed locally, compressed data is generated.

[0167] Compressed data is written into a multi-tiered storage structure, which includes various storage units such as near-memory cache, on-chip cache, and off-chip memory. This tiered storage system dynamically adjusts storage distribution based on data access frequency and latency, enabling layered utilization of storage resources and bandwidth optimization. Data is transmitted within this multi-tiered structure along data transmission paths defined in the scheduling strategy. These paths are generated by the scheduling decision module and designed in conjunction with task allocation schemes and real-time operational status to ensure optimal data flow under bandwidth bottlenecks and resource conflicts.

[0168] After compressed data is transmitted to the nearest-memory processing unit, this unit performs the decompression operation, restoring the data to a state that can be directly processed by the computing unit. The decompression method must correspond one-to-one with the sparse coding scheme to ensure data integrity and decoding efficiency. Next, the task allocation scheme defined by the scheduling strategy maps the initial processing tasks to specific processing units in the systolic array during the execution phase. These specific processing units can be a set of multiplication-accumulation units allocated in the array according to the topology and operator matching method, possessing distributed parallel computing capabilities.

[0169] After the decompressed data is sent to the designated processing unit, the system needs to acquire data characteristics associated with the data, such as complexity, sparsity, and distribution indicators, to dynamically adjust computational behavior. During this stage, the processing unit adjusts the processing precision mode based on the data characteristics. The processing precision mode can include floating-point precision bits (such as FP32, FP16, INT8), fixed-point bit width, dynamic quantization range, etc., and can be flexibly selected based on data complexity and tolerance for error to improve computational efficiency and reduce energy consumption.

[0170] After the specific processing unit completes the precision mode adjustment, the internal multiplication-accumulation unit is activated to perform the initial processing task based on the decompressed data. The multiplication-accumulation unit performs core neural network operations such as parallel matrix calculations and convolution operations in a pipelined manner, outputting the final processing result. After processing is complete, the system will collect the performance data generated throughout the entire execution process. This performance data covers task execution time, energy consumption, and the accuracy of the final processing result, and is used for subsequent optimization feedback.

[0171] This embodiment effectively reduces the storage bandwidth pressure on the data being processed during transmission by performing sparse compression in the nearby storage unit and decompression in the nearby processing unit, thus alleviating the load bottleneck of the multi-level storage system. The scheduling strategy combines the task allocation scheme and data transmission path to ensure a high degree of matching between processing units and data, thereby avoiding idle computing power or data blockage. During execution, the processing precision mode is dynamically adjusted according to data characteristics to achieve an adaptive balance between energy consumption and computational accuracy. The collected performance data reflects the execution performance of the processing system and can provide crucial support for subsequent optimization of scheduling strategies and compilation parameters.

[0172] In one embodiment, step S50 above includes:

[0173] S501, establish a first correlation between the performance data and the optimization parameters of the compilation module;

[0174] S502, establish a second correlation between the performance data and the strategy generation parameters of the scheduling decision module;

[0175] S503, the first and second association relationships are processed by the online learning module to generate update parameters for the compilation module and update parameters for the scheduling decision module;

[0176] S504, verify the performance improvement of the compilation module updating parameters and the scheduling decision module updating parameters on the test dataset;

[0177] S505, when the performance improvement exceeds a preset threshold, the optimization parameters of the compilation module are updated to the update parameters of the compilation module through atomic transactions, and the strategy generation parameters of the scheduling decision module are updated to the update parameters of the scheduling decision module through atomic transactions.

[0178] S506 records the updated parameter version information to the version management module.

[0179] In this embodiment, using performance data to dynamically update the parameter configurations in the compilation and scheduling decision modules is key to achieving the adaptive evolution capability of the processing system. Performance data refers to a combination of information collected by the systolic array processing device during actual execution, including metrics such as execution time, energy consumption, and task accuracy. These metrics collectively reflect the system's performance under the current task configuration. To enable performance feedback to guide module parameter updates, a dual correlation between parameters and performance needs to be established first.

[0180] First, we establish the initial correlation between performance data and compiler optimization parameters. These parameters include topology reduction granularity, substructure partitioning strategy, instruction generation path, and node fusion rules. In large-scale computing tasks, fine-grained parameter tuning directly impacts hardware execution efficiency. By using performance data as input features and combining it with historical execution records of the current task under different parameter configurations, we form a mapping function or correlation model to reflect the parameter response characteristics corresponding to a certain type of performance.

[0181] Secondly, a second correlation is established between performance data and the policy generation parameters of the scheduling decision module. These parameters include resource weight coefficients in the task allocation algorithm, path selection heuristics, and threshold settings for congestion avoidance mechanisms. These parameters directly influence the generation of the scheduling policy, while performance data such as bandwidth usage and task latency reflect the bottlenecks of the current scheduling policy. This correlation is established through paired training of historical scheduling behavior and performance results, constructing a quantitative expression model of how parameters affect scheduling outcomes.

[0182] Based on the two types of relationships mentioned above, an online learning module is introduced to perform parameter updates. The online learning module employs a continuous training and inference mechanism, constantly receiving new performance data and updating the parameter influence function, thereby generating update parameters for the compilation module and scheduling decision module within the current cycle. This process can be implemented based on algorithms such as reinforcement learning, adaptive gradient descent, and incremental regression, and an experience caching mechanism is used to avoid offset errors caused by the drift in the distribution of old and new data.

[0183] After generating the updated parameters, their performance improvement on a representative test dataset must be verified. The test dataset should cover multiple dimensions, including task type, data sparsity, and bandwidth status, and multiple simulations should be conducted to reproduce the actual operating scenario. During the simulation, the original and updated parameters are applied separately, and their changes across multiple performance metrics are compared. If the detected performance improvement exceeds a preset performance improvement threshold, the updated parameters are considered valuable in replacing the original parameters.

[0184] During the parameter replacement phase, atomic transactions are used to update parameters in both modules. Atomic transactions ensure that parameter update operations are uninterrupted and cannot be partially executed in a concurrent environment, avoiding inconsistencies in parameter states during the update process. Parameter updates in the compilation module and the scheduling decision module are committed as two interdependent sub-transactions; if either fails, the entire process is rolled back, thus ensuring the stability and consistency of task processing.

[0185] After the update is complete, the new parameter configuration is written to the version management module. The version management module records metadata such as parameter version number, update timestamp, applicable task scope, and performance improvement report, and supports rollback mechanism and version comparison function so that the system can restore the version when it detects performance regression or abnormal behavior.

[0186] Example Description: In a high-load medical imaging diagnostic center, to achieve real-time processing and abnormal area identification of CT scan data, the system receives a set of three-dimensional chest CT image data to be processed. This data serves as input to run a trained neural network model for lung lesion detection on a pulse array processing device, assisting doctors in early screening for lung nodules.

[0187] First, the system retrieves a lung lesion recognition model built on a deep convolutional neural network and parses its topology through the compilation module, including the input layer, multiple convolutional modules, residual connections, attention mechanism, and output prediction layer. During parsing, multiple duplicate standard convolutional modules are identified in the model topology, which are redundant processing nodes. To improve computational efficiency, these consecutive identical convolutional structures are merged, and the connection relationships between neural network layers are updated accordingly to form an optimized topology. Subsequently, the compilation module obtains the cell layout parameters of the current systolic array, including the number of rows and columns, the distribution of multiplication and accumulation units, and the accessible cache structure. Based on this hardware layout, the optimized topology is divided into multiple logical processing substructures and mapped to the hardware instruction set supported by the systolic array. Further, a genetic algorithm is used for task mapping optimization, allocating the instruction set to specific hardware units to generate initial processing tasks adapted to the target chip.

[0188] Next, the system loads 3D CT image data, which is then fed into the feature extraction module as the data to be processed. This module loads a multimodal feature extraction network pre-trained based on image and structural descriptors to extract complexity features (such as spatial gradient distribution), sparsity features (such as the sparse distribution rate of non-zero voxels in the spatial domain), and data distribution features (such as histogram balance and grayscale dynamic range) from the input data. Each feature dimension is assigned a corresponding influence factor, such as increasing the complexity weight for images with high noise levels. After weighted processing, the system normalizes the overall feature set and performs confidence interval verification on the feature value set, filtering out high-confidence outliers and generating an effective feature set representing the characteristics of the image data.

[0189] Before inference, to rationally allocate computing resources and avoid processing bottlenecks, the system enters a scheduling strategy generation process. In this stage, the task queue length of each processing unit in the systolic array chip is monitored in real time, along with the current bandwidth status of the data transmission channels. For example, during peak diagnostic periods, some units may have tasks waiting to be processed, while some storage paths may experience channel bottlenecks. The system integrates these states with the newly generated data characteristics and the initial processing tasks to form a current operational status profile. This profile is then fed into the scheduling decision module, which, through its scheduling generation network, comprehensively considers factors such as resource idleness, task computation load, and data locality to generate task allocation schemes and data transmission paths. After generation, the system further verifies whether the task allocation results exceed the hardware resource load threshold and integrates the verified schemes to output the final scheduling strategy.

[0190] During the execution phase, sparse slices of CT image data, as determined by the encoding module, are sent to the nearest-memory storage unit. After compression using the sparse coding module, they are written to a multi-level cache system. The system moves the data within the storage structure according to the data path specified in the scheduling strategy and decompresses it in the nearest-memory processing unit. Subsequently, the system sends the decompressed data and neural network model instructions to the target processing units according to the task allocation scheme. In each processing unit, the previously generated set of valid data features is read to dynamically adjust the calculation accuracy mode. For example, low-precision fast processing is used for highly sparse images, while high-precision inference mode is activated for areas with abnormally high density, thereby activating the multiplication and accumulation unit in a differentiated manner to calculate and output the predicted lesion distribution map. Simultaneously, the system records performance indicators such as execution time, energy consumption, and prediction accuracy for each task, forming a performance dataset.

[0191] After completing the inference task, the system automatically enters the update process, using performance data to build a correlation model between the optimization parameters of the compilation module and the parameters generated by the scheduling strategy. An online learning module handles these two types of correlations, dynamically deriving the optimal update direction for the parameters. Subsequently, a test dataset is used in a simulated environment to verify whether updating the parameters improves inference performance. Once the improvement is confirmed to exceed a threshold, the system uses an atomic transaction mechanism to update the module parameters, ensuring uninterrupted system operation and recording the parameter version in the version management module for future traceability or rollback.

[0192] Ultimately, doctors receive the processing results on the diagnostic terminal interface, including the location of high-risk nodules in the lung slices, and combine this with the image hotspot heatmap returned by the system for further analysis and diagnosis. The entire process achieves closed-loop optimization from model task deployment, resource-aware scheduling, adaptive precision control, to inference result generation and parameter evolution, comprehensively improving the performance of the medical image processing system in terms of real-time performance, energy efficiency, and accuracy, and adapting to the realistic needs of complex data processing and resource constraints in healthcare scenarios.

[0193] In a real-time credit approval scenario for a commercial bank's risk control platform, the system needs to assess the risk level of received enterprise-level credit application data. To support in-depth modeling and feature interaction understanding of potential risk factors in credit data, the system loads a risk identification model jointly constructed based on graph neural networks (GNN) and time series analysis, and prepares to deploy it to a pulse array processing device for low-latency, high-throughput real-time inference.

[0194] First, the model's topology is parsed by the compilation module. This topology includes graph convolutional layers, attention-weighted layers, transaction sequence analysis paths, and a dynamic context fusion module. During topology analysis, redundant structures are identified among multiple transaction state evaluation nodes, which are processing units that can be merged. Through structural rewriting and connection mapping adjustments, a compact and clearly defined optimized topology is formed. The compilation module then calls the current row and column dimension configuration of the systolic array, the distribution of multiplication and accumulation units, and the local storage block strategy. Combining the dependency strength between nodes, a graph partitioning-based optimization node clustering algorithm is used to divide the processing structure into multiple compact, highly cohesive processing substructures. Further, an instruction set is generated and mapped to physical units using a genetic optimization method to construct the initial processing task.

[0195] Before processing, the system performs multimodal feature extraction on credit application data. Data sources include enterprise structure information, account transaction records, contract text summaries, and third-party credit scores, exhibiting strong sparsity and structural heterogeneity. A pre-trained multimodal feature extraction network extracts correlation density features from the structured data, semantic complexity features from the text, and constructs a sliding window fluctuation distribution index for time-series behavioral data. Each type of feature is assigned a corresponding influence factor; for example, complexity weights are added to high-frequency, small-amount transaction anomalies, generating a weighted feature set. After normalization, the system performs confidence interval verification and anomaly pruning, retaining only high-confidence features for scheduling inference.

[0196] In response to the current peak processing cycle of the bank's risk control system, the system monitors the real-time operating status of the deployed pulse array chips, dynamically collects the task queue length of each processing unit, and obtains the transmission bandwidth of multiple data channels. The system constructs an operational status profile, which, along with the task structure and data characteristics, is input into the scheduling decision module. The scheduling strategy model combines resource availability, task complexity, and the matching priority of data characteristics within a specific unit to output the optimal task allocation scheme and the data transmission scheme with low-conflict paths. If a resource constraint conflict is detected between a task and the allocated processing unit, the system automatically corrects the task mapping path and generates an executable scheduling strategy.

[0197] During the execution phase, the highly sparse dimensional portion of the financial data is first compressed in the nearest-memory unit. This involves sparse encoding of the sparse feature embedding matrix and storing it in a hierarchical caching system. Following a scheduling strategy, the system transmits the compressed data across multiple layers, including main memory, shared buffers, and nearest-memory units, and then decodes and restores it in the nearest-memory computing unit. Subsequently, the task logic unit receives the decompressed data and normalized feature set after the load instruction, adjusting the processing precision based on the data characteristics: mixed-precision computation is used for structures with stable data distribution but high model sensitivity, while low-precision computation is used to improve throughput for low-sensitivity or low-value input dimensions. Each unit then initiates a multiplication-accumulation module to perform model inference, outputting risk level labels, influence factor heatmaps, and confidence scores.

[0198] The system records the execution time, energy consumption, and prediction deviation for each task, generating a complete performance dataset. The online learning module then constructs two correlation models between the performance data and the compilation and scheduling parameters to evaluate the impact of this processing on hardware resource efficiency and inference accuracy. Parameter update vectors are output through incremental training, and their feasibility is verified using a test set. If performance improvement is significant, the system synchronously updates the control parameters of the compilation and scheduling modules through atomic operations and records the new version parameters to the version manager for future auditing, version rollback, and horizontal migration.

[0199] Ultimately, the system outputs enterprise credit risk level, credit limit suggestions, and transaction network visualization maps to the credit approval system, providing banks with risk control support capabilities that are highly real-time, robust, and capable of handling high concurrency. It enables the software and hardware collaborative deployment of complex network structure models in real-time financial technology business, significantly improving risk control efficiency and judgment accuracy, and effectively reducing business latency and computing resource waste.

[0200] This embodiment establishes a dual correlation between performance data and system module parameters, enabling the processing system to learn and adapt in real time. The compilation module optimizes processing unit allocation and code generation paths based on task execution results, while the scheduling decision module dynamically adjusts the scheduling algorithm according to runtime bottlenecks, thereby improving task execution efficiency and resource utilization. The online learning module ensures the continuity and adaptability of parameter updates, while the atomic transaction mechanism and version management module together constitute a consistency and traceability guarantee system during parameter evolution. Overall, this mechanism enables the processing system to continuously optimize its structure and behavior based on runtime feedback, significantly improving the system's computational performance and stability in heterogeneous task scenarios.

[0201] In one embodiment, a pulsating array scheduling processing device is provided, which corresponds one-to-one with the pulsating array scheduling processing method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the pulse array scheduling processing device of the present invention. The modules include a compilation module 10, a feature extraction module 20, a scheduling decision module 30, an execution module 40, and a feedback optimization module 50. Detailed descriptions of each functional module are as follows:

[0202] The compilation module 10 is used to acquire the data processing model and compile the data processing model into an initial processing task according to the hardware architecture characteristics of the pulsating array.

[0203] Feature extraction module 20 is used to acquire data to be processed and analyze the data features of the data to be processed through feature extraction module;

[0204] The scheduling decision module 30 is used to monitor the real-time operating status of the pulse array processing device, and input the real-time operating status, the data characteristics and the initial processing task into the scheduling decision module to generate a scheduling strategy that includes task allocation and data transmission path;

[0205] The execution module 40 is used to execute the scheduling strategy to drive the pulsating array to complete the processing of the data to be processed, generate the processing result, and collect the performance data during the processing.

[0206] The feedback optimization module 50 is used to update the optimization parameters of the compilation module and the strategy generation parameters of the scheduling decision module based on the performance data.

[0207] In one embodiment, the compilation module 10 is specifically used for:

[0208] Obtain the topology of the data processing model;

[0209] Identify redundant processing nodes in the topology;

[0210] Merge consecutive redundant processing nodes of the same type in the topology and update the topological connection relationship of the topology to generate an optimized topology;

[0211] Based on the cell layout of the pulsating array, the optimized topology is divided into multiple processing substructures;

[0212] Convert each processing substructure into a pulse array hardware instruction set;

[0213] Using a genetic algorithm task mapping strategy, the hardware instruction set is allocated to the multiplication and accumulation processing unit of the systolic array during the compilation phase to generate the initial processing task.

[0214] In one embodiment, the compilation module 10 is specifically used for:

[0215] Analyze the row and column dimension parameters of the cell layout of the pulsating array;

[0216] Analyze the strength of data dependencies between processing nodes in the optimized topology;

[0217] Based on the row and column dimension parameters and the processing capability parameters of the multiplication and accumulation units in the pulsating array, a node clustering algorithm is used to group the processing nodes in the optimized topology into multiple node groups.

[0218] Each node group is mapped to a processing substructure that adapts to the physical structure of the pulsating array based on the row and column dimension parameters.

[0219] Identify high-dependency processing nodes in the optimized topology whose data dependency strength exceeds a preset dependency threshold;

[0220] The boundary partitioning of the processing substructure is optimized to maintain the locality of highly dependent processing nodes;

[0221] Verify the balance of processing load distribution in each processing substructure;

[0222] When the load distribution balance is lower than the preset balance threshold, adjust the node group allocation and re-verify.

[0223] When the load distribution balance reaches the preset balance threshold, the final set of processing substructures is output.

[0224] In one embodiment, the feature extraction module 20 is specifically used for:

[0225] Load the pre-trained multimodal feature extraction model;

[0226] The data to be processed is input into the multimodal feature extraction model, and the complexity and sparsity features of the data to be processed are extracted by the multimodal feature extraction model, and the data distribution features of the data to be processed are quantified.

[0227] Assign a first influence factor, a second influence factor, and a third influence factor to the complexity feature, the sparsity feature, and the data distribution feature, respectively;

[0228] The complexity feature is adjusted using the first influencing factor to generate a weighted complexity feature;

[0229] The sparsity features are adjusted using the second influencing factor to generate weighted sparsity features;

[0230] The data distribution characteristics are adjusted using the third influencing factor to generate a weighted distribution characteristic;

[0231] The weighted complexity feature, weighted sparsity feature, and weighted distribution feature are integrated to generate a weighted feature set;

[0232] The weighted feature set is normalized to generate normalized feature values;

[0233] Verify the validity of the normalized feature values ​​within the confidence interval, filter out feature values ​​that exceed the confidence interval, and generate a set of valid feature values;

[0234] The data feature description is generated using the set of valid feature values.

[0235] In one embodiment, the scheduling decision module 30 is specifically used for:

[0236] Monitor the task queue length of each processing unit in the pulse array processing device;

[0237] Monitor the real-time bandwidth status of the data transmission channel in the pulse array processing device;

[0238] By integrating the task queue length and real-time bandwidth status, a real-time running status profile is generated.

[0239] The real-time running status profile, the data features, and the initial processing task are input into the scheduling decision module, and the scheduling decision model generates a task allocation scheme and a data transmission path.

[0240] Verify whether the task allocation scheme complies with the resource constraints of the pulse array processing device;

[0241] By integrating the validated task allocation scheme and the data transmission path, a scheduling strategy is generated.

[0242] In one embodiment, the execution module 40 is specifically used for:

[0243] In the near-memory storage unit, a sparse coding module is used to compress the high-sparse data in the data to be processed, generating compressed data;

[0244] The compressed data is stored in a multi-level storage structure;

[0245] The compressed data is transmitted in the multi-level storage structure according to the data transmission path in the scheduling strategy.

[0246] The compressed data is decompressed in the nearby storage unit to generate decompressed data.

[0247] According to the task allocation scheme in the scheduling strategy, the initial processing tasks are distributed to specific processing units of the pulsating array during the execution phase.

[0248] The decompressed data is input into the specific processing unit;

[0249] The data features are received in the specific processing unit;

[0250] The processing precision mode is adjusted according to the data characteristics in the specific processing unit.

[0251] The multiplication and accumulation unit in the specific processing unit is started with the adjusted processing precision mode, and the initial processing task is executed using the decompressed data to generate the processing result.

[0252] The execution time, energy consumption, and accuracy of the processing results of the data acquisition and processing tasks are used as performance data.

[0253] In one embodiment, the feedback optimization module 50 is specifically used for:

[0254] Establish a first correlation between the performance data and the optimization parameters of the compilation module;

[0255] Construct a second correlation between the performance data and the policy generation parameters of the scheduling decision module;

[0256] The first and second association relationships are processed through the online learning module to generate update parameters for the compilation module and update parameters for the scheduling decision module.

[0257] Verify the performance improvement of the compilation module's parameter update and the scheduling decision module's parameter update on the test dataset;

[0258] When the performance improvement exceeds a preset threshold, the optimization parameters of the compilation module are updated to update the parameters of the compilation module through atomic transactions, and the strategy generation parameters of the scheduling decision module are updated to update the parameters of the scheduling decision module through atomic transactions.

[0259] Record the updated parameter version information to the version management module.

[0260] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a pulse array scheduling processing method on the server side.

[0261] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a pulse array scheduling processing method on the user side.

[0262] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0263] A data processing model is obtained, and the data processing model is compiled into an initial processing task according to the hardware architecture characteristics of the pulsating array through a compilation module.

[0264] Acquire the data to be processed, and analyze the data features of the data to be processed through the feature extraction module;

[0265] Monitor the real-time operating status of the pulse array processing device, and input the real-time operating status, the data characteristics, and the initial processing task into the scheduling decision module to generate a scheduling strategy that includes task allocation and data transmission path;

[0266] The scheduling strategy is executed to drive the pulsating array to process the data to be processed, generate processing results, and collect performance data during the processing.

[0267] Based on the performance data, update the optimization parameters of the compilation module and the strategy generation parameters of the scheduling decision module.

[0268] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0269] A data processing model is obtained, and the data processing model is compiled into an initial processing task according to the hardware architecture characteristics of the pulsating array through a compilation module.

[0270] Acquire the data to be processed, and analyze the data features of the data to be processed through the feature extraction module;

[0271] Monitor the real-time operating status of the pulse array processing device, and input the real-time operating status, the data characteristics, and the initial processing task into the scheduling decision module to generate a scheduling strategy that includes task allocation and data transmission path;

[0272] The scheduling strategy is executed to drive the pulsating array to process the data to be processed, generate processing results, and collect performance data during the processing.

[0273] Based on the performance data, update the optimization parameters of the compilation module and the strategy generation parameters of the scheduling decision module.

[0274] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0275] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0276] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0277] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for scheduling and processing pulsed arrays, characterized in that, Includes the following steps: A data processing model is obtained, and the data processing model is compiled into an initial processing task according to the hardware architecture characteristics of the pulsating array through a compilation module. Acquire the data to be processed, and analyze the data features of the data to be processed through the feature extraction module; Monitor the real-time operating status of the pulse array processing device, and input the real-time operating status, the data characteristics, and the initial processing task into the scheduling decision module to generate a scheduling strategy that includes task allocation and data transmission path; The scheduling strategy is executed to drive the pulsating array to process the data to be processed, generate processing results, and collect performance data during the processing. Based on the performance data, update the optimization parameters of the compilation module and the strategy generation parameters of the scheduling decision module.

2. The pulsating array scheduling processing method as described in claim 1, characterized in that, Obtain the data processing model, and compile the data processing model into an initial processing task based on the hardware architecture characteristics of the systolic array using a compilation module, including: Obtain the topology of the data processing model; Identify redundant processing nodes in the topology; Merge consecutive redundant processing nodes of the same type in the topology and update the topological connection relationship of the topology to generate an optimized topology; Based on the cell layout of the pulsating array, the optimized topology is divided into multiple processing substructures; Convert each processing substructure into a pulse array hardware instruction set; Using a genetic algorithm task mapping strategy, the hardware instruction set is allocated to the multiplication and accumulation processing unit of the systolic array during the compilation phase to generate the initial processing task.

3. The pulsating array scheduling processing method as described in claim 2, characterized in that, Based on the cell layout of the pulsating array, the optimized topology is divided into multiple processing substructures, including: Analyze the row and column dimension parameters of the cell layout of the pulsating array; Analyze the strength of data dependencies between processing nodes in the optimized topology; Based on the row and column dimension parameters and the processing capability parameters of the multiplication and accumulation units in the pulsating array, a node clustering algorithm is used to group the processing nodes in the optimized topology into multiple node groups. Each node group is mapped to a processing substructure that adapts to the physical structure of the pulsating array based on the row and column dimension parameters. Identify high-dependency processing nodes in the optimized topology whose data dependency strength exceeds a preset dependency threshold; The boundary partitioning of the processing substructure is optimized to maintain the locality of highly dependent processing nodes; Verify the balance of processing load distribution in each processing substructure; When the load distribution balance is lower than the preset balance threshold, adjust the node group allocation and re-verify. When the load distribution balance reaches the preset balance threshold, the final set of processing substructures is output.

4. The pulsating array scheduling processing method as described in claim 1, characterized in that, Acquire the data to be processed, and analyze the data features of the data to be processed through the feature extraction module, including: Load the pre-trained multimodal feature extraction model; The data to be processed is input into the multimodal feature extraction model, and the complexity and sparsity features of the data to be processed are extracted by the multimodal feature extraction model, and the data distribution features of the data to be processed are quantified. Assign a first influence factor, a second influence factor, and a third influence factor to the complexity feature, the sparsity feature, and the data distribution feature, respectively; The complexity feature is adjusted using the first influencing factor to generate a weighted complexity feature; The sparsity features are adjusted using the second influencing factor to generate weighted sparsity features; The data distribution characteristics are adjusted using the third influencing factor to generate a weighted distribution characteristic; The weighted complexity feature, weighted sparsity feature, and weighted distribution feature are integrated to generate a weighted feature set; The weighted feature set is normalized to generate normalized feature values; Verify the validity of the normalized feature values ​​within the confidence interval, filter out feature values ​​that exceed the confidence interval, and generate a set of valid feature values; The data feature description is generated using the set of valid feature values.

5. The pulsating array scheduling processing method as described in claim 1, characterized in that, The real-time operating status of the pulse array processing device is monitored, and the real-time operating status, the data characteristics, and the initial processing task are input into the scheduling decision module to generate a scheduling strategy that includes task allocation and data transmission paths, including: Monitor the task queue length of each processing unit in the pulse array processing device; Monitor the real-time bandwidth status of the data transmission channel in the pulse array processing device; By integrating the task queue length and real-time bandwidth status, a real-time running status profile is generated. The real-time running status profile, the data features, and the initial processing task are input into the scheduling decision module, and the scheduling decision model generates a task allocation scheme and a data transmission path. Verify whether the task allocation scheme complies with the resource constraints of the pulse array processing device; By integrating the validated task allocation scheme and the data transmission path, a scheduling strategy is generated.

6. The pulsating array scheduling processing method as described in claim 1, characterized in that, The scheduling strategy is executed to drive the pulsating array to process the data to be processed, generate processing results, and collect performance data during the processing, including: In the near-memory storage unit, a sparse coding module is used to compress the high-sparse data in the data to be processed, generating compressed data; The compressed data is stored in a multi-level storage structure; The compressed data is transmitted in the multi-level storage structure according to the data transmission path in the scheduling strategy. The compressed data is decompressed in the nearby storage unit to generate decompressed data. According to the task allocation scheme in the scheduling strategy, the initial processing tasks are distributed to specific processing units of the pulsating array during the execution phase. The decompressed data is input into the specific processing unit; The data features are received in the specific processing unit; The processing precision mode is adjusted according to the data characteristics in the specific processing unit. The multiplication and accumulation unit in the specific processing unit is started with the adjusted processing precision mode, and the initial processing task is executed using the decompressed data to generate the processing result. The execution time, energy consumption, and accuracy of the processing results of the data acquisition and processing tasks are used as performance data.

7. The pulsating array scheduling processing method as described in claim 1, characterized in that, Based on the performance data, update the optimization parameters of the compilation module and the policy generation parameters of the scheduling decision module, including: Establish a first correlation between the performance data and the optimization parameters of the compilation module; Construct a second correlation between the performance data and the policy generation parameters of the scheduling decision module; The first and second association relationships are processed through the online learning module to generate update parameters for the compilation module and update parameters for the scheduling decision module. Verify the performance improvement of the compilation module's parameter update and the scheduling decision module's parameter update on the test dataset; When the performance improvement exceeds a preset threshold, the optimization parameters of the compilation module are updated to update the parameters of the compilation module through atomic transactions, and the strategy generation parameters of the scheduling decision module are updated to update the parameters of the scheduling decision module through atomic transactions. Record the updated parameter version information to the version management module.

8. A pulse array scheduling processing device, characterized in that, The pulse array scheduling processing device includes: The compilation module is used to obtain the data processing model and compile the data processing model into an initial processing task according to the hardware architecture characteristics of the systolic array. The feature extraction module is used to acquire the data to be processed and analyze the data features of the data to be processed. The scheduling decision module is used to monitor the real-time operating status of the pulse array processing device, and input the real-time operating status, the data characteristics and the initial processing task into the scheduling decision module to generate a scheduling strategy that includes task allocation and data transmission path; The execution module is used to execute the scheduling strategy to drive the pulsating array to complete the processing of the data to be processed, generate the processing result, and collect the performance data during the processing. The feedback optimization module is used to update the optimization parameters of the compilation module and the strategy generation parameters of the scheduling decision module based on the performance data.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a systolic array scheduling process stored in the memory and executable on the processor, wherein when executed by the processor, the systolic array scheduling process implements the steps of the systolic array scheduling method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a pulsating array scheduling processing program, which, when executed by a processor, implements the steps of the pulsating array scheduling processing method as described in any one of claims 1-7.

Citation Information

Cited By

  • Fine-grained quantization matrix multiplication device and method based on systolic array

    CN121479115A