In-memory computing integrated parallel processing system and method
Through the integrated parallel processing method of storage and computing, the problems of high latency and inefficiency under the separation of storage and computing architecture are solved, and data processing is efficient and intelligent, the accuracy of data analysis and system response speed are improved, and the utilization rate of storage and computing resources is optimized.
Patent Information
- Application Number
- CN202510653728.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing separation of storage and computing architecture leads to high latency between data storage and computing, which is inefficient, unable to meet real-time computing needs, lack of effective feature fusion and optimization strategies, and difficult to adapt to the rapidly changing data environment, resulting in limited accuracy and efficiency of data analysis, and traditional methods cannot achieve efficient flow control and cache integration, resulting in access delays and waste of resources.
By obtaining the original data for topological skeleton projection, generating multi-scale spatial fusion data, determining the weight relationship matrix for gradient optimization and adjustment, mining the memory unit activates the chain interaction between mapped data, performing recursive self-calibration and predicting heterogeneous access channels, reconstructing the data access link, performing multi-layer collaborative flow control cache integration and task reorganization, and performing flow scheduling to eliminate parallel conflicts.
It realizes the efficiency and intelligence of data processing, improves the processingability and accuracy of data, optimizes the access efficiency, reduces resource waste, improves the system's response speed and overall processing capabilities, enhances the system's flexibility and adaptability, and promotes the improvement of data processing capabilities in a high-performance computing environment.
Smart Images

Figure CN120179606B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to an in-memory computing integrated parallel processing system and method. Background Art
[0002] The architecture of separating storage and computing leads to high latency between data storage and computing. Existing methods often face the problem of low efficiency when dealing with complex data. Traditional storage systems cannot meet the requirements of real-time computing. Data is prone to bottlenecks during transmission, affecting the overall performance and response speed. Especially when dealing with multi-scale data, there is a lack of effective feature fusion and optimization strategies, which limits the accuracy and efficiency of data analysis. In addition, existing technologies have deficiencies in dynamic data self-calibration and chained interaction mining, and cannot make full use of the interconnection relationship between in-memory computing units, resulting in the ineffective mining of the correlation between data. The implementation of recursive self-calibration is difficult, and there is a lack of a flexible adjustment mechanism, making it difficult to adapt to the rapidly changing data environment. Especially in multi-level data access, traditional methods often cannot achieve efficient flow control and cache integration, resulting in access latency and resource waste. Summary of the Invention
[0003] Based on this, it is necessary to provide an in-memory computing integrated parallel processing system and method to solve at least one of the above technical problems.
[0004] To achieve the above object, the in-memory computing integrated parallel processing method includes the following steps:
[0005] Step S1: Obtain the original stored data; perform topological skeleton projection on the original stored data to obtain a projected skeleton structure; perform feature hierarchical integration on the original stored data according to the projected skeleton structure to generate multi-scale spatial fusion data;
[0006] Step S2: Determine the weight relationship matrix based on the multi-scale spatial fusion data; perform gradient optimization adjustment on the multi-scale spatial fusion data through the weight relationship matrix, and perform memory-computing unit matching to obtain multiple in-memory computing unit activation mapping data;
[0007] Step S3: Mine the chained interaction correlation data between the in-memory computing unit activation mapping data; perform recursive self-calibration on the original stored data according to the chained interaction correlation data to generate self-calibration updated data;
[0008] Step S4: Predict the heterogeneous access channels based on the self-calibration updated data, and reconstruct each hierarchical data access link through the predicted heterogeneous access channels; perform multi-level collaborative flow control cache integration according to the reconstructed each hierarchical data access link to obtain access latency elimination and optimization data;
[0009] Step S5: Reorganize the dependency path tasks for the access latency elimination optimized data to generate task reorganization parallel processing data; perform pipelining scheduling on the task reorganization parallel processing data to execute the elimination of parallel conflicts, thereby obtaining parallel conflict decoupled data.
[0010] The present invention generates a projected skeleton structure by obtaining the original stored data and performing topological skeleton projection, which provides a clear spatial framework for subsequent data processing. The feature hierarchical integration realizes multi-dimensional analysis of the original data. The generated multi-scale spatial fusion data improves the processability and accuracy of the data while retaining important features. The weight relationship matrix determined based on the multi-scale spatial fusion data provides a scientific basis for data optimization. The gradient optimization adjustment makes the matching between calculation and storage more efficient. The generation of activation mapping data of multiple memory and computing units lays the foundation for parallel processing. Mining the chain interaction correlation data between the activation mapping data of memory and computing units further improves the correlation and processing efficiency of the data. The recursive self-calibration can dynamically adjust the original stored data, and the generated self-calibration updated data ensures the consistency and accuracy of the data. Predicting the heterogeneous access channels based on the self-calibration updated data and reconstructing each hierarchical data access link optimize the data flow path. The reconstructed multi-layer collaborative flow control cache integration effectively eliminates the access latency. The generated access latency elimination optimized data reduces resource waste while improving the access efficiency. The dependency path task reorganization creates conditions for the parallel processing of tasks. The generated task reorganization parallel processing data improves the system's response speed during execution. The pipelining scheduling mechanism effectively eliminates conflicts in parallel processing. The finally obtained parallel conflict decoupled data improves the overall processing capacity of the system. The overall method realizes the high-efficiency and intelligence of data processing under the framework of memory and computing integration, provides reliable support for the execution of complex computing tasks, promotes the development of new memory and computing architectures, enhances the flexibility and adaptability of the system, promotes the improvement of data processing capabilities in high-performance computing environments, optimizes the utilization rate of storage and computing resources, and provides new ideas and technical paths for future multi-task parallel processing.
[0011] The present invention also provides a memory and computing integrated parallel processing system for executing the memory and computing integrated parallel processing method as described above. The memory and computing integrated parallel processing system includes:
[0012] A multi-scale feature fusion module for obtaining the original stored data; performing topological skeleton projection on the original stored data to obtain a projected skeleton structure; and hierarchically integrating the features of the original stored data according to the projected skeleton structure to generate multi-scale spatial fusion data;
[0013] The weight matching activation module is used to determine the weight relationship matrix based on the multi-scale spatial fusion data; optimize and adjust the gradient of the multi-scale spatial fusion data through the weight relationship matrix, and perform memory-computation unit matching, so as to obtain multiple memory-computation unit activation mapping data;
[0014] The interactive chain self-calibration module is used to mine the chain interactive correlation data between the memory-computation unit activation mapping data; recursively self-calibrate the original stored data according to the chain interactive correlation data, and generate self-calibration update data;
[0015] The channel reconstruction flow control integration module is used to predict the heterogeneous access channels based on the self-calibration update data, and reconstruct each hierarchical data access link through the predicted heterogeneous access channels; perform multi-layer collaborative flow control cache integration according to the reconstructed hierarchical data access links, so as to obtain access latency elimination optimization data;
[0016] The task decoupling scheduling module is used to reorganize the dependent path tasks of the access latency elimination optimization data to generate task reorganization parallel processing data; perform pipelining scheduling on the task reorganization parallel processing data to execute and eliminate parallel conflicts, so as to obtain parallel conflict decoupling data.
[0017] Through the multi-scale feature fusion module, the present invention combines the acquisition of the originally stored data with the topological skeleton projection to form a clear projection skeleton structure. Based on this, the hierarchical integration of features realizes the multi-dimensional analysis of data. The generated multi-scale spatial fusion data improves the data processing efficiency while retaining the key features. The weight matching activation module provides a scientific basis for the gradient optimization of data based on the weight relationship matrix generated from the multi-scale spatial fusion data. The effective matching of the memory and computing units improves the resource utilization rate. The activation mapping data of multiple memory-computation units realizes the basis of parallel processing through optimization and adjustment. The interactive chain self-calibration module enhances the correlation between data by mining the chain interaction correlation data between the activation mapping data of the memory-computation units. The updated data generated by recursive self-calibration ensures the consistency and accuracy of the original data. The channel reconstruction flow control integration module predicts the heterogeneous access channels based on the self-calibration updated data. The reconstructed hierarchical data access links optimize the data flow path. The multi-layer collaborative flow control cache integration formed after reconstruction effectively eliminates the access delay. The generated optimized data significantly improves the access efficiency. The task decoupling scheduling module creates conditions for parallel processing by reorganizing the dependent path tasks of the optimized data with eliminated access delay. The generated task reorganization parallel processing data improves the system response speed during execution. The pipeline scheduling mechanism effectively eliminates the conflicts in parallel processing. The finally obtained parallel conflict decoupling data improves the processing capacity of the overall system. The entire system realizes efficient and intelligent data processing within the framework of memory-computation integration, provides reliable support for the execution of complex computing tasks, promotes the development of new memory-computation architectures, enhances the flexibility and adaptability of the system, optimizes the utilization of storage and computing resources, provides an innovative technical path for future multi-task parallel processing. The high efficiency and accuracy of the overall process lay a solid foundation for the application in large-scale data processing scenarios, promote the improvement of data processing capabilities in high-performance computing environments, and drive the wide application and development of memory-computation integration technology. Brief Description of the Drawings
[0018] Figure 1 It is a schematic diagram of the step flow of a memory-computation integration parallel processing method;
[0019] Figure 2 It is a schematic diagram of the detailed implementation step flow of step S2;
[0020] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments
[0021] The technical method of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0022] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0023] It should be understood that although terms such as "first" and "second" may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed associated items.
[0024] To achieve the above object, please refer to Figures 1 to 2 , a memory-computation integrated parallel processing method, including the following steps:
[0025] Step S1: Obtain the original stored data; perform topological skeleton projection on the original stored data to obtain a projected skeleton structure; perform feature hierarchical integration on the original stored data according to the projected skeleton structure, thereby generating multi-scale spatial fusion data;
[0026] Step S2: Determine the weight relationship matrix based on the multi-scale spatial fusion data; perform gradient optimization adjustment on the multi-scale spatial fusion data through the weight relationship matrix, and perform memory-computation unit matching, thereby obtaining multiple memory-computation unit activation mapping data;
[0027] Step S3: Mine the chain interaction correlation data between the memory-computation unit activation mapping data; perform recursive self-calibration on the original stored data according to the chain interaction correlation data, and generate self-calibration update data;
[0028] Step S4: Predict the heterogeneous access channels based on the self-calibration updated data, and reconstruct each hierarchical data access link through the predicted heterogeneous access channels; perform multi-layer collaborative flow control cache integration according to the reconstructed hierarchical data access links, so as to obtain access latency elimination optimized data;
[0029] Step S5: Reorganize the dependent path tasks for the access latency elimination optimized data to generate task reorganization parallel processing data; perform pipelining scheduling on the task reorganization parallel processing data to execute and eliminate parallel conflicts, so as to obtain parallel conflict decoupled data.
[0030] The present invention projects the original stored data through obtaining the original stored data and performing topological skeleton projection, and the generated projection skeleton structure provides a clear spatial framework for subsequent data processing. The feature hierarchical integration realizes the multi-dimensional analysis of the original data. The generated multi-scale spatial fusion data improves the processability and accuracy of the data while retaining important features. The weight relationship matrix determined based on the multi-scale spatial fusion data provides a scientific basis for data optimization. The gradient optimization adjustment makes the matching between calculation and storage more efficient. The generation of activation mapping data of multiple memory-computation units lays the foundation for parallel processing. Mining the chain interaction correlation data between the activation mapping data of memory-computation units further improves the correlation and processing efficiency between data. The recursive self-calibration can dynamically adjust the original stored data, and the generated self-calibration updated data ensures the consistency and accuracy of the data. Predicting the heterogeneous access channels based on the self-calibration updated data and reconstructing each hierarchical data access link optimizes the data flow path. The reconstructed multi-layer collaborative flow control cache integration effectively eliminates the access latency. The generated access latency elimination optimized data reduces resource waste while improving the access efficiency. The dependent path task reorganization creates conditions for the parallel processing of tasks. The generated task reorganization parallel processing data improves the system's response speed during execution. The pipelining scheduling mechanism effectively eliminates conflicts in parallel processing. The finally obtained parallel conflict decoupled data improves the overall processing ability of the system. The overall method realizes the high-efficiency and intelligence of data processing under the framework of memory-computation integration, provides reliable support for the execution of complex computing tasks, promotes the development of new memory-computation architectures, enhances the flexibility and adaptability of the system, promotes the improvement of data processing ability in a high-performance computing environment, optimizes the utilization rate of storage and computing resources, and provides new ideas and technical paths for future multi-task parallel processing.
[0031] In the embodiment of the present invention, the memory-computation integrated parallel processing method includes the following steps:
[0032] Step S1: Obtain the original stored data; perform topological skeleton projection on the original stored data to obtain a projection skeleton structure; perform feature hierarchical integration on the original stored data according to the projection skeleton structure, so as to generate multi-scale spatial fusion data;
[0033] In this embodiment, when obtaining the original stored data, a high-speed acquisition interface based on the DDR4 bus communication protocol is used to access the data stream input by heterogeneous devices to the FPGA (Field Programmable Gate Array) buffer control module. The data sampling period is set to 5 nanoseconds, and batch data collection is performed through a 64-bit wide bus. The collected data is organized into original data blocks in units of 128 bytes and accessed to the internal buffer of the FPGA using the AXI bus protocol. The buffer is configured as a dual-port FIFO structure to maintain the stability of continuous data input. When performing topological skeleton projection on the data blocks in the buffer, the first principal axis direction of each data block in the three-dimensional feature space is extracted through three-dimensional principal component analysis PCA (Principal Component Analysis), and this principal axis direction is used as the projection principal vector. All data points are projected along this direction and their two-dimensional mappings on the orthogonal plane are recorded to form an initial skeleton graph structure. Based on this skeleton structure, a feature hierarchy is constructed. The original projection graph is downsampled 3 times spatially using the multi-scale Gaussian pyramid method, with a downsampling factor of 2 for each layer, and local geometric feature points at each scale are extracted respectively. The feature maps at each scale are combined into a three-dimensional feature tensor by constructing local extreme points in the scale space. Finally, this tensor is point-to-point mapped with the original data block to construct multi-scale space fusion data, and this data structure records the spatial positions, relative skeleton distances, scale levels where they are located, and the index relationship with the original data of the feature points at each level according to the index.
[0034] Step S2: Determine the weight relationship matrix based on the multi-scale space fusion data; perform gradient optimization adjustment on the multi-scale space fusion data through the weight relationship matrix, and perform memory-computation unit matching, so as to obtain activation mapping data of multiple memory-computation units;
[0035] In this embodiment, when determining the weight relationship matrix based on the multi-scale spatial fusion data constructed above, the sparse tensor encoding method is used to encode the feature point information at each scale into sparse vectors, and the similarity score is calculated through a dual index based on the spatial angle between points and the relative scale difference. The similarity score formula uses the Euclidean distance plus the scale ratio to form a mixed weight function, where the weighting coefficient of the distance term is set to 0.7, and the weighting coefficient of the scale difference is set to 0.3. Finally, the mixed weight values between all pairs of feature points are recorded in the sparse matrix to form the weight relationship matrix. The dimension of this matrix is N×N, where N is the total number of all valid feature points in the fusion data. When performing gradient optimization adjustment on the fusion data based on this matrix, the L-BFGS (Limited-memory Broyden-Fletcher-Goldfarb-Shanno) algorithm is used for constrained optimization. The objective function is set to minimize the sum of the squares of the feature point reconstruction errors. The initial value of the gradient step size is set to 0.01, and the maximum number of iterations is set to 200 times. Through iterative convergence, the optimized multi-scale fusion tensor is obtained. When performing the memory-computation unit matching operation on this tensor, the fusion data is grouped and sorted according to the spatial density and access frequency. The sorting basis is the number of times each feature point appears in the original data index plus its depth value in the skeleton structure. Finally, the top K high-weight points are selected and mapped to the SRAM. The value of K is set to 128, and the corresponding mapping operation is completed using the DMA controller. After the SRAM is bound to the FPGA computing core, all the mapped data is converted into the active state, and the active mapping data of the memory-computation unit is generated, which includes the active index table, the data access frequency table, and the tensor mapping pointer structure.
[0036] Step S3: Mine the chained interaction correlation data between the active mapping data of the memory-computation unit; perform recursive self-calibration on the original stored data according to the chained interaction correlation data to generate self-calibration update data;
[0037] In this embodiment, when mining the chain interaction correlation data based on the activation mapping data of the memory-computation unit, a graph convolutional network (GCN) structure is used for processing. First, all nodes in the activation mapping data are constructed into an undirected graph structure. The edge weights between nodes are set according to the non-zero terms in the weight relationship matrix generated in the previous stage. The number of interactions attached to each edge is obtained through a time-series access simulation experiment. The simulation process is set to continuously write and read 128 times, record the access switching frequency between adjacent nodes, and use it as the interaction strength parameter of the edge. Based on this graph structure, a 3-layer GCN network is built, with the output dimension of each layer being 64. The ReLU (Rectified Linear Unit) is used as the activation function. By training the graph network model, a high-interaction-frequency chain structure is identified. The chain structure is defined as a node sequence with continuous access dependency relationships. After the GCN output, all node sequences in the interaction chain are extracted as chain interaction correlation data, and at the same time, their interaction directions, access frequencies, and node-to-node time delays are recorded. When performing recursive self-calibration on the original stored data based on this interaction correlation structure, a multi-round reverse mapping strategy is adopted. The chain sequence is mapped back to the original data storage structure, and the arrangement order of the data in the storage block is readjusted according to the interaction direction. Consistency verification is performed for each adjustment. The verification criterion is that the access interval of three consecutive groups of data in the same interaction chain does not exceed two memory access cycles. If it does not meet the criterion, a data exchange operation is performed to complete the rearrangement within the entire data block, and the updated data structure is marked as self-calibration updated data. Each data block is appended with an interaction chain sequence index and a calibration timestamp.
[0038] Step S4: Predict heterogeneous access channels based on the self-calibration updated data, and reconstruct each hierarchical data access link through the predicted heterogeneous access channels; perform multi-layer collaborative flow control cache integration according to the reconstructed hierarchical data access links, so as to obtain access latency elimination and optimized data;
[0039] In this embodiment, when constructing a heterogeneous access channel prediction model based on the generated self-calibration update data, the long short-term memory network (LSTM, Long Short-Term Memory) is used to learn the access history sequence of the self-calibration data. The input sequence is set as the access interval sequence of each 128-byte data block and the type label of the interaction chain. The LSTM structure is set with two hidden layers, each with 128 neurons, trained using the RMSprop optimizer, and the loss function uses the mean squared error. The training data comes from 512 groups of typical self-calibration sequences and their historical access paths. The prediction result is the heterogeneous channel type number to be assigned to each data block. The type numbers include high-throughput SRAM channels, low-latency DRAM channels, and large-capacity Flash channels. According to the prediction result, the original hierarchical data access link is reconstructed, and the data blocks are respectively assigned to the corresponding access paths according to the predicted channel types. When constructing the new link structure, a hash index mechanism is used to attach the channel type identifier and the channel start address to each data block. After completing the link reconstruction, a multi-level cooperative flow control cache integration operation is performed to impose flow control consistency constraints on each cache layer in the reconstructed access chain. Each cache uses the LRU (Least Recently Used) algorithm for page table management. The cache page size is fixed at 4KB, and the write-back policy between cache layers is set so that the write-back latency does not exceed two scheduling cycles. Finally, the access structure after cache integration is generated and marked as access latency elimination optimization data. This data structure includes a channel allocation table, an access latency distribution matrix, and a cache synchronization status record.
[0040] Step S5: Reorganize the dependent path tasks for the access latency elimination optimization data to generate task reorganization parallel processing data; perform pipelining scheduling on the task reorganization parallel processing data to execute and eliminate parallel conflicts, thereby obtaining parallel conflict decoupled data.
[0041] In this embodiment, when performing dependency path task recombination on access latency elimination optimized data, first, the channel mapping order and cache synchronization status in the optimized data are extracted to construct a directed graph structure for task execution. Each node in the graph represents a computational task for a data block, and the edges represent its dependency relationships and channel transmission order. The Topological Sort algorithm is used to rearrange the graph structure to ensure that all data blocks are reordered according to the shortest latency path. After recombination, task recombination parallel processing data is generated, which is saved in the form of an execution flow array. Each array element contains a computational task index, the identifier of the dependent task, the required channel number, and the task trigger latency. Subsequently, a pipelining scheduling operation is performed on this parallel processing data. The scheduling process adopts a priority-based polling strategy, where the priority is determined by the criticality of the path where the task is located (i.e., the node with the maximum path latency). One scheduling window is allocated per clock cycle, and the window size is 8 tasks. The pipelining scheduler sequentially schedules non-conflicting tasks and sends them to the parallel computing array. If two tasks access the same channel or the dependency relationship is not satisfied, they are postponed to the next cycle. All scheduling conflict situations are recorded in the conflict matrix, and resource decoupling is performed according to the conflict matrix. The decoupling operation is to introduce an intermediate buffer for tasks and re-calibrate the task trigger conditions. After completion, the final parallel conflict decoupled data is output, and the data format is a task execution sequence, a scheduling cycle comparison table, and a resource isolation flag index table.
[0042] Preferably, step S1 includes the following steps:
[0043] Step S11: Perform channel domain partitioning processing on the originally stored data to obtain channel distribution segment data, where the channel partitioning granularity range is set to pixel block dimensions from 8×8 to 64×64;
[0044] Step S12: Perform local topological tensor mapping processing on the channel distribution segment data to generate basic topological mapping data, where the mapping dimension is set to a 3rd-order tensor structure, and the element dimension range is 16 - 512;
[0045] Step S13: Perform skeletonized structure projection on the basic topological mapping data to obtain a projected skeleton structure, where the structure connectivity threshold is set to 0.35 - 0.65, and the skeleton sparsity rate is controlled within 15% - 40%;
[0046] Step S14: Perform scale weight coupling according to the projected skeleton structure and nest it into a skeleton scale label;
[0047] Step S15: Perform hierarchical feature compression and fusion on the originally stored data according to the skeleton scale label to generate hierarchical feature integration data, where the fusion depth is set to 4 to 12 layers, and the compression ratio is controlled within 1:4 - 1:12;
[0048] Step S16: Perform spatial redundancy regularization processing on the hierarchical feature integration data to obtain multi-scale spatial fusion data, where the redundancy regularization window size is set to 5×5 to 11×11, and the residual filtering threshold is set to 0.02 - 0.12.
[0049] In this embodiment, the operation process of dividing the original stored data in the channel domain is as follows: First, the original stored data is loaded into the data acquisition unit. The data is expressed in the form of a three-dimensional matrix with a dimension size of 256×256×128. The first two dimensions correspond to the spatial range of the image, and the third dimension is the number of channels. During the division process, a fixed sliding window operation is used to set each scanned block as a pixel block of 32×32. The channel division is sliced into groups of 16 channels each, and the segmentation and cropping in the spatial and channel dimensions are completed in sequence, thereby obtaining the channel distribution segment data. A total of 64 spatial region blocks are generated, each region corresponding to 8 groups of channel segments. The data dimension of each group of channel segments is 32×32×16. By using the region aggregation method based on channel energy sorting, the channels with similar signal-to-noise ratios and intensity distributions are combined into the same segment, and the average channel density distribution matrix is recorded on each group of segments as the subsequent mapping basis. The channel distribution segment data is input into the tensor processing module, and a tensor representation is constructed for each segment using a three-dimensional tensor reconstruction engine. The tensor dimension is set to a third-order structure, that is, it contains information in three directions: horizontal, vertical, and channel. The unit value of each tensor is derived from the normalization result of the original channel pixel values. The minimum value of each dimension in the tensor is set to 16, and the maximum value does not exceed 512. After tensor assembly, a homogeneous filling operation is performed on the tensor boundary by the structure regularization module to ensure the consistency of the form and size between each group of data. The unified output size is set to 32×32×64, and this structure constitutes the basic topological mapping data. The basic topological mapping data is input into the skeletonization structure module, and the main connected path in the tensor structure is extracted using the structure path extraction method. First, a channel structure diagram is generated by the adjacency degree calculation method, and then according to the set structure connectivity threshold of 0.5. Delete all weak connection paths, only retain the connection relationships between strongly connected channels, further set the skeleton sparsity rate to 25%. During the entire sparse processing process, sort the path priorities according to the pixel gradient intensity inside each tensor block, only retain the top 75% of the connection nodes in the gradient sorting, and delete the remaining paths to form a sparse skeleton structure. This structure is output in a two-dimensional projection manner, and the dimension is maintained within the same spatial range as the input tensor. Send the obtained projected skeleton structure into the scale coupling module for scale weight annotation processing, use a multi-scale fusion strategy to construct skeleton scale labels, and superimpose the corresponding structure complexity labels on each skeleton node. This label is derived from the channel activity statistics at different resolutions. Set three groups of scale ranges corresponding to small-scale, medium-scale, and large-scale structures respectively, and determine the label weights of each node according to the probability distribution of the skeleton appearing in these three types of structures. Finally, reorganize the skeleton structure into a nested structure according to the node scale labels, and form the output data of the skeleton scale label containing three types of weight information. Send the skeleton scale label into the feature fusion module, combine the original stored data to construct a multi-layer fusion structure, use a hierarchical convolution compression processing method, and group and compress the original data channels step by step according to the skeleton scale label. The fusion process is set to 8 layers deep. The first layer processes 128-channel original data, and each subsequent layer further aggregates its compression result. The compression ratio of each layer is 1:4, 1:5, 1:6 until 1:12 in turn. After each layer of compression, use a multi-channel attention mechanism to weightedly integrate the outputs of each compression layer to generate the final hierarchical feature integration data. The dimension of this data is 128×128×48, retaining the main feature channel information under different scale skeletons. Send the hierarchical feature integration data into the redundancy regularization module for multi-scale spatial regularization operations. Use a sliding window method to scan each region in the spatial dimension. Set the window size to 9×9 and the sliding step size to 3 each time. Calculate the residual redundancy value within each window region, mark the redundant part lower than the set threshold of 0.08 as an invalid region, and use a low-rank filling strategy to recover and compress this region, retaining the backbone feature data with a lower redundancy. After completing the redundancy regularization processing, output the final multi-scale spatial fusion data. This data serves as the input basis data for subsequent weight relationship matrix extraction and activation mapping calculation, and the dimension remains 128×128×48, with the redundancy controlled below 5% of the total feature volume.
[0050] Preferably, step S2 includes the following steps:
[0051] Step S21: Perform memory mapping expansion on the multi-scale spatial fusion data to obtain the memory mapping structure expansion data;
[0052] Step S22: Identify the patterns of the memory mapping structure expansion data and classify the identified patterns according to similarity to obtain the classified spatial feature patterns;
[0053] Step S23: Construct a weight matrix based on the classified spatial feature pattern to generate a weight relationship matrix; perform gradient optimization adjustment on the weight relationship matrix and perform structured sparsity projection to obtain structured sparse mapping data;
[0054] Step S24: Perform memory addressing optimization based on the structured sparse mapping data and calculate the unit activation mapping process based on the addressed optimization data to obtain multiple memory-computation unit activation mapping data.
[0055] In this embodiment, the specific operation of memory mapping expansion of multi-scale spatial fusion data is to load the multi-scale spatial fusion data output in the previous step into the mapping control module, the fusion data dimension is 128×128×48, and in the mapping expansion process, by setting the memory linear expansion rule, the three-dimensional tensor structure is flattened into a one-dimensional continuous vector stream in row priority order, and a fixed-length segment encoding strategy is used to set the length of each expansion unit to 6144 floating-point data units. Each data unit occupies 4 bytes of memory space, forming a total of 1024 groups of memory mapping fragments. In the expansion process, a direct mapping structure is used to delimit the mapping address range to a 32-bit address space. The address mapping strategy uses an intra-page offset in conjunction with a high-order address switch to ensure that the expansion form of the high-dimensional feature dimension in the memory structure maintains the original spatial continuity. At the same time, a partitioned circular cache mechanism (cache loop Partitioning) caches the feature maps of different channels in four-way parallel cache areas respectively, thereby forming memory mapping structure expansion data, and sends the memory mapping structure expansion data to the pattern recognition and classification unit. The spatial clustering method is used to identify the repeated structural patterns in the data. First, the local self-similarity measure is used to calculate the vector cosine similarity of the adjacent areas for each group of expanded data. The similarity threshold is set to 0.85, and the data segments with cosine angles less than 30 degrees are identified as the initial similar pattern set. Then, the density-based spatial clustering algorithm (DBSCAN) is used to cluster the initial pattern set. The minimum number of samples is set to 6 and the maximum distance threshold is set to 1.2. After the clustering is completed, spatial feature pattern sets with different similarity levels are obtained, which are divided into three categories: A, B, and C. Among them, the category A pattern is full structural matching, the category B pattern is channel vector morphology matching, and the category C pattern is high-frequency area texture. For similarity matching, the three types of patterns are output as differently numbered classification spatial feature patterns, and their starting offsets and length ranges in the memory address are marked. The classification spatial feature patterns are then imported into the weight construction module, and a weight matrix is constructed using the vector weighted integration method. Each type of spatial feature pattern is set as a core node in the weight matrix. Edge connection weights are constructed for the pattern segments around each node through adjacency weight mapping. The edge weights are derived from the aforementioned similarity scores and calculated in combination with the channel synergy coefficient. After the weight matrix is constructed, a gradient optimization operation is performed on the matrix. The edge weights of each node are adjusted using a three-round iterative method. The first round uses linear scaling adjustment, the second round performs boundary fitting adjustment, and the third round performs directional consistency adjustment. The optimized matrix is then subjected to structured sparsity projection processing. In this process, the sparsity is set to 65%, the highest weight connection edges are retained, and those below the set threshold of 0 are deleted.For the edge connection relationship of 2, output structured sparse mapping data. This sparse structure is expressed in the form of an index matrix, recording the effective connection paths and activation weight values of each feature pattern. Pass the structured sparse mapping data into the memory addressing optimization unit, and re-adjust the physical addressing strategy according to the sparse connection relationship. First, reconstruct the access addresses of consecutive feature patterns into a shared addressing area through a compressed address mapping table, and use a displacement address mapping mechanism to bind the same cache mapping segment to multiple activation nodes to avoid repeated access. During the adjustment process, use 1024 bytes as the minimum address remapping unit, adopt a three-level page table structure to record the compressed mapping relationship, and add a reverse addressing index table at the end of the structure for quickly locating the original mapping offset. After completing the addressing optimization, send the mapping result into the activation mapping control module, and use the mapping control engine to perform activation processing on the data segments within each valid address area for the storage and computing units. Set that there are 128 parallel processing units for the activation processing, and each processing unit corresponds to activating a sparse structure node. Uniformly distribute the processing tasks through the control engine and start the matching activation tasks in the physical array. Each activation process uses 4-channel data input for parallel expansion, and finally outputs multiple activation mapping data for the storage and computing units. The dimension of the activation data is 128×48, and each group of data contains complete activation paths and corresponding processing unit number information.
[0056] Preferably, mining the chained interaction correlation data between the activation mapping data of the storage and computing units in step S3 includes:
[0057] Monitor the interaction channels of the activation mapping data of the storage and computing units to obtain interaction channel probe data;
[0058] Extract the activation coupling degree from the interaction channel probe data to generate activation coupling degree correlation data;
[0059] Perform chained fusion mapping according to the activation coupling degree correlation data to obtain chained interaction correlation data.
[0060] Especially importantly, extracting the activation coupling degree from the interaction channel probe data includes:
[0061] Screen the excitation frequency between channels of the interaction channel probe data to obtain channel excitation frequency data;
[0062] Perform synchronous activation interval registration based on the channel excitation frequency data to generate registered activation window data;
[0063] Perform coupling strength quantization mapping on the registered activation window data to generate activation coupling degree mapping data;
[0064] Perform threshold partition regularization according to the activation coupling degree mapping data to obtain activation coupling degree correlation data.
[0065] In this embodiment, the operation process of monitoring the interaction channels of the activation mapping data of the memory-computation unit is as follows: load the activation mapping data of the memory-computation unit generated in the previous stage into the channel probe control unit. The dimension of this activation data is 128×48, and each dimension corresponds to an independent computing unit status signal. In the channel monitoring operation, set every 4 consecutive activation units as a group of interaction unit clusters. Each group of clusters is divided into 32 groups and accessed in parallel to the interaction channel listening link. Use a capacitive coupling probe to record the signal switching frequency and data transfer sequence between channels in real time. The listening period is set to 64 clock units, and the length of each unit is 12.5 nanoseconds. Sample the channel switching signals of all interaction clusters within each period, establish a binary state encoding for each change in the signal rising and falling edges, and finally output an interaction state matrix. This matrix takes the time axis as the horizontal axis and the interaction cluster number as the vertical axis, recording the activation state encoding of each channel in each cluster at each time node. The obtained interaction state matrix is the interaction channel probe data. The operation of screening the excitation frequency between channels for the interaction channel probe data includes: first, input the interaction state matrix into the frequency analysis unit, count the activation times of each channel in each group of interaction clusters within 512 consecutive periods, use a sliding window method with a 64-period unit sliding window, and a sliding step of 16 periods each time. Calculate the excitation frequency value of each channel in each window. The excitation frequency value is expressed as the number of activations per unit time divided by the total length of the window, and the result is floating-point percentage data. Mark the channels with a frequency value higher than the set threshold of 0.65 as strongly activated channels, and the rest as weakly activated channels. At the same time, based on the co-occurrence times of the activation of channel pairs, establish a channel excitation co-occurrence frequency table. The operation of synchronizing the activation interval registration based on the channel excitation frequency data includes: for all strongly activated channels that meet the condition that the co-occurrence frequency is greater than 0.The channels of 45 are extracted, and registration operations are performed based on their activation timestamp vectors. The activation window width is set to 16 cycles, and the activation time vectors of each pair of channels are synchronously detected by the phase alignment method. The phase alignment method adopts the Euclidean distance matching plus time offset compensation strategy, with a maximum allowable offset of ±3 cycles. Finally, the synchronous activation registration windows of each pair of channels in each time period are output. The registration window data is stored in the form of triples, including the channel pair number, start time, and end time. Multiple synchronous registration activation window sequences are constructed in the entire data sequence. The operation of performing coupled strength quantization mapping on the registration activation window data includes inputting all registration window data into the coupled strength calculation module. First, the active time overlap degree is calculated for each registration window. This overlap degree is defined as the ratio of the effective time length of the two channels within the synchronous activation interval to the total registration window length. Then, the channel weight coefficient adjustment value is introduced, which is weighted and corrected based on the previous activation frequency value. Finally, the coupled strength value is calculated by multiplying the overlap degree by the weight coefficient. The coupled strength values of all registration windows are organized into a coupled strength mapping matrix. This matrix uses the channel pair as the row index and the registration window number as the column index, and the elements are coupled strength values. The data format is fixed-point floating-point representation with a precision of 32 bits. The operation of performing threshold partitioning and normalization based on the coupled strength mapping data includes setting the coupled strength threshold level partitioning rules. Specifically, above 0.8 is the high coupling area, 0.5 to 0.8 is the medium coupling area, 0.3 to 0.5 is the low coupling area, and below 0.3 is the weak coupling area. Each element in the coupled strength mapping matrix is labeled with a level to form a level label matrix, and the channel pairs with the same level are clustered into a group to construct an activation coupling degree association set. This set is represented by a structured vector, and each structure unit contains the channel pair number, activation time period index, coupling level label, and synchronous window number. After completing the construction of the activation coupling degree association data, the process of performing chained fusion mapping operations is as follows: The topological path construction method is adopted to search for channel sequences with continuous coupling paths in the coupling degree association set. For each path, the transfer coupling direction and strength gradient direction between nodes are calculated, and a chained interactive channel graph is constructed based on the path continuity. The channel graph is encoded in the form of a sparse graph and embedded into the activation structure of the current memory and computing unit as a parallel computing reference structure. Finally, the chained interactive association data is output. The chained data format is represented by a structured multi-directional graph model. Each node contains the activation channel number, coupling level, channel position offset, and channel coupling direction identifier. Each edge records the coupling path code number and edge coupling strength parameter.
[0066] Preferably, the recursive self-calibration of the original stored data according to the chained interactive association data in step S3 includes:
[0067] Inferring interactive pulse data according to the chained interactive association data;
[0068] Perform temporal reconstruction based on interactive pulse data to generate temporally interactive reconstructed data;
[0069] Verify the associated context through the temporally interactive reconstructed data to obtain verified interactive associated data;
[0070] Perform feature residual decoupling on the verified interactive associated data to obtain residual decoupled data;
[0071] Perform autoregressive feature recursion on the residual decoupled data to generate recursive feature data;
[0072] Determine the attenuation state of the verified interactive associated data based on the recursive feature data;
[0073] Perform synchronous gain adjustment on the verified interactive associated data based on the attenuation state to generate self-calibrated update data.
[0074] In this embodiment, the operation of inferring interactive pulse data based on chain interaction associated data includes loading the chain interaction associated data into an interactive structure decoding module for inferring the logic flow. This module extracts the activation states between channels segment by segment in the time domain based on the interactive path order identified in the chain diagram, and extracts the state change mapping between the starting activation node and the ending response node of the channels by identifying the continuous activation conduction structure between the channels. For each path, neural structure analogy deconstruction is performed in the order of channel activation, and by identifying the mutation segments of the activation states, it is determined whether there are inferable interactive pulses. Then, according to the cumulative distribution trend of the activation state differences between adjacent channels in the path, a channel pulse indication label is output. The paths with continuous activation responses and meeting the mutation trigger characteristics are marked as strong response paths, and based on this, an interactive pulse data structure is constructed. Each structural unit contains a path number, a channel sequence, an inferred pulse type identifier, a time index label, and a description of the activation sequence summary features. The operation of performing time series reconstruction based on the interactive pulse data includes inputting the interactive pulse data into an index-driven synchronous rearrangement module, and performing start-stop synchronous scanning processing on the activation segments of all included pulse activation paths. First, the channels in each path are numbered in chronological order, and the original activation signal segments corresponding to each number are extracted. The original activation sequence is sliced uniformly using a time window with a fixed width. The length of the time window is set using a preset parameter and according to the event period corresponding to the inferred pulse. Each activation segment is divided into multiple small segments and is attached with relative time tags. After aligning all the segment data at a unified reference time, a time series activation matrix structure driven by the channel sequence is generated. Subsequently, the cross nodes between adjacent paths are identified as interactive points, and time marking signals are inserted into the activation matrix to form a multi-path synchronous interaction structure. The final output result is time series interaction reconstruction data organized by path, and the structure contains a channel index, a time period number, a summary of the signal segment features, and an interaction mark index. The operation of verifying the associated context through the time series interaction reconstruction data includes inputting the reconstructed data into a verification module for path consistency verification. The module performs context backtracking processing based on the activation time order of the channels in each path, identifies the activation delay relationship between consecutive channels in the path, and determines whether it conforms to the conduction characteristic logic defined in the previous chain interaction associated data. By comparing the start and end times of the activation time periods and the expected interaction intervals in each path, if there is a significant deviation in the delay time between the channels within a path from the range defined in the preset conduction logic model, then that path is marked as an inconsistent path, otherwise it is marked as a consistent path. At the same time, a path consistency matrix is constructed for the verification results of all paths. This matrix contains a path number, a verification status label, a deviation interval time period, a deviation node number, and an activation delay index feature. The final output structured verification data is the verified interaction associated data. The operation of performing feature residual decoupling on the verified interaction associated data includes,Map the verification status of each path to its actual activation time distribution in the reconstructed data. Use the pre-stored standard activation template as a control model to compare the timing differences between the actual activation structure of the path and the standard template. Extract the activation start point, end point, duration, and activation sequence number of each channel to form a feature vector. Calculate the residual information through the differences between vectors, and organize it into a residual vector group in units of paths. Deconstruct the vector group through dimensional decomposition, convert the cross-coupling residuals between all channels into several independent channel response residual sets, and perform structured storage on the response residual sets respectively. Each residual set corresponding to a channel contains activation duration offset, activation sorting offset, delay interval offset, and synchronization mismatch indicators. After structural integration, the output is residual decoupled data. The operation of autoregressive feature recursion on the residual decoupled data includes dividing the residual data of each channel into continuous time series according to its time period label, performing a residual trend scanning operation on each segment of data according to a fixed-length time sliding window, identifying whether there is a phenomenon of continuous offset accumulation in each segment, and judging the trend direction by comparing whether there is a gradually increasing trend in the offset degree distribution in the front and back windows. Then construct a residual recursion trajectory according to the trend direction. The trajectory contains the time period number, offset direction, offset acceleration label, historical cumulative residual amplitude classification identifier, and activation frequency change summary. Integrate the trajectories of all channels to generate recursion feature data. Each record unit in this data contains channel identification, recursion direction, time index, and recursion trend classification code. The operation of determining the attenuation state of the verification interaction correlation data according to the recursion feature data includes inputting the recursion feature data into the channel state evaluation module. First, classify the channels involved in each path according to their recursion direction identifiers. Mark the channels with continuously changing negative recursion directions as the decreasing category, mark the channels with continuously changing positive recursion directions as the enhancing category, and mark the channels with multiple consecutive time periods marked as stable as the stable category. Summarize the marking information of all channels to generate a path-level status index table. This index table records the path number, channel number, status label, corresponding time period label, residual level category, and the conclusion of the synchronization status comparison between channels. At the same time, identify the channel combinations that are frequently marked as the decreasing category status in the coupling relationship and mark them as potential attenuation coupling paths. Finally, all the mapping results of the status and path relationships form an attenuation state data set. The operation of synchronously adjusting the gain of the verification interaction correlation data based on the attenuation state includes inputting the attenuation state data set into the gain control engine module, performing enhancement instruction allocation processing on all the path channels in the decreasing state, assigning a predefined gain enhancement factor to each decreasing category channel in the path, and applying this factor to the activation amplitude field of the corresponding channel in the path structure. At the same time, extract the channel pairs marked as attenuation coupling in all the coupling paths, and use the dual-channel synchronization strategy to set the enhancement ratio of their interaction amplitudes, and perform synchronous proportional adjustment on the amplitude field while ensuring that the synchronous activation intervals are consistent.Complete parameter replacement for the general reconstruction module of all enhanced channels, and repackage the adjusted structure into self-calibrated updated data. In each update node structure, record the activation amplitudes before and after the update, adjustment parameters, channel numbers, path numbers, and time period labels. Reorganize all self-calibrated updated data in a chained structure for storage.,
[0075] Preferably, predicting heterogeneous access channels based on the self-calibrated updated data in step S4, and reconstructing each hierarchical data access link through the predicted heterogeneous access channels includes:
[0076] Extract the multi-domain channel attribute spectrum from the self-calibrated updated data;
[0077] Construct a channel cluster grouping structure based on the multi-domain channel attribute spectrum;
[0078] Sort out the corresponding relationship between channel clusters and data blocks according to the storage-computation hierarchy based on the channel cluster grouping structure;
[0079] Decompose the corresponding relationship between channel clusters and data blocks into multiple dynamically reconfigurable links;
[0080] Predict heterogeneous access channels based on the dynamically reconfigurable links;
[0081] Perform adaptive reconstruction on the predicted heterogeneous access channels based on the preset storage capacity to generate node adaptation links;
[0082] Simulate the cache limit extension boundary through the node adaptation links, and determine the preferred cache layout based on the simulated cache limit extension boundary;
[0083] Perform hierarchical link reconstruction and integration according to the preferred cache layout to obtain the reconstructed hierarchical data access links.
[0084] In this embodiment, the multi-domain channel attribute spectrum extracted from the calibration update data integrates access data from different sources using multi-source data fusion technology, including key attributes such as device identifier, storage type, current bandwidth, historical latency, error rate, etc. The multi-dimensional attribute data is normalized using the Z-score normalization method to make the mean of each attribute 0 and the standard deviation 1. A multi-domain channel attribute matrix is constructed, and the matrix dimension is the number of devices multiplied by the number of attributes. The first K principal components are extracted through principal component analysis, and the value of K is determined according to the principle that the cumulative contribution rate exceeds 85%. The K-means clustering algorithm is used to divide the attribute spectrum into three categories: basic storage cluster, cache cluster, and computing acceleration cluster according to device type and storage medium. The weighted average attribute value of each cluster is calculated, and the weight is determined according to the device online duration. Finally, the result is saved in the binary attribute spectrum file format, including a file header description and a data segment. The file header defines the attribute type and unit, and the data segment organizes the attribute values by cluster. Based on the multi-domain channel attribute spectrum, a channel cluster grouping structure is constructed, a hierarchical clustering tree is established for each type of channel in the attribute spectrum file, and the Ward minimum variance method is used to calculate the inter-cluster distance. The merging stops when the increase in intra-cluster variance exceeds the threshold At this time, The value is 1.5 times the initial intra-cluster variance. The resulting clusters are divided into three categories according to the storage-computing architecture: high-performance computing clusters, large-capacity storage clusters, and balanced clusters. A globally unique identifier GUID is assigned to each cluster, which is generated by a combination of timestamp, machine ID, and serial number. An adjacency matrix is established for the channels in each cluster. The matrix elements are calculated using the Jaccard similarity coefficient, and the threshold is set to 0.6. The threshold is determined from the interval [0.5, 0.7] using the elbow rule to form a weighted directed graph structure. This structure is stored in an adjacency table, where entries contain the pointing node, weight value, and connection type tag. According to the channel cluster grouping structure, the channel cluster-data block correspondence is sorted out according to the storage-computing hierarchy, and a storage-computing hierarchy mapping table is constructed. The table entries include data block identifiers, hot and cold attribute tags, access frequency counters and priority levels. The access frequency counters are updated using the LRU (Least Recently Used) strategy, and the time window is set to 5 minutes. The hot and cold tags are divided according to the median access frequency. The tags above half of the value are marked as hot, otherwise they are cold. For each data block, the three channel clusters with the highest correlation are calculated. The correlation is calculated using TF-IDF weighted calculation. The word frequency is counted by access records, and the inverse document frequency is the ratio of the total number of data blocks in the cluster to the total number of data blocks in the system. The data block-channel cluster association matrix is generated and stored in CSR (Compressed Sparse Row) format to save space. An association log is created to record change operations in timestamp order. Each log contains the operation type, the data block involved, the associated cluster identifier and timestamp before and after the operation. The channel cluster-data block correspondence is decomposed into multiple dynamically reconfigurable links. A graph traversal algorithm is implemented based on the association matrix. The improved Dijkstra algorithm is used to calculate the optimal data flow path. The edge weight is determined by the bandwidth-delay product. The delay is the moving average of the last 10 measurements, and the bandwidth is a real-time detection value. The path calculation result is stored as a link descriptor object, which contains the path sequence, estimated delay, and bandwidth guarantee level. The health of the channels in the path is evaluated. Indicators include error rate, load rate, and temperature. An upper health threshold is set. Channels exceeding the upper limit trigger a rerouting mechanism. A backup path pool mechanism is used to calculate and maintain three backup paths for each link in advance. The healthiest path is selected during rerouting. Link descriptors are stored according to the storage layer to which they belong. The computing layer link uses the metadata area of the high-performance storage medium, and the storage layer link uses the log area of the SSD. Heterogeneous access channels are predicted based on dynamically reconfigurable links. The time series prediction method is used to analyze the historical traffic data of the link. The Holt-Winters three-parameter exponential smoothing model is selected, and the smoothing coefficient Set them to 0.3, 0.2, and 0.1 respectively. Fit the parameters from historical data using the least squares method to predict the traffic distribution within the next time window. The size of the time window is set to 15 minutes. Adjust the link bandwidth allocation ratio according to the prediction results. Adopt a dynamic weight adjustment algorithm to prioritize ensuring the link resources with predicted traffic growth exceeding 20%. The bandwidth adjustment step size is set to 10 Mbps, and the cooling period after each adjustment is 30 seconds to prevent frequent fluctuations. Synchronously update the bandwidth parameters in the link descriptor. Recover the resources of the links with predicted traffic less than 50% of the current allocation, and transfer the idle bandwidth to the high-load links, with the transfer amount not exceeding 75% of the idle bandwidth. Adaptively reconstruct the predicted heterogeneous access channels based on the preset storage capacity to generate node-adapted links. Read the node hardware configuration parameters, including the number of CPU cores, memory capacity, storage capacity, and network bandwidth, and establish a resource capacity matrix. The matrix elements are the proportions of each resource to the total resources. Calculate the resource gap based on the resource matrix and the predicted traffic. Set the resource over-allocation coefficient to 1.2 to generate a set of reconstruction strategies. The strategies use a greedy algorithm to prioritize meeting the requirements of critical links, and the critical links are defined by grading according to the business importance. During the reconstruction process, maintain the bandwidth ratio between the storage layer and the computing layer between 3:1 and 5:1. Generate a node-adapted link configuration file, organized in JSON format, containing the link ID, source node, target node, allocated bandwidth, QoS level, and priority label. The configuration file is stored in the shared memory area for each module to access. Simulate the cache limit extension boundary through the node-adapted link, and determine the optimal cache layout based on the simulated cache limit extension boundary. Establish a cache simulation environment, use a pseudo-random number generator to generate access patterns, and the access frequency distribution conforms to the Zipf law, with the parameter α set to 0.85, representing the intensity of the long-tail effect. Simulate the cache capacity increasing gradually from 100 MB to 10 GB, with an increase of 1 GB at each level, and record the change curve of the cache hit rate at different capacities. Use the binary search method to locate the capacity corresponding to the 95% quantile of the hit rate as the cache limit boundary. When optimizing the cache layout, classify the data blocks by temperature. The hot data is managed using the LRU strategy, and the cold data uses the LFU (Least Frequently Used) strategy. The cache block size is set to 4 KB, aligned to the memory page boundary. The distribution strategy uses the consistent hashing algorithm, and the number of virtual nodes is set to 200 times the number of physical nodes to reduce the data migration amount when the nodes change.Hierarchical link reconstruction and integration are performed according to the preferred cache layout to obtain reconstructed hierarchical data access links. The cache levels are divided into three levels: L1, L2, and L3. The L1 cache is adjacent to the computing node, with a capacity of 10% of the total cache. L2 is the intermediate layer with a capacity of 50%, and L3 is the remote layer with a capacity of 40%. Data prefetch thresholds between cache levels are set based on the preferred layout. The threshold from L1 to L2 is set at 80% utilization, and from L2 to L3 is set at 70%. When reconstructing the hierarchical links, priority control is implemented for cross-layer transmissions. The lower limit of the exclusive bandwidth ratio for high-priority links is 25%, and the upper limit is 50%. End-to-end delay measurement is performed when establishing the link, using the three-way handshake protocol and RTT (Round-Trip Time) estimation. If the delay exceeds the threshold, path reconstruction is triggered, and the threshold is set at 1.5 times the average delay. The information of the reconstructed link is stored in a distributed key-value database, where the key is the global identifier of the link, and the value is the serialized object of the link parameter configuration.
[0085] Preferably, the multi-layer collaborative flow control cache integration according to the reconstructed hierarchical data access links in step S4 includes:
[0086] Perform node priority reconstruction on the reconstructed hierarchical data access links to generate reconstructed data access links;
[0087] Perform multi-source scheduling coupling processing on the reconstructed data access links to generate collaborative scheduling links, where the range of the channel bandwidth balance coefficient for multi-source scheduling coupling processing is set to 0.4 - 0.8;
[0088] Perform cache path cascade integration on the collaborative scheduling link data to obtain flow control cache integration data. During the cache path cascade integration process, the cache block granularity is set to 512KB to 2MB, and the hierarchical cascade does not exceed 4 layers of structure;
[0089] Suppress node queuing jitter based on the flow control cache integration data to generate a stable output cache, where the jitter suppression operation is performed based on the rule that the maximum transient queuing delay does not exceed 5ms;
[0090] Perform delay trajectory difference cancellation processing through the stable output cache data to obtain access delay elimination and optimization data.
[0091] In this embodiment, historical bandwidth utilization data of each node is collected. The bandwidth usage in the past hour is obtained through the monitoring system, and the bandwidth peak value and average utilization rate of each node are recorded. At the same time, the current load status of the node is monitored, including CPU usage rate, memory occupancy rate, and cache hit rate. These data are collected in real time by the distributed monitoring probe and updated every 10 seconds. According to these metrics, the comprehensive priority score of each node is calculated. Using the method of weighted summation, the bandwidth utilization rate accounts for 40% of the weight, the current load status accounts for 30% of the weight, and the cache hit rate accounts for 30% of the weight. The calculated priority scores are sorted. The top 20% of the nodes are marked as high-priority, the middle 60% are marked as medium-priority, and the bottom 20% are marked as low-priority. The nodes in the reconstructed data access link are marked with priorities. High-priority nodes are preferentially allocated high-speed channel resources, medium-priority nodes are allocated medium-speed channel resources, and low-priority nodes are allocated low-speed channel resources, generating reconstructed data access link information with priority markings. A channel bandwidth monitoring mechanism is established. The real-time bandwidth usage data of each channel, including the used bandwidth and the remaining bandwidth, is collected every 50 milliseconds. The bandwidth difference rate between channels is calculated. When the difference rate exceeds 15%, a multi-source scheduling coupling operation is triggered. The initial balance coefficient is 0.6, and the optimal value is found between 0.4 and 0.8 through a dynamic adjustment mechanism. During the adjustment process, the change in the overall throughput of the monitoring system is monitored. After each adjustment, wait for 10 milliseconds to observe the effect, and record the balance coefficient with the highest throughput as the current optimal value. According to the determined balance coefficient, the bandwidth resources of each channel are allocated. High-priority links are allocated a larger proportion of the bandwidth, medium-priority links are allocated a medium proportion, and low-priority links are allocated a smaller proportion, but ensure that the minimum service bandwidth of each link is not less than 100 Mbps, generating collaborative scheduling link metadata containing the balance coefficient and the corresponding bandwidth allocation scheme. The collaborative scheduling link data is implemented with cache path cascading integration to obtain flow control cache integration data. During the cache path cascading integration process, the cache block granularity is set to 512 KB to 2 MB, and the hierarchical cascade does not exceed 4-layer structure, constructing a multi-level cache path structure. The first layer is the storage layer cache, the second layer is the computing layer distributed cache, and the third layer is the node local cache. According to the heat of data access, data is allocated to different levels of caches. A heat grading mechanism is adopted. Data with an access frequency higher than the set threshold is marked as hot data and allocated to the first layer cache. Data with a moderate access frequency is marked as medium-temperature data and allocated to the second layer cache. Data with a lower access frequency is marked as cold data and allocated to the third layer cache. The maximum capacity of each layer of cache is set. The capacity of the first layer cache accounts for 50% of the total capacity, the second layer accounts for 30%, and the third layer accounts for 20%. The LRU replacement strategy is adopted to manage cache blocks within each layer of cache. When the number of cache blocks exceeds the limit, the least recently used data block is eliminated. Different differential compression strategies are implemented for cache data at different levels. Hot data is not compressed to ensure access speed, and medium-temperature data uses a fast compression algorithm.Cold data uses a high compression ratio algorithm to generate a flow control cache integrated data object containing a cache path structure and parameters. Based on the flow control cache integrated data, node queuing jitter is suppressed to generate a stable output cache. The jitter suppression operation is performed based on the rule that the maximum transient queuing delay does not exceed 5 ms. A node queuing status monitoring mechanism is established, and the queue length and average service time data are collected every 10 milliseconds to calculate the transient queuing delay. When the detected transient queuing delay exceeds 5 ms, the jitter suppression mechanism is started. A dynamic window adjustment strategy is adopted, and the optimal service window size is calculated according to the current queue length and service time. The window size adjustment range is limited within ±15% of the current value. Group scheduling is implemented for queues exceeding the threshold, and they are divided into several groups according to the packet arrival time. The maximum service volume of each group does not exceed 20% of the total queue length. The data of each group is processed in turn using the polling method to ensure relatively balanced service time for each packet. A stable output cache index table is generated to record the physical location and service timestamp of each data block in the buffer. Through the stable output cache data, delay trajectory difference cancellation processing is performed to obtain access delay elimination optimized data. A delay trajectory prediction model is established, and the data block arrival time series of the previous 10 time windows is collected as the model input to predict the data block arrival time at the next moment. The predicted arrival time is compared with the actual arrival time, and the time difference is calculated. When the difference exceeds the set threshold, a delay compensation operation is performed, and different compensation strategies are selected according to the size of the difference. When the difference is small, the prefetch strategy is adopted to read the data 1 millisecond in advance. When the difference is moderate, the double buffer mechanism is used to load the data in parallel. When the difference is large, the fast channel switching mechanism is triggered to re-obtain the data from the standby high-speed channel. An optimized access delay log is generated to record the original delay, compensated delay, and final access timestamp of each data block.
[0092] Preferably, the task reorganization of the access delay elimination optimized data in step S5 includes:
[0093] Trace the data flow path of the access delay elimination optimized data, and analyze the dependency path characteristics in the access delay elimination optimized data based on the data flow path;
[0094] Identify bottleneck nodes according to the dependency path characteristics;
[0095] Determine the critical dependency path based on the bottleneck nodes;
[0096] Decouple and reorganize the task units of the critical dependency path to obtain decoupled and reorganized task data;
[0097] Perform dependency grouping aggregation on the decoupled and reorganized task data to generate dependency grouping aggregation data;
[0098] Perform parallel mapping distribution based on the dependency grouping aggregation data to obtain task reorganization parallel processing data.
[0099] In this embodiment, the data flow path of the access latency elimination optimized data is traced, and the dependency path features in the access latency elimination optimized data are analyzed based on the data flow path. The timestamps and node identifiers of the access latency elimination optimized data when flowing in the system are collected. The sequence of nodes passed by each data packet and the time consumption are recorded through a distributed tracing system. The timestamp information is embedded in the data packet header with a precision reaching the nanosecond level. The unique identifier of the data packet is calculated using a hash function, and the hash algorithm adopts SHA-256 to ensure the uniqueness of the identifier. A data flow path graph is constructed, where nodes represent processing units, edges represent the data flow direction, and the edge weight is the transmission delay. 5000 data flow samples are collected, and features such as path length, node hops, and transmission delay are analyzed. The key path dependency patterns are extracted, and the importance of nodes is calculated using an improved PageRank algorithm. The number of iterations is set to 20 times, and the damping factor is set to 0.85. Long dependency paths and cyclic dependencies are identified. The data flow path graph is stored in an adjacency list structure, and each table entry records the target node, transmission delay, and data volume size. Bottleneck nodes are identified according to the dependency path features. By analyzing the node delay and throughput metrics in the data flow path graph, a bottleneck node identification criterion is set, and the average node delay exceeds the system average delay by 1.Nodes with a throughput less than one-fifth of the average are determined as potential bottleneck nodes. Collect the performance metrics of each node, including average latency, maximum latency, throughput, and packet loss rate. Calculate the weighted score of each node, with the weight distribution being 60% for latency and 40% for throughput. The score threshold is dynamically adjusted based on historical data. Use a sliding window algorithm to calculate the statistical metrics of the last 1000 data packets, with a window sliding step of 50. Conduct a secondary verification on nodes exceeding the threshold to confirm the bottleneck status by increasing probing data packets. Mark the confirmed bottleneck nodes and record their IDs, types, and processing capabilities. Store the results in a distributed key-value database, with the key being the node ID and the value being a structure containing performance metrics and processing suggestions. Determine the critical dependency paths based on the bottleneck nodes. Extract all paths containing the identified bottleneck nodes from the data flow path graph. Traverse the path graph using a depth-first search algorithm with a depth limit of 10 levels. Collect paths with a length exceeding 5 hops as candidate critical paths. Calculate the comprehensive weight of each candidate path, which is determined by the total path latency, the number of bottleneck nodes, and the data volume size. Latency accounts for 50%, the number of nodes accounts for 30%, and the data volume accounts for 20%. Sort the candidate paths and select the top 20 with the highest weights as the critical dependency paths. Mark the determined critical paths and record the path IDs, start and end nodes, and the included bottleneck nodes. Generate a critical path metadata descriptor, which contains the path topology structure, node attributes, and weight calculation parameters. Store the descriptor in an in-memory database with an expiration time of 24 hours and recalculate after timeout. Decouple and reorganize the task units on the critical dependency paths to obtain decoupled and reorganized task data. Divide the tasks on the critical dependency paths into independent units according to the processing logic. Use a control flow graph analysis method to identify the boundaries of task units, with the division granularity set to a single function call or data processing stage. Establish a task unit dependency relationship matrix, where the matrix elements are the data dependency strengths between task units, and the strengths are divided into three levels: strong, medium, and weak. Reorganize the task units according to the dependency relationship matrix. Task units with strong dependency relationships are retained in the same processing unit, and task units with medium and weak dependency relationships are split across units. Maintain data consistency during the reorganization process and use a distributed lock mechanism to ensure the atomicity of critical data operations. Store the reorganized task units in a task queue in the execution order. The queue uses a priority structure, with the priority of urgent tasks higher than that of ordinary tasks. The priority is set according to the task latency sensitivity, and the sensitivity metrics are dynamically evaluated by the system. Generate a decoupled and reorganized task data object, which contains task unit IDs, dependency relationships, and execution order. Group and aggregate the decoupled and reorganized task data according to dependencies to generate dependency-grouped and aggregated data. Group the task units according to the dependency relationship strength between them. Task units with strong dependency relationships are divided into the same group, and task units with medium and weak dependency relationships are assigned to different groups. Use a hierarchical clustering algorithm for grouping, with the distance metric being the reciprocal of the dependency relationship strength and the clustering radius set to 0.5. Control the number of task units in each group between 5 and 20, calculate the overall dependency characteristics of each group of task units, including average dependency intensity, maximum dependency distance, and data transfer volume, establish an inter-group communication matrix, record the data interaction volume and latency between different groups, and implement the merger and split optimization of task units according to group characteristics and the communication matrix. Merge adjacent groups with large communication volume and split groups with high communication density and independent processing logic to generate a dependency grouping aggregation data object. The object records the group ID, the included task units, the dependency relationship graph, and communication characteristics. Based on the dependency grouping aggregation data, perform parallel mapping and distribution to obtain task reorganization parallel processing data. Assign an independent processing resource pool to each dependency group. The resource pool includes CPU cores, GPU accelerators, and memory blocks. The resource allocation ratio is determined according to the grouping task characteristics. Allocate more CPU and GPU resources for compute-intensive tasks and a larger memory capacity for data-intensive tasks. Use a polling scheduling algorithm to assign grouping tasks to available processing nodes. The node selection priority is determined by resource utilization, network bandwidth, and processing power. The weight distribution is 30% for resource utilization, 40% for bandwidth, and 30% for processing power. Set the task parallelism parameter. The number of parallel tasks does not exceed twice the number of CPU cores and is not less than 2. Perform thread-level parallel optimization on tasks assigned to the same node. The number of threads matches the number of tasks, and the work-stealing algorithm is used between threads to balance the load. Package the task execution plan after parallel mapping into a task reorganization parallel processing data object. The object includes a task allocation table, parallel parameters, and a scheduling strategy. The data is stored in the shared memory area.
[0100] Especially important is to identify bottleneck nodes according to the dependency path characteristics, including:
[0101] Archive the path delay gradient of the dependency path characteristic data to generate a path delay distribution;
[0102] Extract the path delay distribution of the path delay distribution to obtain delay peak indexing data;
[0103] Perform path overlap metric detection on the delay peak indexing data to obtain path conflict density data;
[0104] Perform node resource load mapping based on the path conflict density data to generate node load occupancy data;
[0105] Perform the maximum residence-waiting ratio screening according to the node load occupancy data to obtain bottleneck node data.
[0106] In this embodiment, the timestamps and node identification information in the dependency path feature data are collected to construct a dataset containing at least 10,000 records. Each record includes a path ID, a starting node, an ending node, and a transmission delay. The transmission delays of each path are sorted in chronological order, and the delay change amount between adjacent time points is calculated to form a delay gradient sequence. The sliding window algorithm is used to process the delay gradient sequence. The window size is set to 50 time points, and the step size is 10. The average value and standard deviation of the delay gradient within the window are calculated, and the calculation results are archived and stored according to the path ID to form a path delay gradient archive. The archive data structure adopts the key-value pair form, where the key is the path ID, and the value is a tuple containing the average gradient, standard deviation, and gradient sequence. The archived path delay gradient data is grouped and statistically analyzed according to the path ID, and the delay distribution histogram of each path is calculated. The number of bins in the histogram is set to 50, and the range covers the minimum to maximum delay values of all paths to generate a path delay distribution dataset. The path delay distribution is extracted to obtain the delay peak indexing data. For each path in the path delay distribution dataset, the bin interval with the largest count in the delay distribution histogram is located, and the midpoint value of this interval is marked as the delay peak of the path. The bimodal detection algorithm is used to identify the delay distribution with multimodal characteristics. When multiple peaks are detected and the height ratio of the sub-peak to the main peak exceeds 0.At 3 o'clock, mark this path as a multi-peak path. For a single-peak path, directly record the position of the main peak. For a multi-peak path, record all peak positions and their relative heights, and construct a delayed peak index table. The table structure includes the path ID, the main peak delay value, the secondary peak delay value (if any), and the kurtosis coefficient. The kurtosis coefficient is calculated using the fourth-order central moment formula, with the denominator being the fourth power of the standard deviation and the numerator being the sum of the fourth powers of the deviations. Conduct a path overlap metric detection on the delayed peak indexing data to obtain the path conflict density data, and construct a path overlap detection matrix. The rows and columns of the matrix represent different paths, and the matrix element values represent the degree of overlap of the delayed peaks of the corresponding path pairs. The overlap degree calculation formula is the proportion of the overlap time in the total time when the absolute value of the peak delay difference between two paths is less than the set threshold τ. The threshold τ is initially set to 5 milliseconds and dynamically adjusted according to the network environment within the range of [3, 8] milliseconds. Traverse all path pairs to calculate the overlap degree, forming an N×N-dimensional overlap matrix, where N is the total number of paths. Sum the overlap matrix by row to obtain the total conflict time of each path, and divide it by the total monitoring time to obtain the path conflict density index. Construct a path conflict density vector, with the vector elements being the conflict density values of each path. After sorting from high to low density, store it in the distributed cache system Memcached. Based on the path conflict density data, perform node resource load mapping to generate node load occupancy data, establish a node resource monitoring system, and collect data on the CPU usage rate, memory occupancy rate, and network bandwidth utilization rate of each node. The sampling frequency is set to once per second, and a sliding window is used to calculate the average and peak resource usage in the last 60 seconds. Map the paths in the path conflict density vector to the corresponding start and end nodes, and accumulate the conflict density values of all paths passing through the node to obtain the total conflict load of the node. Combine the node resource monitoring data to calculate the node comprehensive load index. The calculation formula is Comprehensive Load =. × Conflict Load + × CPU Usage Rate + × Memory Occupancy Rate + × Bandwidth Utilization Rate, Weight Coefficient = 0.4, = 0.3, = 0.2, = 0.1, normalize the calculated comprehensive load index and map it to the interval [0, 1] to generate a node load occupancy data table. The table structure includes node ID, comprehensive load index, resource utilization rates, and conflict load contribution rates. The data is stored in the distributed time-series database InfluxDB, and the data retention policy is set to 30 days. Screen the bottleneck node data according to the node load occupancy data. Define the node residence time as the average processing time of the data packet in the node, and the waiting time as the average waiting time of the data packet in the node queue. Collect the residence time and waiting time data of each node through the distributed tracing system, calculate the residence-waiting ratio RWR = residence time / waiting time for each node, set the RWR threshold range to [0.5, 2.0], mark the nodes outside this range as abnormal nodes, sort the nodes within the normal range in descending order of the comprehensive load index, select the node with the highest load index and an RWR value close to the lower limit of the threshold as the candidate bottleneck node, further analyze the resource usage trend of the candidate node, calculate the moving standard deviation of the CPU, memory, and bandwidth utilization rates in the last 10 minutes, and confirm that the resource fluctuation is too large when the standard deviation exceeds 30% of its respective mean. Finally, determine the node with the highest comprehensive load index and stable resource fluctuations as the bottleneck node, and generate a bottleneck node data table containing node ID, comprehensive load index, RWR value, resource fluctuation index, and confirmation timestamp.
[0107] Preferably, the pipelining scheduling of the task reorganization parallel processing data in step S5 and the execution of eliminating parallel conflicts include:
[0108] Divide the task reorganization parallel processing data into pipeline segments to generate preliminary pipeline segment division data;
[0109] Perform resource affinity matching based on the preliminary pipeline segment division data to obtain resource affinity matching data;
[0110] Perform parallel timing optimization on the task reorganization parallel processing data according to the resource affinity matching data to generate timing optimization scheduling data;
[0111] Detect potential conflict clusters in the timing optimization scheduling data, and strip and reconstruct the conflict paths based on the potential conflict clusters to obtain conflict stripping and reconstruction data;
[0112] Fuse the conflict stripping and reconstruction data into a conflict-free path to obtain parallel conflict decoupling data.
[0113] In this embodiment, the task dependency graph in the parsed task reorganization parallel processing data is analyzed, the execution time and resource requirement parameters of each task node are extracted, the pipeline segments are divided by using the method based on critical path analysis. All paths from the start node to the end node in the task dependency graph are traversed, the total execution time of each task node on each path is calculated, the longest path is determined as the critical path, and the task nodes on the critical path are evenly sliced into several pipeline segments according to the execution time. The number of task nodes included in each segment is dynamically adjusted according to the number of system hardware resources. For example, when the system has 8 computing units, the number of task nodes included in each segment is set to 1 / 8 of the total number of tasks. The task nodes on the non-critical path are allocated to each pipeline segment by using the backward push strategy to ensure that the difference in the total execution time of each pipeline segment does not exceed 15%. A preliminary pipeline segment division data structure is generated, including the pipeline segment number, the list of task nodes included, the expected execution time, and the total resource requirement. Based on the preliminary pipeline segment division data, resource affinity matching is performed to obtain resource affinity matching data. The system hardware resource configuration information is collected, including the number of CPU cores, the number of GPU accelerators, the memory capacity, and the cache size. A resource feature matrix is established, where the rows of the matrix represent different types of hardware resources and the columns represent the resource requirement characteristics of each pipeline segment. The cosine similarity algorithm is used to calculate the matching degree between each pipeline segment and the hardware resources. The similarity calculation formula is the dot product of vectors divided by the product of the vector norms. The resource affinity threshold range is set to [0.7, 1.0]. When the similarity is lower than 0.7, it is marked as a low-affinity pipeline segment. For the low-affinity pipeline segments, a resource reallocation strategy is implemented. The computing unit with more idle resources is preferentially selected for migration, and the task dependency relationship is kept unchanged during the migration process. The greedy algorithm is used to select the migration target node, and the node that can maximize the affinity is selected each time. A resource affinity matching data table is generated, recording the pipeline segment number, the matching hardware resource number, the affinity score, and the migration status flag. According to the resource affinity matching data, parallel timing optimization is performed on the task reorganization parallel processing data to generate timing optimization scheduling data. A task execution timeline model is constructed, and the pipeline segments are allocated to the corresponding hardware resources according to the resource affinity matching results. The dynamic programming algorithm is used to optimize the task scheduling order on each resource. The state transition equation is defined as the total completion time of the task sequence on the current resource being equal to the completion time of the previous state plus the execution time of the current task. The constraint conditions are the task dependency relationship and the resource exclusivity requirement. For tasks with data dependency relationships, a pipeline parallel strategy is implemented, and the pipeline depth parameter is set to 3, that is, each task is split into 3 stages and executed in parallel on the pipeline. The earliest start time and the latest end time of each pipeline segment are calculated to generate a scheduling plan in the form of a Gantt chart, which is converted into a timing optimization scheduling data structure, including the time slice number, the allocated pipeline segment number, the start time, and the end time. Potential conflict cluster detection is performed on the timing optimization scheduling data, and conflict path stripping and reconstruction are performed based on the potential conflict clusters to obtain conflict stripping and reconstruction data.A conflict detection matrix is established. The rows and columns of the matrix respectively represent pipeline segments on different time slices. The matrix element value of 1 indicates that there is a resource competition or data dependence conflict for the corresponding pipeline segment on the time slice. The scan line algorithm is used to traverse the time slice sequence in the scheduling data to detect whether there is a conflict between pipeline segments on the same resource on adjacent time slices. The conflict threshold parameter is set that the number of resource contentions exceeds 2 times. When a conflict is detected, the relevant pipeline segments are marked as potential conflict clusters. Path analysis is performed on the potential conflict clusters to identify the resource types and data dependence paths where the conflicts occur. The path stripping strategy is adopted to migrate the conflicting pipeline segments to idle resources or adjust the execution time, and the task dependence relationship remains unchanged during the migration process. For the reconstructed scheduling scheme, the earliest start time and the latest end time of each pipeline segment need to be recalculated, and a conflict stripping reconstruction data table is generated, recording the original pipeline segment number, the newly allocated pipeline segment number, the time slice number after migration, and the conflict resolution mark. The conflict stripping reconstruction data is fused into a conflict-free path to obtain parallel conflict decoupled data, and a conflict-free scheduling graph model is constructed. The nodes represent pipeline segments, and the edges represent task dependence relationships or resource allocation constraints. The depth-first search algorithm is used to traverse the scheduling graph to detect whether there is a loop structure. When a loop is detected, a loop-breaking operation is performed, and the edge with the smallest weight in the loop is selected for deletion. The weight is defined as the product of the task execution time and the resource competition degree. The path fusion operation is performed on the reconstructed scheduling graph to merge the pipeline segments without dependence relationships to be executed on different resources in the same time slice to improve resource utilization rate. The task load balance on each resource is maintained during the merging process, and the difference does not exceed 10%. A parallel conflict decoupled data structure is generated, including the fused time slice number, the set of allocated pipeline segments, the execution order of each pipeline segment, and the resource allocation scheme. The data is stored as an XML format file. The root node records the scheduling period and the total number of resources, and the child nodes record the allocation information of each pipeline segment in the order of time slices. The attributes include the execution order and the dependence relationship mark.
[0114] The present invention also provides a memory-computation integrated parallel processing system for executing the memory-computation integrated parallel processing method as described above. The memory-computation integrated parallel processing system includes:
[0115] A multi-scale feature fusion module for obtaining the original stored data; performing topological skeleton projection on the original stored data to obtain a projection skeleton structure; and performing feature hierarchical integration on the original stored data according to the projection skeleton structure to generate multi-scale spatial fusion data;
[0116] A weight matching activation module for determining a weight relationship matrix based on the multi-scale spatial fusion data; performing gradient optimization adjustment on the multi-scale spatial fusion data through the weight relationship matrix, and performing memory-computation unit matching to obtain multiple memory-computation unit activation mapping data;
[0117] The interactive chain self-calibration module is used to mine the chained interactive correlation data between the activation mapping data of the memory-computation units; recursively self-calibrate the original stored data according to the chained interactive correlation data to generate self-calibration updated data;
[0118] The channel reconstruction flow control integration module is used to predict heterogeneous access channels based on the self-calibration updated data, and reconstruct each hierarchical data access link through the predicted heterogeneous access channels; perform multi-layer collaborative flow control cache integration according to the reconstructed hierarchical data access links to obtain access latency elimination and optimization data;
[0119] The task decoupling scheduling module is used to reorganize the dependency path tasks of the access latency elimination and optimization data to generate task reorganization parallel processing data; perform pipeline scheduling on the task reorganization parallel processing data to execute and eliminate parallel conflicts, thereby obtaining parallel conflict decoupling data.
[0120] In this invention, through the multi-scale feature fusion module, the acquisition of the original stored data is combined with the topological skeleton projection to form a clear projection skeleton structure. Feature hierarchical integration realizes multi-dimensional analysis of data on this basis. The generated multi-scale spatial fusion data improves the efficiency of data processing while retaining key features. The weight matching activation module provides a scientific basis for the gradient optimization of data based on the weight relationship matrix generated from the multi-scale spatial fusion data. The effective matching of the memory and computing units improves resource utilization. The activation mapping data of multiple memory-computation units realizes the basis of parallel processing through optimization and adjustment. The interactive chain self-calibration module enhances the correlation between data by mining the chained interactive correlation data between the activation mapping data of the memory-computation units. The updated data generated by recursive self-calibration ensures the consistency and accuracy of the original data. The channel reconstruction flow control integration module predicts heterogeneous access channels based on the self-calibration updated data, and the reconstructed hierarchical data access links optimize the data flow path. The multi-layer collaborative flow control cache integration formed after reconstruction effectively eliminates access latency, and the generated optimized data significantly improves the access efficiency. The task decoupling scheduling module creates conditions for parallel processing by reorganizing the dependency path tasks of the access latency elimination and optimization data. The generated task reorganization parallel processing data improves the system's response speed during execution. The pipeline scheduling mechanism effectively eliminates conflicts in parallel processing. The finally obtained parallel conflict decoupling data improves the overall system's processing ability. The entire system realizes efficient and intelligent data processing within the framework of memory-computation integration, provides reliable support for the execution of complex computing tasks, promotes the development of new memory-computation architectures, enhances the flexibility and adaptability of the system, optimizes the utilization of storage and computing resources, provides an innovative technical path for future multi-task parallel processing. The high efficiency and accuracy of the overall process lay a solid foundation for the application in large-scale data processing scenarios, promote the improvement of data processing capabilities in high-performance computing environments, and drive the wide application and development of memory-computation integration technology.
[0121] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be encompassed within the present invention.
[0122] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.
Claims
1. A parallel processing method integrating storage and computing, characterized in that: The following steps are involved: Step S1: Obtain original stored data; Perform topological skeleton projection on the original stored data to obtain the projection skeleton structure; perform feature hierarchical integration on the original stored data according to the projection skeleton structure to generate multi-scale spatial fusion data; Step S2: determining a weight relationship matrix based on multi-scale spatial fusion data; The multi-scale spatial fusion data is gradient optimized and adjusted through the weight relationship matrix, and memory-computing unit matching is performed to obtain activation mapping data of multiple storage and computing units; Wherein step S2 comprises the following steps: Step S21: performing memory mapping expansion on the multi-scale spatial fusion data to obtain memory mapping structure expansion data; Step S22: identifying patterns of the expanded data of the memory mapping structure, and classifying the identified patterns according to similarity to obtain classification space feature patterns; Step S23: constructing a weight matrix based on the classification space feature pattern to generate a weight relationship matrix; performing gradient optimization adjustment on the weight relationship matrix and performing structured sparsity projection to obtain structured sparse mapping data; Step S24: performing memory addressing optimization according to the structured sparse mapping data, and computing unit activation mapping process based on the addressing optimization data, thereby obtaining multiple memory computing unit activation mapping data; Step S3: mining chain-type interactive association data between the activated mapping data of the storage and computing units; performing recursive self-calibration on the original stored data according to the chain-type interactive association data to generate self-calibration update data; wherein mining chain-type interactive association data between the activated mapping data of the storage and computing units includes: Perform interactive channel monitoring on the storage and computing unit activation mapping data to obtain interactive channel probe data; Extracting activation coupling from interactive channel probe data to generate activation coupling correlation data; Perform chain fusion mapping based on activation coupling correlation data to obtain chain interaction correlation data; Step S4: predicting heterogeneous access channels based on the self-calibration update data, and reconstructing each hierarchical data access link through the predicted heterogeneous access channels; performing multi-layer collaborative flow control cache integration based on the reconstructed hierarchical data access links, thereby obtaining access delay elimination optimization data; Step S5: reorganize the dependent path tasks on the access delay elimination optimization data to generate task reorganization parallel processing data; perform pipeline scheduling on the task reorganization parallel processing data to eliminate parallel conflicts, thereby obtaining parallel conflict decoupling data.
2. The storage and computing integrated parallel processing method according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: performing channel domain division processing on the original stored data to obtain channel distribution segment data, wherein the channel division granularity range is set to a pixel block dimension of 8×8 to 64×64; Step S12: performing local topological tensor mapping processing on the channel distribution segment data to generate basic topological mapping data, with the mapping dimension set to a 3rd-order tensor structure and an element dimension range of 16-512; Step S13: Perform skeletonized structural projection on the basic topological mapping data to obtain a projected skeleton structure, wherein the structural connectivity threshold is set to 0.35-0.65 and the skeleton sparsity rate is controlled at 15%-40%; Step S14: Scale weight coupling is performed according to the projected skeleton structure and nested into a skeleton scale label; Step S15: performing hierarchical feature compression and fusion on the original stored data according to the skeleton scale label, thereby generating hierarchical feature integration data, wherein the fusion depth is set to 4 to 12 layers, and the compression ratio is controlled at 1:4-1:12; Step S16: Perform spatial redundancy regularization on the hierarchical feature integration data to obtain multi-scale spatial fusion data, wherein the redundant regularization window size is set to 5×5 to 11×11, and the residual filter threshold is set to 0.02-0.
12.
3. The storage and computing integrated parallel processing method according to claim 1, characterized in that: The recursive self-calibration of the original stored data according to the chain-type interactive correlation data in step S3 includes: Inferring interaction spike data from chained interaction correlation data; Perform time series reconstruction based on interactive pulse data to generate time series interactive reconstruction data; Reconstruct the data verification correlation context through time series interaction, thereby obtaining verification interaction correlation data; Perform feature residual decoupling on the verification interaction correlation data to obtain residual decoupling data; Perform autoregressive feature recursion on the residual decoupling data to generate recursive feature data; determining a decay state of verification interactive correlation data according to the recursive characteristic data; Synchronous gain adjustment is performed on the calibration cross-correlation data based on the attenuation state to generate self-calibration update data.
4. The storage and computing integrated parallel processing method according to claim 1, characterized in that: Predicting heterogeneous access channels based on the self-calibration update data in step S4 and reconstructing each hierarchical data access link by predicting the heterogeneous access channels includes: Extracting multi-domain channel attribute spectra from calibration update data; Construct channel cluster grouping structure based on multi-domain channel attribute spectrum; According to the channel cluster grouping structure and storage-computing level, the correspondence between channel clusters and data blocks is sorted out; Decompose the channel cluster-data block correspondence into multiple dynamically reconfigurable links; Predicting heterogeneous access channels based on dynamically reconfigurable links; Based on the preset storage capacity, the predicted heterogeneous access channel is adapted and reconstructed to generate a node adaptation link; Simulating a cache limit extension boundary through a node adaptation link, and determining an optimal cache layout based on the simulated cache limit extension boundary; The hierarchical links are reconstructed and integrated according to the preferred cache layout to obtain reconstructed hierarchical data access links.
5. The storage and computing integrated parallel processing method according to claim 1, characterized in that: The multi-layer collaborative flow control cache integration based on the reconstruction of each hierarchical data access link in step S4 includes: Reconstruct the node priority of each hierarchical data access link to generate a reconstructed data access link; Perform multi-source scheduling coupling processing on the reconstructed data access link to generate a cooperative scheduling link, where the channel bandwidth balance coefficient range is set to 0.4-0.8; Implement cache path cascade integration on the cooperative scheduling link data to obtain flow control cache integration data. During the cache path cascade integration process, the cache block granularity is set to 512KB to 2MB, and the hierarchical cascade structure does not exceed 4 layers. Based on the flow control cache, data is integrated to suppress node queue jitter, thereby generating a stable output cache. The jitter suppression operation is performed based on the rule that the maximum transient queue delay does not exceed 5ms. By stabilizing the output cache data and performing delay trace difference offset processing, access delay elimination optimization data is obtained.
6. The storage and computing integrated parallel processing method according to claim 1, characterized in that: The step S5 of reorganizing the dependent path tasks for the access delay elimination optimization data includes: Track the data flow path of the access delay elimination optimization data, and analyze the dependency path characteristics in the access delay elimination optimization data based on the data flow path; Identify bottleneck nodes based on dependency path characteristics; Determine the critical dependency path based on the bottleneck node; Decouple and reorganize the task units of the key dependency paths to obtain decoupled and reorganized task data; Perform dependency grouping aggregation on the decoupled and reorganized task data to generate dependency grouping aggregation data; Based on the dependency grouping and aggregation data, parallel mapping distribution is performed to obtain task reorganization and parallel processing data.
7. The storage and computing integrated parallel processing method according to claim 1, characterized in that: In step S5, the task reorganization and parallel processing data are pipelined and executed to eliminate parallel conflicts, including: Divide the task reorganization parallel processing data into pipeline segments and generate preliminary pipeline segment division data; Perform resource affinity matching based on the preliminary segmentation data of the pipeline to obtain resource affinity matching data; Perform parallel timing optimization on task reorganization and parallel processing data based on resource affinity matching data to generate timing optimized scheduling data; Detect potential conflict clusters on the timing optimization scheduling data, and perform conflict path stripping and reconstruction based on the potential conflict clusters to obtain conflict stripping and reconstruction data; The conflict stripping and reconstructed data are fused into conflict-free paths to obtain parallel conflict decoupled data.
8. A parallel processing system integrating storage and computing, characterized in that: For executing the storage-computing integrated parallel processing method according to claim 1, the storage-computing integrated parallel processing system comprises: The multi-scale feature fusion module is used to obtain the original stored data; perform topological skeleton projection on the original stored data to obtain the projection skeleton structure; and perform feature hierarchical integration of the original stored data according to the projection skeleton structure to generate multi-scale spatial fusion data; The weight matching activation module is used to determine the weight relationship matrix based on the multi-scale spatial fusion data; the multi-scale spatial fusion data is gradient optimized and adjusted through the weight relationship matrix, and memory-computing unit matching is performed to obtain activation mapping data of multiple memory-computing units; The interactive chain self-calibration module is used to mine the chain-like interactive correlation data between the activation mapping data of the storage and computing units; recursively self-calibrate the original stored data based on the chain-like interactive correlation data to generate self-calibration update data; The channel reconstruction flow control integration module is used to predict heterogeneous access channels based on self-calibration update data and reconstruct the access links of each hierarchical data through the predicted heterogeneous access channels. Based on the reconstructed access links of each hierarchical data, multi-layer collaborative flow control cache integration is performed to obtain access delay elimination optimization data. The task decoupling scheduling module is used to reorganize the dependent path tasks of the access delay elimination optimization data to generate task reorganization parallel processing data; perform pipeline scheduling on the task reorganization parallel processing data to eliminate parallel conflicts, thereby obtaining parallel conflict decoupling data.
Citation Information
Patent Citations
Service prediction method, device and equipment based on LNM large numerical model
CN119089156A
Real-time data processing method and chip of novel computing low-power AI processor
CN119806849A