Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

757 results about "Memory footprint" patented technology

Memory footprint refers to the amount of main memory that a program uses or references while running. The word footprint generally refers to the extent of physical dimensions that an object occupies, giving a sense of its size. In computing, the memory footprint of a software application indicates its runtime memory requirements, while the program executes. This includes all sorts of active memory regions like code segment containing (mostly) program instructions (and occasionally constants), data segment (both initialized and uninitialized), heap memory, call stack, plus memory required to hold any additional data structures, such as symbol tables, debugging data structures, open files, shared libraries mapped to the current process, etc., that the program ever needs while executing and will be loaded at least once during the entire run.

High-efficiency computer software system resource scheduling method, system, equipment and medium

The invention discloses a high-efficiency computer software system resource scheduling method, system and device and a medium, and the method comprises the steps: deploying a hardware counter in an operating system, and monitoring the CPU utilization rate, memory occupation, disk I / O, network bandwidth, the use condition of data aggregation system resources and the resource demand of each process or thread in real time; based on monitoring data, analyzing a current resource supply and demand state by adopting an advanced algorithm, and predicting a future resource demand trend; according to an analysis result, a resource allocation strategy is dynamically adjusted, it is ensured that a high-priority task obtains needed resources in time, and meanwhile fairness of a low-priority task is considered; data feedback after scheduling execution is collected, a scheduling algorithm is continuously optimized, and the scheduling efficiency is improved; the system, the equipment and the medium are used for realizing the high-efficiency computer software system resource scheduling method. The problems that an existing computer system is low in resource utilization rate, poor in key task stability and high in operation and maintenance cost are solved.
Owner:冯译萱

Edge computing node dynamic load balancing method based on multi-agent system

The invention discloses an edge computing node dynamic load balancing method based on a multi-agent system, and relates to the technical field of node load balancing. According to the edge computing node dynamic load balancing method based on the multi-agent system, the current computing power communication data and the operation environment data of each edge computing node are collected, the operation load index of each node is analyzed, and whether an overload node exists or not is judged; if an overload node exists, a plurality of nodes adjacent to the overload node are obtained, a task migration candidate set is screened out, finally, a target node is determined from the candidate set, and task migration operation is executed; the operation load index of the edge computing node is dynamically quantified by collecting and comprehensively analyzing multiple parameters such as the CPU utilization rate, the memory occupancy rate, the communication delay, the temperature rise rate, the I / O blocking rate, the power consumption fluctuation rate, the cache write-in waiting time and the vibration disturbance amplitude in real time. Compared with a traditional single index or fixed threshold mode, the method can identify the node overload state more accurately and timely.
Owner:YANCHENG AGRICULTURAL SCIENCE & TECHNOLOGY VOCATIONAL COLLEGE

Target recognition model reasoning optimization method and device

The invention provides a target recognition model reasoning optimization method and device, and the method comprises the steps: firstly carrying out the structural analysis and sensitivity evaluation of a pre-training model, extracting the structural features of each network layer, activating the distribution features, carrying out the quantitative sensitivity scoring, and constructing a data set reflecting the hierarchical features and fault-tolerant capability; and querying a quantitative configuration knowledge base based on the data set to generate a heterogeneous quantitative strategy. Layered low-bit quantization is executed according to the strategy, and a layered weighted loss function is introduced to carry out quantization perception training, so that precision loss caused by bit width compression is effectively compensated. According to the method, through hierarchical heterogeneous quantification, the model recognition precision is preserved to the maximum extent while high compression ratio and reasoning acceleration are achieved, and particularly, the performance of a high-sensitivity layer is protected. The generated heterogeneous quantitative model remarkably reduces memory occupation and power consumption, is suitable for an edge hardware platform with limited resources, forms a set of complete automatic process from analysis and configuration to training compensation, and has good universality and engineering practical value.
Owner:CHINA WEAPON EQUIP RES INST

A method of optimizing linear transformation

A method and system for optimizing compute runtime and memory footprint of a linear transformation process are provided. The method includes determining a set of optimal rotation parameters, wherein the optimal rotation parameters provide an optimal tradeoff between runtime compute resources and a memory footprint for a runtime execution of the linear transformation process; initializing the linear transformation process to run a boosting technique with the determined set of optimal rotation parameters, wherein the boosting technique, when executed at runtime as part of the linear transformation process, performs at least one iteration that yields rotated ciphertexts, and wherein the at least one iteration is based on the determined set optimal rotation parameters and at least one key switching key (KSK); and loading the initialized linear transformation process to an internal memory of a hardware accelerator.
Owner:CHAIN REACTION LTD

FPGA superposition processor acceleration system and method based on state space duality

The invention belongs to the field of machine learning, and discloses an FPGA (Field Programmable Gate Array) superposition processor acceleration system and method based on state space duality, which comprises a sparse predefined data acquirer, a reconfigurable systolic array, a partial sum cache, an element-by-element operation cache, a function calculation module, an on-chip memory management module and the like. Zero elements are eliminated through the sparse predefined data acquirer, redundant data are reduced, and redundant calculation is remarkably reduced; the reconfigurable systolic array flexibly supports multiple calculation modes, and the utilization rate of hardware resources is increased; the part and the cache realize cross-cycle accumulation and element-by-element operation cache integration results, the function calculation module completes nonlinear operation, and the on-chip memory management module optimizes result cache, so that on-chip operation of SSD calculation is ensured, off-chip memory access is remarkably reduced, and reasoning efficiency and energy efficiency are improved; by adopting the system, the memory occupation is effectively reduced, the element-by-element calculation efficiency is improved, and the sparse calculation redundancy is reduced, so that the reasoning process of the Mamba2 model is remarkably accelerated.
Owner:NINGBO ORIENTAL UNIV OF TECH (TEMPORARY NAME)

Retraining-free pruning and recombination method and system for sparse expert hybrid large model

The invention discloses a retraining-free pruning and recombination method for a sparse expert hybrid large model, and belongs to the technical field of large model compression and optimization. The method aims at solving the problems that due to the fact that an existing sparse expert hybrid (SMoE) model needs to load all expert parameters, memory occupation is too high, and deployment is difficult. According to the method, firstly, redundant experts are identified and pruned based on routing activation statistics; then, decomposing the pruned experts into neuron-level functional fragments, and redistributing the fragments to the reserved experts according to structural similarity; and finally, original fragments and newly distributed fragments are merged in the reserved experts through a weighted clustering algorithm, so that compact experts with fewer parameters and stronger expression ability are reconstructed. According to the method, fine-grained operation is carried out at the neuron level, the inherent representation conflict and dislocation problems among experts are effectively solved, the performance of the compressed model is remarkably improved, and reliable technical support is provided for deploying a large-scale SMoE model.
Owner:ZHEJIANG UNIV

Drawing method for making three-dimensional model based on three-dimensional laser point cloud

The invention relates to the technical field of three-dimensional models, in particular to a drawing method for making a three-dimensional model based on three-dimensional laser point clouds, which comprises the following steps of: sequentially reading point cloud data files of the three-dimensional laser point clouds, inserting coordinate values of each point into an octree structure, and when octree node buffer areas in a memory are full, drawing a three-dimensional laser point cloud into the octree structure; and writing the node data in the buffer area into a hard disk, and traversing the hard disk octree node set from bottom to top. According to the method, the three-dimensional laser point cloud data is inserted step by step, and the hierarchical aggregated octree structure is utilized, so that the point cloud data processing efficiency is effectively improved, and the memory occupation pressure is reduced; representative points in an octree structure are fused with an average normal to construct a continuous function field, a grid model is adaptively generated in a multi-detail-level mode, and the detail retaining capacity and drawing precision of the model are improved. Furthermore, accurate identification and positioning of topological features are realized by adopting calculation of a cell complex sequence and a topological feature noise rank.
Owner:SHANDONG ZHIWEI SURVEY PLANNING & DESIGN CO LTD

Computing power service dynamic resource allocation method and system applied to AI model training

The invention provides a computing power service dynamic resource allocation method and system applied to AI model training. The method comprises the following steps: firstly, collecting real-time computing power resource use data (including computing node load, memory occupation and data transmission delay) and model training state data (including training task stage identification, model parameter updating frequency and training data batch processing progress) in AI model training; generating a computing power resource demand association feature set, constructing a computing power resource dynamic allocation decision model including resource allocation priority judgment, adjustment amplitude calculation and scheduling opportunity selection units based on the set, and outputting a computing power resource allocation scheme (including computing node number, memory capacity and data transmission bandwidth adjustment instructions) through the model. Resources are scheduled according to the scheme, and new data are collected to update the feature set, so that dynamic and accurate allocation of computing power resources is realized, and the resource utilization rate and the training efficiency are improved.
Owner:SICHUAN BOCHUANGHUI FRONTIER TECH CO LTD +1

Main Monitor service pressure optimization method in distributed storage cluster

The invention provides a main Monitor service pressure optimization method in a distributed storage cluster, and the method comprises the steps: S1, collecting the load state information of each Monitor node in real time through a monitoring module, and enabling the load state information to comprise the CPU utilization rate, the memory occupancy rate and the network bandwidth occupancy rate; s2, a load balancing module calculates the load weight of each Monitor node based on the load state information, and dynamically adjusts a message distribution strategy according to the load weight; s3, the multi-stage message distribution module distributes the state updating message to a plurality of Monitor nodes for parallel processing according to the message distribution strategy; and S4, the state synchronization module maintains the state consistency among the Monitor nodes through a consistency protocol, and regularly verifies the data version of each node. According to the main Monitor service pressure optimization method in the distributed storage cluster provided by the embodiment of the invention, the load pressure of the main Monitor node in the distributed storage upgrading process is effectively reduced, and the system message processing efficiency and stability are improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Computing resource scheduling method and device, electronic equipment and storage medium

The invention provides a computing resource scheduling method and device, electronic equipment and a storage medium, and the method comprises the steps that a computing task request issued by a user is received, and the computing task request comprises a computing unit mask parameter; a scheduling instruction is generated according to the calculation task request, the scheduling instruction is sent to a command processing module, and the scheduling instruction comprises the calculation unit mask parameters; and analyzing the calculation unit mask parameter in the scheduling instruction by a command processing module, and distributing a calculation task to a target calculation unit set according to an analysis result. According to the method and the device, static binding of masks and task flows / queues in a traditional scheme is replaced by dynamic integration of the mask parameters and the scheduling instructions, the situation that independent task flows / queues are created for each mask combination is avoided, and memory occupation and system management overhead can be remarkably reduced.
Owner:SUZHOU YIZHU INTELLIGENT TECH CO LTD

Data verification method and device for redundant disk array, equipment and medium

The invention discloses a data verification method and device for a redundant disk array, equipment and a medium, and relates to the technical field of data verification, when a data verification value of target data is generated, the target data is firstly divided into a plurality of data blocks, and correspondingly reference data is also divided into one-to-one corresponding reference blocks; then, the unit verification value of each data block is determined by taking the data block as a unit, and the data verification value of the target data is obtained through block-by-block calculation, so that the data verification of the target data is realized according to the data verification value, the memory occupation in the data verification value generation process is reduced, the storage requirement of temporary data is reduced, and the data verification efficiency is improved. The high memory overhead of processing the whole target data at a time is avoided, and the calculation efficiency of the data verification value is improved. The size of the data block can be flexibly adjusted according to requirements, the application range is wide, the flexibility is high, the check and problem positioning of the stored data are facilitated, and the troubleshooting of abnormal conditions such as data damage of the stored data is facilitated.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

System and method for large-scale video access and AI reasoning enhancement

The invention discloses a large-scale video access and AI reasoning enhancement system and method, belongs to the technical field of video processing, and aims to solve the technical problem of how to realize large-scale video stream efficient access, resource elastic scheduling and abnormity self-healing in a complex environment. Comprising a resource dynamic scheduling module, a decoding and reasoning control module, a resolution dynamic processing module, a multi-stage shared memory transmission module, an exception self-healing and fault-tolerant module and a task parallel execution module, and resource allocation is optimized and calculated by dynamically binding a CPU core and GPU hardware unit isolation; the resolution is dynamically adjusted based on the scene algorithm precision requirement, and invalid calculation is reduced; a stream pushing and frame rate decoding strategy is controlled by using a Redis mark, and memory occupation is reduced by combining long and short queues; transmission coding and decoding are reduced by adopting a shared memory, and the cross-process interaction efficiency is improved; and task self-healing is realized through dual anomaly detection and process-level heartbeat monitoring.
Owner:INSPUR QILU SOFTWARE IND

Mass small file reading optimization method, system, equipment and medium

The invention relates to the technical field of big data, and discloses a massive small file reading optimization method, system and device and a medium, and the method comprises the following steps: scanning small files in a distributed file system by using a Spark calculation engine to obtain metadata information of the small files; based on the metadata information, small files are classified, and the small files with the same type and data relevance are classified into one group; the resource states of cluster nodes are monitored in real time, the resource states comprise the processor utilization rate, the memory occupancy rate and the I / O load, the load score of each node is calculated, and small file task quotas are dynamically allocated; executing small file merging operation in parallel by utilizing a Spark memory calculation engine to generate a merged file containing a plurality of original small files; creating and maintaining index information in an external database, and recording the position and attribute of each original small file in the combined file; and providing a read-write interface through the special service layer, and positioning and accessing specific small file contents in the combined file.
Owner:GUANGDONG POLYTECHNIC NORMAL UNIV

Display device and sequence animation playing method

The invention provides a display device and a sequence animation playing method, and the method can respond to a sequence animation playing instruction, obtain an updating period of a sequence animation, read a target bitmap object from a bitmap set associated with the sequence animation according to the updating period, and update an image frame displayed by an animation view according to the target bitmap object. Wherein the sequence animation comprises a plurality of continuous image frames, the bitmap set comprises a plurality of bitmap objects, the bitmap objects are used for representing the image frames of the sequence animation loaded into the running memory from the memory in advance, and the bitmap set is periodically reset according to the sequence animation stored in the memory; the target bitmap object is used for representing an image frame to be displayed by the animation view in the current updating period. According to the method, the bitmap set in the memory can be periodically loaded, the occupation optimization of the memory is realized, and the problem that the occupation of the memory is continuously increased when the sequence animation is played is solved.
Owner:HISENSE VISUAL TECH CO LTD

Deep learning inference engine tensor optimization method and system for GPU acceleration

The invention relates to the technical field of data processing, and provides a deep learning inference engine tensor optimization method and system for GPU acceleration. Performing asymmetric quantization on the first weight tensor of the original model to generate a second weight tensor, and updating the original model based on the second weight tensor to obtain an updated model; blocking the data tensor according to a preset dimension to generate input sub-blocks, and loading the input sub-blocks and a convolution kernel in the update model to a shared memory; calling the input sub-blocks from the shared memory, and processing the input sub-blocks of different batches in parallel based on a single kernel function formed by convolution kernels to generate a processing result; and based on the GPU utilization rate and the request queue length corresponding to the to-be-processed data, performing dynamic output adjustment to output a processing result. The performance and efficiency of a deep learning inference engine are improved, memory occupation and calculation delay are reduced, and powerful support is provided for deep learning application of GPU acceleration.
Owner:SHENZHEN HUANDONG INTELLIGENT TECH CO LTD

Structured sparsity guided training in an artificial neural network

A novel and useful system and method of improved power performance and lowered memory requirements for an artificial neural network based on packing memory utilizing several structured sparsity mechanisms. The invention applies to neural network (NN) processing engines adapted to implement mechanisms to search for structured sparsity in weights and activations, resulting in a considerably reduced memory usage. The sparsity guided training mechanism synthesizes and generates structured sparsity weights. A compiler mechanism within a software development kit (SDK), manipulates structured weight domain sparsity to generate a sparse set of static weights for the NN. The structured sparsity static weights are loaded into the NN after compilation and utilized by both the structured weight domain sparsity mechanism and the structured activation domain sparsity mechanism. The application of structured sparsity lowers the span of search options and creates a relatively loose coupling between the data and control planes.
Owner:HAILO TECH LTD

Intelligent screen multi-terminal collaborative interaction system and method based on edge computing

The invention relates to the technical field of intelligent terminals, in particular to an intelligent screen multi-terminal collaborative interaction system and method based on edge computing. The intelligent screen multi-terminal collaborative interaction system based on edge calculation comprises a terminal display unit, an edge calculation unit and an intelligent screen display unit, according to the method, feature information is extracted by integrating multiple factors, terminal display requirements and importance are accurately reflected, important information is judged according to priority permission levels, judgment is performed according to the feature information when the levels are the same or no level exists, the result is more reasonable, random or average distribution display is avoided, meanwhile, the memory occupancy rate of the smart screen is monitored in real time, and the user experience is improved. In addition, the historical information features are continuously updated, and the intelligence and adaptability of the whole system are improved.
Owner:ZHONGKE WANLI (SHENZHEN) TECHNOLOGY CO LTD

Multi-modal data active recommendation method based on cross-modal dynamic fusion and equipment resource awareness

The invention provides a multi-modal data active recommendation method based on cross-modal dynamic fusion and equipment resource perception. The method is applied to the technical field of data processing. The method comprises the following steps: acquiring multi-modal data through a smart home device end; performing modal feature extraction on the multi-modal data to generate a voice feature matrix, an environment probability matrix and an equipment state vector matrix; adopting a cross-modal dynamic fusion algorithm, combining time dimension and space dimension weights, and carrying out adaptive weighted fusion on the voice feature matrix, the environment probability matrix and the equipment state vector matrix; the CPU load and the memory occupancy rate of the central control equipment are detected in real time through the equipment resource monitoring module, and a data calculation strategy is dynamically adjusted; and based on the fused multi-modal feature matrix, generating an active recommendation instruction through user portrait modeling and knowledge graph association, and driving the smart home device to execute adaptive adjustment. In this way, the recommendation efficiency can be improved.
Owner:ULTIMATE IOT (HENAN) TECHNOLOGY LTD +1

Page rendering method and device, terminal equipment and readable storage medium

The invention relates to the technical field of page rendering, and discloses a page rendering method and device, terminal equipment and a readable storage medium. According to the page rendering method, page structure information of a to-be-rendered page is subjected to blocking processing, and an initial data block set is obtained; carrying out lazy loading processing on the initial data block set to obtain a target data block set; performing parameter analysis on the target data block set to generate structured rendering data; performing paging decision according to the structured rendering data to obtain page paging information; constructing a page rendering instruction set according to the page paging information; and performing drawing processing based on the page rendering instruction set to generate a visual page rendering result. According to the page rendering method, the page rendering efficiency and the cross-page typesetting accuracy are improved while the memory occupation is effectively reduced.
Owner:KINCHENG BANK OF TIANJIN CO LTD

Translation model training method, information translation method, system and related product

The embodiment of the invention discloses a translation model training method, an information translation method, an information translation system and a related product. The training method comprises the following steps: performing low-rank decomposition on an original matrix to obtain an initial low-rank matrix; performing iterative training on the initial low-rank matrix by using each source language text and each target language text to obtain a new low-rank matrix different from the initial low-rank matrix; and updating partial model parameters of the large language model by using the new low-rank matrix to generate a translation model for translating the to-be-translated text. The initial low-rank matrix is introduced to approximately simulate all parameters of the large language model, the number of parameters needing to be trained can be reduced, the calculation complexity and memory occupation in the training process can be reduced, the cost is reduced, meanwhile, the large language model can be efficiently and finely adjusted, and therefore the translation model which can be specially suitable for translation tasks is generated; and the manual workload and translation time are reduced.
Owner:SANGFOR TECH INC

CXL memory fault tolerance method, and server system, storage medium and electronic device

PCT designated stageWO2025227987A1TransmissionRedundant hardware error correctionMemory faultsMemory footprint
A CXL memory fault tolerance method, and a server system, a storage medium and an electronic device. The method comprises: acquiring parameter values of a group of operating parameters of CXL memory devices in a CXL memory device group, wherein the group of operating parameters are used for representing the operating states of the corresponding CXL memory devices; on the basis of the acquired parameter values of the group of operating parameters, predicting the operating states of the CXL memory devices in the CXL memory device group; and when it is predicted that there is an abnormal memory device operating abnormally in the CXL memory device group, performing controlling to execute a migration operation on memory data in the abnormal memory device, so as to migrate the memory data in the abnormal memory device to a target memory device operating normally in the CXL memory device group. By means of the present application, the problem of CXL memory fault tolerance methods in the prior art of the memory utilization rate of a server being low due to a hot standby memory occupying a server slot is solved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Optimization method and device of embedded linked list container, electronic equipment and storage medium

The invention discloses an optimization method and device for an embedded linked list container, electronic equipment and a storage medium, and relates to the technical field of computers, the method comprises the steps that a linked list node is embedded into a host data structure to serve as a member variable, extra memory occupation of a pointer domain in a traditional linked list is omitted, and the optimization efficiency is improved. Physical storage of the nodes and the host data structure is continuous, so that cache missing can be reduced; meanwhile, the stability of the linked list under high concurrency is guaranteed through atomized insertion and deletion operations, and a traditional independent node structure and a non-atomized operation are not adopted; the technical problems that in the prior art, a traditional linked list node pointer domain occupies an extra memory, so that expenditure is increased, discontinuous node physical distribution causes cache missing, access delay is increased, and system abnormal safety and operation reliability are affected by iterator failure under high concurrency can be solved. And the technical effects of reducing the memory overhead, reducing the access delay and improving the abnormal safety and the operation reliability of the system are achieved.
Owner:JINAN INSPUR DATA TECH CO LTD

Big data hierarchical encryption filtering mechanism based on TEE, query optimization method, system, equipment and medium

The invention relates to a TEE-based big data hierarchical encryption filtering mechanism and query optimization method, system, equipment and medium, and the method comprises the following steps: firstly, in the process of writing Parquet data into a disk, performing hierarchical encryption processing on the data in the TEE through a user-defined Spark data source interface, and generating a multi-level metadata index; then, based on data hierarchical encryption processing and the generated multi-level metadata index, data query is carried out, and in the query execution process, a Spark query engine firstly receives an SQL query request of a user, and parses a plaintext predicate in a query condition into a ciphertext predicate; then, the query execution process enters a layered encryption filtering stage, and partition-level, file-level, row-group-level and column-level screening operations are executed in sequence, so that the ciphertext decryption calculation amount in the TEE is reduced; the system, the equipment and the medium realize the big data hierarchical encryption filtering mechanism and the query optimization based on the big data hierarchical encryption filtering mechanism and the query optimization method of the TEE; according to the method, the EPC memory pressure caused by full-disk decryption of a traditional TEE scheme is avoided, Parquet data query can be efficiently executed under the condition of relatively low memory occupation, and finally, the balance of privacy protection and query optimization is realized.
Owner:XIDIAN UNIV

Dynamic audio buffer management method and system and medium

The invention relates to the technical field of software, and discloses a dynamic audio buffer management method and system and a medium. The method comprises the following steps: presetting a white list library of a plurality of application scenes, and matching a package name of a current foreground running application with the white list library to determine a current application scene; acquiring system performance state parameters in real time, wherein the parameters comprise at least one of a CPU occupancy rate, a memory occupancy rate and a temperature control and frequency limiting state; generating a buffer adjustment variable according to a statistical result of the audio lagging times; inputting the current application scene, the system performance state parameter and the buffer area adjustment variable into a preset strategy model, and outputting a buffer area gear value; and dynamically adjusting the size of a buffer area of an audio driving layer according to the gear value so as to adapt to the performance change of the system. According to the invention, the problem of lagging or delay caused by a fixed buffer area can be solved.
Owner:SHANGHAI LONGCHEER TECH CO LTD

Text coding method and device, computer equipment and storage medium

The invention relates to the technical field of artificial intelligence, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a text coding method and device, computer equipment and a storage medium, and the method comprises the following steps: obtaining an input text sequence, and generating a key vector matrix corresponding to the input text sequence, the key vector matrix comprises a query matrix, a key matrix and a value matrix; obtaining the attention degree of each element of the input text sequence, and pruning the query matrix based on the attention degree and the query matrix; and based on the query matrix after pruning processing, the key matrix and the value matrix, calculating an attention score of the input text sequence so as to encode the input text sequence. According to the method and the device, the technical problem that the operand of attention calculation and memory occupation cannot be reduced while accurate text coding cannot be ensured in the prior art is solved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Dynamic resource scheduling method and system based on reinforcement learning

The embodiment of the invention provides a dynamic resource scheduling method and system based on reinforcement learning, and the system comprises a state sensing module which is used for collecting and preprocessing the resource state data of each node in a cluster, and the resource state data at least comprises a CPU utilization rate, a memory occupancy rate, a network bandwidth utilization rate, a task queue length and a node load; the action decision module is used for outputting a scheduling action according to the current state representation vector, and the reward calculation module is used for calculating a reward value according to an actual operation result of the system; the model training module is used for training the strategy model by using a reinforcement learning algorithm, and the deployment optimization module is used for deploying the trained strategy model to a production environment. Complicated and changeable workloads and resource states in the cloud environment can be automatically dealt with, the manual intervention cost is remarkably reduced, and the scheduling efficiency and accuracy are improved. And a comprehensive and unified environment perception capability can be constructed.
Owner:HUANENG ZHAOCAI DIGITAL TECHNOLOGY CO LTD +1

Redundant code detection method, electronic equipment, storage medium and program product

The invention discloses a redundant code detection method, electronic equipment, a storage medium and a program product, and relates to the technical field of software development. An enhanced abstract syntax tree is generated according to a full analysis result of the metadata information containing the dynamic dependency relationship between the codes, and then a code dependency graph is constructed according to the enhanced abstract syntax tree. Generating a dynamic behavior sequence according to dynamic behavior features extracted in a code operation process, and performing multi-modal feature fusion on the code dependency graph and the dynamic behavior sequence by using a graph neural network and time convolutional network hybrid model to obtain redundancy probabilities corresponding to nodes in the code dependency graph; the technical problems that the omission ratio is too high, the granularity is too coarse and misjudgment is serious are solved, and the technical effects that the detection precision of the redundant codes is improved, the detection accuracy of the redundant codes is improved, and memory occupation is reduced are achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

KV-cache streaming for improved performance and fault tolerance in generative model serving

A method of serving a generative transformer model includes determining a batch size to use in processing inference requests and allocating at least one prompt pipeline and at least on token pipeline to the generative transformer model to process the batch of inference requests. The number of prompt pipelines and the number of token pipelines, and the depths of the pipelines are determined based on the batch size, an average prompt length, a cache requirement per stage, and a memory footprint of model weights for the generative model per stage using a resource allocator component of the model serving system. Cache streaming is used to stream prompt cache from prompt pipelines to token pipelines to generate tokens. Cache streaming involves gather-copy operations which may be performed using compute kernels.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

RISC-V simulation resource dynamic generation method and system

The invention belongs to the technical field of integrated circuit simulation verification, particularly relates to an RISC-V simulation resource dynamic generation method and system, and solves the problems of instruction consistency and variable-length instruction truncation through a metadata double-table structure and a cross-boundary instruction splicing mechanism. Virtualized two-stage translation support and abnormal injection are realized through a recursive multiple hit detection and dynamic attribute bit modification mechanism. According to the method, complete storage is replaced with lightweight metadata, memory occupation is remarkably reduced, the problem that memory occupation is linearly increased along with time is solved, simulation efficiency and consistency are improved, and the method is suitable for full-system verification of a high-performance RISC-V processor.
Owner:SHANDONG UNIV

Audio test data processing method, system and equipment based on zero-copy NIO and dynamic sliding window

The invention discloses an audio test data processing method, system and device based on zero-copy NIO and a dynamic sliding window, and the method comprises the following steps: S1, building a memory mapping file channel, and directly mapping an audio file to a process address space; s2, constructing a double-buffer processing pipeline, and realizing asynchronous decoupling of a collection thread and a calculation thread by adopting a producer-consumer mode; and S3, a dynamic sliding window mechanism is adopted to maintain a real-time audio data interval, and the window size is dynamically adjusted according to a television CPU load. According to the technical scheme of the invention, the problems of high memory occupation, high detection delay and low cross-application data reading efficiency existing in audio testing of a television system in a multi-application concurrent scene are solved.
Owner:PANOVASIC TECHNOLOGY CO LTD