Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

60 results about "Computational RAM" patented technology

Computational RAM or C-RAM is random-access memory with processing elements integrated on the same chip. This enables C-RAM to be used as a SIMD computer. It also can be used to more efficiently use memory bandwidth within a memory chip.

Compute in memory three-dimensional non-volatile NAND memory for neural networks with weight and input level expansions

A non-volatile memory device for performing compute in memory operations for a neural network uses a three dimensional NAND architecture. Multi-bit weight values are stored encoded as sets of threshold voltages for sets of memory cells. A weight value is stored in multiple memory cells on the same word line and connected between a bit line and a source line, each of the memory cells programmed to one of multiple threshold voltages. When multiplying an input value with the weight value, the word line is biased so that, for at least one of the threshold voltages, the memory cell will be in the linear operation region. Input values are encoded as a set of one or more voltage levels applied to a corresponding set of bit lines, each bit line connected memory cells also storing the weight value, connected to the word line, and connected to the source line.
Owner:SANDISK TECHNOLOGIES LLC

Static random-access memory (SRAM) device and related SRAM-based compute-in-memory devices

An SRAM cell includes a first inverter cross-coupled to a second inverter. The first inverter includes a first pull-up transistor and a first pull-down transistor, having coupled drains that define a first storage node. The SRAM cell further includes a first N-type pass-gate transistor having a first drain coupled to a write bit line, a first source coupled to the first storage node, and a first gate coupled to a first write word line. The SRAM cell further includes a first P-type pass-gate transistor having a second drain coupled to the write bit line and a second source coupled to the first storage node. The SRAM cell further includes a P-type transistor having a third drain, coupled to a second gate of the first P-type pass-gate transistor, a third source coupled to a second write word line, and a third gate coupled to an enable signal.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Multi-channel data acquisition system and method based on FPGA (Field Programmable Gate Array)

The invention relates to the technical field of signal processing, and discloses a multichannel data acquisition system and method based on an FPGA (Field Programmable Gate Array), an acquisition module drives a global counter by using a global synchronous clock, writes acquired data into an annular buffer area with a physical address and a counter value in a linear mapping relationship, and establishes implicit time index storage. When a trigger event occurs, the system broadcasts the locked global trigger timestamp, and the acquisition module backtracks and reads historical data according to the global trigger timestamp, and packages the historical data into a sparse matrix type data packet in combination with the channel validity mask. And after receiving the data packet, the data processing module directly calculates a memory mapping address by using global time information in the data packet, and writes a data load into a corresponding position of the waveform reconstruction buffer area. According to the method, through strict binding of the physical address and the absolute time, independent time label redundancy is eliminated, high-precision synchronization of distributed multiple channels is ensured, and out-of-order automatic in-situ recombination and efficient waveform reconstruction of data of a receiving end are achieved.
Owner:CHANGCHUN TESTING MASCH RES INST

Adaptive gateway flow scheduling method based on multi-dimensional information

The invention discloses a multi-dimensional information-based adaptive gateway traffic scheduling method, which comprises the following steps of: acquiring server performance indexes, network delay, bandwidth and geographic position data in real time, and establishing a historical database; performance load weights are generated through weighted calculation of CPU, memory and disk indexes, and routing priorities are dynamically adjusted by adopting exponential function mapping; performing normalization processing on network bandwidth, delay and geographic distance, and calculating a service access weight to optimize a transmission path in combination with a step function; analyzing historical traffic data based on an LSTM model to predict a future trend, and generating a traffic prediction weight through a sigmoid function to realize pre-equalization; and combining the three types of weights, weighting to generate dynamic routing configuration, eliminating abnormal nodes in real time in combination with a health check mechanism, triggering high-load early warning and linking a predictive balancing strategy. According to the method, multi-dimensional index collaborative decision-making is realized, routing distribution is dynamically optimized, network delay is effectively reduced, and system throughput and stability in a high-concurrency scene are improved.
Owner:河北省体育局射击射箭运动中心(河北省军事体育运动学校) +1

Network security functions for dynamic construction and programmatic placement

A system and method are provided for placing network functions among respective locations in a network. The locations at which the network functions are placed can be nodes and network devices within the network. These nodes can be selected, e.g., based on which network devices have available capacity and or specialized hardware (e.g., accelerator sin a data processing units (DPUs)) that is optimized for particular network functions. The network functions can include an inline network function that is provisioned directly in a data plane of one of the network devices (e.g., in-lined directly in a hardware offload device without a virtual machine and without a container). The decision of where to place the network functions can be based on a performance metric (e.g., representing available computational / memory resources at the network nodes) and / or a network-function metric (e.g., representing consumed computational / memory resources by the network functions) to improve system performance.
Owner:CISCO TECHNOLOGY INC

Edge server load evaluation method based on analytic hierarchy process and performance loss

The invention relates to the technical field of edge computing resource management, in particular to an edge server load evaluation method based on analytic hierarchy process and performance loss. According to the method, firstly, edge servers are divided into a common type, a computing type, a memory type and an I / O type according to heterogeneous types; aiming at each type of edge server, constructing a judgment matrix by adopting an analytic hierarchy process and solving a feature vector so as to calculate load weight coefficients of CPU (Central Processing Unit), memory and disk I / O (Input / Output) resources; calculating a real-time load based on the weight coefficient, wherein the real-time load is the weighted sum of the resource utilization rates; the performance loss is further calculated according to the nonlinear relation between the real-time load and the total load ratio; and finally, restraining the real-time load through the load threshold value, and evaluating the performance state of the server in combination with the performance loss. According to the method, the accuracy and reliability of load evaluation of the heterogeneous edge server are effectively improved, the resource utilization rate is optimized, overload of the server is avoided, and the system stability is enhanced.
Owner:GUIZHOU INST OF TECH

High-concurrency virtual service transfer scheduling method and system based on cloud edge collaborative architecture

The invention discloses a high-concurrency virtual service transfer scheduling method and system based on a cloud edge collaborative architecture, and relates to the technical field of cloud computing and edge computing collaboration, and the method comprises the steps: constructing a service gene analysis model at an edge node, responding to a virtual service access request, extracting the resource consumption characteristics of a service flow in real time, and calculating a service separation potential energy index. The virtual business is dynamically decoupled into a connection anchoring phase and a computing power free phase based on the index; establishing a load inertia monitoring mechanism, modeling resource consumption of edge nodes into a kinematics system, collecting speed and acceleration of load change in real time, calculating system pressure momentum, and generating a graded unloading instruction when the momentum exceeds a preset critical collapse threshold value; starting a convergent differential injection mechanism, pre-constructing a logic shadow at a cloud end, and calculating a convergence ratio of a memory dirty page generation rate to a network transmission rate in real time; the method has the advantages of being high in practicability, having the pre-judgment capacity and being capable of achieving lossless switching.
Owner:JIANGSU AIYUSEN INFORMATION TECHNOLOGY CO LTD

Fully analog compute-in-memory architecture for neural networks

A Fully Analog STate-Space Compute-In-Memory (FAST-CIM) architecture implements cascaded CIM arrays that maintain continuous analog signal flow throughout neural network computations. Sequential data patches are processed through cascaded CIM arrays, where a feedforward CIM array performs initial vector-matrix multiplication using analog computations within memory cells, and the analog output flows directly through a gain circuit to a recurrent CIM array without digital conversion. The gain circuit converts analog current signals to voltage signals, enabling direct analog connection between the cascaded arrays. The recurrent CIM array executes State-Space Model (SSM) computations, processing the cascaded analog signals along with previous state information maintained by a State Write and Propagate (SWAP) circuit that stores states using capacitive elements. The cascaded array configuration eliminates analog-to-digital and digital-to-analog converters between processing stages, achieving reduced power consumption and latency compared to conventional implementations that require digital interfaces between CIM arrays.
Owner:GEORGIA TECH RES CORP

Apparatus, engine, system and method for predictive analytics in a manufacturing system

A predictive analytics apparatus, engine, system and method capable of providing real time analytics in a manufacturing system that may include a data input capable of receiving raw data output from at least one machine operable to effect the manufacturing system embodiments, and a processor to execute code from a computing memory. The code may comprise an adaptor to push the received raw data to a database to processed data; an extractor to extract the processed data from the database; predictive analytics to receive the extracted processed data and apply thereto a predictive model comprised of target data for the at least one machine, and to provide feedback to the at least one machine to modify performance of the at least one machine based on the application of the predictive model; and a visualizer capable to provide at least a visualization of the feedback and the performance.
Owner:JABIL INC

A large file export optimization method and system based on real-time memory monitoring

The application relates to the technical field of data processing, and discloses a large file export optimization method and system based on real-time memory monitoring. The method comprises the following steps: monitoring a memory state in real time and calculating a memory growth rate; taking the growth rate as an independent risk dimension, and giving a warning when the rate exceeds a threshold value and the memory usage rate does not reach a threshold value; performing adjustment in response to the warning, including dynamically updating the throughput control of the data shard size, and triggering the degradation mode to reduce the export field to the data precision control of the core field; and attaching an instruction page to the export file during degradation. Through trend prediction and dynamic control, the application reduces the risk of memory overflow and guarantees core data export.
Owner:山东齐鲁壹点传媒有限公司 +1

A kinetic solution method and system for multi-scale particle transport simulation

This invention belongs to the technical field of multi-scale particle transport simulation, and discloses a kinetic solution method and system for multi-scale particle transport simulation. The method includes: a particle kinetic model based on discrete velocity form; selecting discrete points in velocity space; calculating the distribution function along the direction of the discrete points in velocity space at the center of the grid according to a compact scheme; accumulating the contribution values ​​of the distribution function to macroscopic physical quantities; scanning the downstream grid sequentially based on the direction of the discrete points in velocity space; changing the discrete points in velocity space and repeating the operation until the contribution values ​​of the grid at all discrete points in velocity space are accumulated, and obtaining the final macroscopic physical quantities; outputting the calculation results of macroscopic physical quantities when the convergence condition is met. This invention discloses a kinetic solution method and system for multi-scale particle transport simulation that balances computational memory, accuracy, and efficiency, and is particularly suitable for multi-scale particle transport simulation with a large number of discrete points in velocity space and strong heterogeneity.
Owner:HUAZHONG UNIV OF SCI & TECH

Data processing method, distributed system, electronic device, and storage medium

Embodiments of the present disclosure provide a data processing method, a distributed system, an electronic device and a storage medium. The method is applied to a distributed system comprising a master node and a plurality of slave nodes, and comprises: the master node determining task information allocated to each slave node according to a hierarchical attribute of a target model and a computing power state of the plurality of slave nodes, wherein the task information is used to specify a model layer processed by the corresponding slave node; each slave node performing weight quantization processing of the corresponding model layer based on the allocated task information to obtain complete quantized weights of the target model. The method realizes full-link collaborative optimization from computing, memory to storage, significantly improves the execution efficiency of weight quantization tasks and the utilization rate of hardware resources, and provides technical support for efficient deployment of large-scale models.
Owner:SHANGHAI BIREN TECH CO LTD

Cross-data center large model training system architecture and resource allocation method and system

PendingCN121957895AImplement collaborative trainingEfficient collaborative utilizationResource allocationBiological modelsWide areaData center
The invention provides a system architecture for cross-data center large model training and a resource allocation method and system, and belongs to the technical field of cross-wide area distributed large model training. According to the method, large-scale model cooperative training across multiple data centers can be realized, and the bottleneck that the computing power of a single data center is limited is broken through. Through unified modeling and scheduling of calculation, memory and network resources, task loads can be intelligently allocated according to hardware performance and network bandwidth of different data centers, and efficient collaborative utilization of computing power resources is realized. The training task of the super-large-scale model can be rapidly completed in the heterogeneous computing power environment, and the training time is remarkably shortened. The provided flexible parallelism degree allocation method can be adaptive to different task and resource conditions, the proportion of data parallelism, model parallelism and pipeline parallelism is automatically adjusted, the parallelism efficiency is improved, and the communication overhead is reduced. A training time estimation function is integrated, the overall time delay and resource requirements can be predicted before task execution, and a basis is provided for scheduling decision making.
Owner:BEIJING JIAOTONG UNIV

A category-based 6g network multi-dimensional resource ai model dynamic deployment optimization method

The application discloses a kind of 6G network multidimensional resource AI model dynamic deployment optimization methods based on category theory, belongs to intelligent collaborative optimization technical field;Method is: the cross-layer consistency dependency of end-to-end AI reasoning service is formalized by functor form;Establish the joint optimization model with long-term average end-to-end delay minimization as target, while being constrained by multidimensional resource and service quality;Convert long-term random optimization problem into time-slot online decision problem;Get AI model dynamic deployment and task scheduling result.The application realizes cross-layer consistency description to task scheduling and model deployment through category theory unified modeling and functor composite mechanism, reduces the inconsistency and redundant constraint caused by hierarchical modeling, improves the structured degree and explainability of joint decision;Under the constraint of multidimensional resources such as calculation, memory, storage and bandwidth, dynamic adaptive optimization is realized, node resource over-limit and load imbalance are effectively avoided, and congestion and queuing delay are reduced.
Owner:NANJING UNIV OF POSTS & TELECOMM

An NTT hardware implementation system based on an FPGA platform

The present application belongs to the technical field of post-quantum cryptography engineering, and specifically relates to an NTT hardware implementation system based on an FPGA platform. The system is used for accelerating the polynomial multiplication operation on a ring with the highest complexity and the longest time consumption in a lattice-based cryptography scheme; the hardware circuit supporting multiple parameters is designed for standard NTT, pruned NTT and hybrid NTT, and specifically includes: a compact butterfly operation unit, a reconfigurable modulo reduction unit and a cross-storage type memory access mode, the polynomial multiplication operation is completed through a multi-parallel acceleration structure and a multiplexed butterfly calculation unit; forward NTT, corresponding point coefficient multiplication and inverse NTT operations can be realized. The butterfly operation unit is internally in a pipeline mode, thereby reducing the time delay of the critical path; the modulo reduction hardware unit uses addition and shift to replace the multiplication operation therein, thereby reducing the consumption of DSP resources; the cross-storage type memory access mode can simultaneously support three NTT calculations, and the memory utilization rate reaches 100%.
Owner:FUDAN UNIVERSITY

Method and apparatus for determining memory access latency, system

PendingCN122332134AStart timeEngineering
This disclosure provides a method, apparatus, and system for determining memory access latency. The method includes: determining the start time of multiple memory access commands based on a single timer; for any one of the memory access commands, in response to detecting the completion of the memory access command, determining the latest value of a timer overflow flag corresponding to the memory access command, wherein the latest value of the timer overflow flag indicates the number of times the timer has reached its maximum value from the start time of the memory access command until its completion; and determining the access latency duration corresponding to the memory access command based on the latest value of the timer overflow flag. This method can accurately calculate the access latency duration corresponding to memory access commands while reducing hardware resource consumption, thereby achieving efficient and accurate memory access latency monitoring.
Owner:MOORE THREADS TECH CO LTD

DATA STORAGE DEVICE AND METHOD FOR OPERATING A DATA STORAGE DEVICE

Data storage device comprising an integrated circuit further comprising a control unit (100) and a storage array (400) of charge-based memory cells, wherein: the storage array (400) comprises a first subsection (410) that can be operated as a storage unit and a second subsection (420) that can be operated as a dosimeter; the control unit (100) is capable of providing a reference current (Iref) and To perform memory access operations in order to access memory with reference to the reference stream (Iref); and the control unit (100) is further capable of analyzing a statistical distribution of read streams (Iread) using memory access operations in the second subsection (420), the analysis comprising the following: counting logical reading errors of the Memory access operations and calibration of the reference current (Iref) depending on a number of counted logical read errors, which is also an indicator of total ionization dose (TID).
Owner:AMS INTERNATIONAL AG

Decoupling type multi-host energy efficiency cooperative control system

The invention is suitable for the technical field of cooperative control, and provides a decoupling type multi-host energy efficiency cooperative control system, which comprises a resource abstraction and dynamic reconstruction module used for performing hardware-level decoupling on computing, memory, storage and accelerator resources of heterogeneous computing nodes, and constructing a dynamic resource pool supporting on-demand combination; the multi-dimensional perception and digital twinborn modeling module is used for collecting power consumption data, performance data, thermodynamic data and business load data of the physical infrastructure in real time and constructing a system-level real-time energy efficiency digital twinborn model; the quantum heuristic collaborative decision module adopts an optimization algorithm based on a quantum annealing principle; the self-adaptive closed-loop execution module is used for converting the joint optimization strategy into a control instruction of underlying hardware and adjusting the running state of the system in real time through a feedback mechanism; and a self-optimized intelligent closed loop is formed by means of real-time feedback, so that a traditional static and isolated energy efficiency management normal form is overturned fundamentally.
Owner:JINSHENG ZHIHE (SHANGHAI) TECHNOLOGY CO LTD

CXL switch board and CXL memory allocation system, method and apparatus

The present application relates to the technical field of computers, and provides a CXL switch board and a CXL memory allocation system, method and apparatus. The CXL switch board is provided with a CXL switch unit and a micro-control unit; the CXL switch unit comprises at least one first port and at least one second port, the first port being used for connecting to a CPU, and the second port being used for connecting to a CXL memory board; after the CXL switch unit is powered on, the micro-control unit is used for acquiring the memory capacity of a local memory corresponding to the CPU; the micro-control unit or the CXL switch unit is used for calculating a CXL memory starting address on the basis of the memory capacity; the CXL switch unit is used for writing the CXL memory starting address into a register of the first port corresponding to the CPU, and using the CXL memory starting address as a base address of a CXL memory to allocate the CPU the CXL memory in the CXL memory board.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

A fast line resistance network analysis method for non-volatile memory arrays

The present invention belongs to the technical field of line resistance network analysis, and discloses a fast line resistance network analysis method for a non-volatile memory array, comprising: establishing KCL equations for electrodes in all rows and columns of the line resistance network of the non-volatile memory array; transforming and decomposing the voltage lines and coefficient matrices in the KCL equations, introducing a fixed-point iteration algorithm, and iteratively calculating the voltage distribution of the top and bottom nodes of the array. The originally complex large-scale linear equations problem is converted into two smaller-scale sub-problems. The two sub-problems are solved by mutual iteration, thereby reducing the computational complexity and computational memory. The method of the present invention can predict and evaluate the impact of IR-drop on the acceleration of neural network calculations during the design phase. Advantages include: high accuracy, capable of simulating the impact of IR-drop with an accuracy of more than 95%; low computational complexity, reducing the computational complexity from O((2mn) 2 ) is reduced to O(mn); the memory load is small, minimizing memory usage while maintaining high accuracy.
Owner:HUAZHONG UNIV OF SCI & TECH

A memory dynamic adjustment method and system, electronic equipment and storage medium

The application relates to the computer technical field and discloses a memory dynamic adjustment method and system, an electronic device and a storage medium, which comprise the following steps: collecting resource performance indexes of a virtual machine and a host computer, forming historical standardized time sequence data, performing normalization processing to obtain standardized time sequence data, filtering abnormal data points exceeding a preset mutation threshold, if it is identified that the virtual machine is in a non-steady-state operation, suspending a prediction process, inputting a double-scale time sequence convolution network model, outputting a future short-term memory usage rate sequence through a short-term prediction branch, outputting a memory change trend label through a long-term trend branch, calculating a memory pressure entropy value based on the future short-term memory usage rate sequence, presetting a multistage constraint rule, generating a memory adjustment instruction, correspondingly triggering a hot plug operation according to the memory adjustment instruction, and executing a rollback mechanism when an adjustment exception is detected. The application can meet the actual needs of efficient, intelligent and self-adaptive dynamic adjustment of virtual machine memory resources.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

A matrix multiplication approximate calculation method based on multi-hash voting mechanism

This application relates to a matrix multiplication approximation calculation method based on a multi-hash voting mechanism. The method includes: a source computing node in the inference system obtains the token activation value matrix of the data to be inferred, which is to be sent to the target computing node where the target expert resides; the source and target computing nodes perform hash compression and voting operations on the token activation value matrix and the expert weight matrix respectively based on multiple preset hash functions to generate corresponding sketch sets; during the All-to-All communication process at the MoE layer, the source computing node sends the sketch set of the token activation value matrix to the target computing node; the target computing node performs approximate calculations based on the rows selected from the intersection of the two sketch sets to obtain the inference result. This method improves the inference efficiency of the inference system by introducing multiple independent hash functions for collaborative sampling and voting, thereby reducing computation, memory usage, network load, and GPU idle waiting time.
Owner:GREEN IND INNOVATION RES INST OF ANHUI UNIV

A memory allocation and recycling method for an Android device

The application aims to provide an efficient memory allocation and recovery method for Android devices, and relates to the technical field of memory processing, which comprises the following steps: step 1: obtaining the running data of the Android device, calculating the memory state entropy based on the running data, and combining the priority matrix and the running data to calculate the memory fragmentation risk degree; step 2: calculating the memory allocation dynamic threshold according to the memory fragmentation risk degree, and calculating the adaptive memory compression rate; step 3: using the adaptive memory compression rate to compress the current unused memory data of the Android device until the number of pages of the active process is lower than the set minimum page threshold; step 4: if the memory fragmentation risk degree exceeds the set fragmentation consolidation threshold, the Android device performs memory fragmentation consolidation, and the remaining memory capacity is evenly allocated to all active processes. The application realizes accurate allocation and efficient recovery of memory resources.
Owner:SHENZHEN RUIJIANG TECH CO LTD

Distributed large language model reasoning adaptive load balancing method based on MAPE-K

The invention discloses a distributed large language model reasoning adaptive load balancing method based on MAPE-K. The method comprises the following steps: S1, multi-dimensional performance monitoring; s2, carrying out multi-factor load analysis; s3, self-adaptive routing planning is carried out; and S4, load balancing execution. According to the method, performance problems can be automatically found without manual intervention, an optimization scheme is formulated, adjustment is implemented, and the method has self-adaption and self-optimization capabilities; the node state is comprehensively evaluated from calculation, memory, queue and tail delay, different types of performance bottlenecks can be accurately distinguished, and a reliable basis is provided for formulating a targeted optimization strategy; intelligent matching is carried out according to request features and node capabilities, appropriate workloads are allocated to appropriate nodes, and resource waste and performance degradation caused by one-step static allocation are avoided; historical experience is accumulated through the knowledge base and used for optimizing future decisions, and the load balancing capacity of the system is continuously improved along with the increase of running time.
Owner:HARBIN INST OF TECH

A lightweight server intelligent scheduling method and system for heterogeneous BIM models

This invention discloses an intelligent scheduling method and system for lightweight servers in heterogeneous BIM models. The method includes: constructing a data storage structure and designing engine information tables, server information tables, and task information tables; obtaining unit task memory consumption and server memory configuration information from the engine information tables and server information tables based on the engine identifier code; calculating the total number of server tasks and the number of successful lightweight tasks from the task information tables and calculating the lightweight success rate; and counting the number of tasks in processing from the cache structure; calculating memory usage, memory occupancy rate, and remaining computing power based on unit task memory consumption, server memory configuration, lightweight success rate, and the number of tasks in processing, and determining a weighted remaining computing power based on the lightweight success rate; and implementing server scheduling based on the calculated memory occupancy rate and weighted remaining computing power. Using this invention, model resource consumption and server remaining computing power can be dynamically matched, improving the average utilization rate of cluster memory.
Owner:DONGHUI (ZHEJIANG) TECHNOLOGY CO LTD

A remote sensing image change detection method based on sparse change self-attention mechanism

The present application relates to a kind of remote sensing image change detection method based on sparse change self-attention mechanism, for computing, memory efficient high-precision unified change detection.The present application combines the principle of deep learning, probability graph theory, proposes a unified probability change modeling theory framework, the joint distribution of random variable in change process is conditionally decomposed, according to different prior assumptions, different factors can be decomposed, these factors are the theoretical representation of the architecture of deep change detection model, further parameterize these decomposition factors using the proposed sparse change self-attention module, so as to obtain specific task adaptive, computationally efficient deep change detection model architecture.The present application can solve the problem that existing architecture design lacks theoretical basis and has high computational complexity, and can realize unified processing of various change detection tasks and rapid change detection of large-scale remote sensing image pairs.
Owner:WUHAN UNIV

Memory allocation method, device, apparatus and storage medium

The application relates to data storage and provides a memory allocation method, device, equipment and storage medium. The method calculates a memory usage rate based on a first occupied capacity of a first-level memory unit, calculates a picture loading rate based on a loaded image of the first-level memory unit, calculates a picture removal rate based on a removed image of the first-level memory unit, performs capacity adjustment processing on the first-level memory unit according to the memory usage rate, the picture loading rate and the picture removal rate, obtains an adjusted first-level memory unit, identifies the picture size of a reused picture in a second-level memory unit, creates a set size according to the picture size, performs capacity division on the second-level memory unit based on the reuse frequency and reuse rate of each set size, and obtains a size memory unit of each set size, so that the memory utilization rate can be improved. In addition, the application also relates to a blockchain technology, and the reuse rate can be stored in the blockchain.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Computing and memory chip packaging structure and device and resource pooling system

The invention provides a computing and memory chip packaging structure and device and a resource pooling system. The memory chip packaging structure comprises: a packaging substrate; the optical interconnection module comprises a photoelectric signal conversion module and a first power receiving and generating chip; a first photonic integrated circuit chip disposed on the package substrate; and a memory chip module disposed on the first photonic integrated circuit chip and including a first electrical interconnection interface. The first power receiving and generating chip comprises a second electrical interconnection interface which is electrically connected with the first electrical interconnection interface through the first photon integrated circuit chip. The first power receiving and generating chip converts a first electric signal from the memory chip module into a second electric signal and converts a third electric signal from the photoelectric signal conversion module into a fourth electric signal sent to the memory chip module. The photoelectric signal conversion module converts the second electric signal into a first optical signal output to the outside and converts a second optical signal received from the outside into a third electric signal.
Owner:HANGZHOU GUANGZHIYUAN TECH CO LTD

Dynamic current limiting method and device, electronic equipment and storage medium

The invention provides a dynamic flow limiting method and device, electronic equipment and a storage medium, and the method comprises the steps: collecting memory indexes of a flow calculation node in real time, the memory indexes comprising a heap memory utilization rate, a direct memory utilization rate and a garbage collection frequency; calculating a memory pressure index according to the heap memory utilization rate, the direct memory utilization rate and the garbage collection frequency; determining a rate adjustment coefficient according to the memory pressure index; and dynamically adjusting the data processing rate of the stream computing node according to the rate adjustment coefficient. According to the scheme, the data processing rate is dynamically adjusted according to the real-time memory pressure, elastic allocation and accurate regulation and control of memory resources are achieved, idle waste of the resources is avoided, system stability is guaranteed, and the memory management efficiency of the streaming computing system is fundamentally improved.
Owner:BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD

Reinforcement learning method and system for training push decoupling and asynchronous overlapping of small language model

The invention discloses a training deduction decoupling and asynchronous overlapping reinforcement learning method and system for a small language model, which is characterized in that a training deduction decoupling and iteration internal asynchronous scheduling method is adopted, and the parallel execution of reasoning and training is realized without changing the semantics and convergence properties of the original GRPO algorithm, so that the learning efficiency is improved. According to training deduction decoupling, reasoning Worker and training Worker are deployed on different GPU resource groups, and physical separation of Actor reasoning and training in the spatial dimension is achieved; the asynchronous scheduling in iteration realizes time overlapping of reasoning and training; the system is composed of a reasoning execution module, a Reference model calculation module, a reward generation module, a training module and the like and a parameter buffer area located in a CPU, and all the modules are unified and coordinated through a scheduling layer. Compared with the prior art, the method has the advantages that a localized, high-throughput and stable reinforcement learning training process is realized under the limited hardware condition by a user, and the system bottlenecks of communication amplification, resource coupling, computing-memory load imbalance and the like generally existing in a training system after SLM reinforcement learning are effectively solved.
Owner:EAST CHINA NORMAL UNIV