Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

214 results about "Request queue" patented technology

Large model batch reasoning and data flow optimization system oriented to MOE architecture

The invention relates to the technical field of project management, in particular to a large-model batch reasoning and data flow optimization system oriented to an MOE architecture. According to the method, a collaborative architecture of the request access module, the environment sensing module, the expert routing engine, the resource scheduling module and the dynamic optimization control module is set, the text length and the subject type are extracted by using the request access module, a basis is provided for accurate routing, and the GPU video memory, the I / O bandwidth and the request queue depth are acquired in real time through the environment sensing module, so that the real-time routing is realized. The system load is comprehensively monitored, meanwhile, an expert sub-network is activated through an expert routing engine according to request features, invalid calculation is avoided, weight loading and resource allocation are managed through a resource scheduling module, the I / O bottleneck is reduced, and finally an optimization strategy is intelligently triggered through a dynamic optimization control module based on routing conflict factors. The problems of large reasoning delay fluctuation and unbalanced resource utilization rate mentioned in the background technology are solved, and stable low-delay response and resource collaborative optimization in a high-concurrency scene is realized.
Owner:VIRTAI TECH BEIJING CO LTD

Video stream concurrent access front-end dynamic scheduling control system and method based on digital twinborn scene

The invention discloses a video stream concurrent access front-end dynamic scheduling control system and method based on a digital twinborn scene, and relates to the technical field of computer software and network communication, the system comprises three core modules: a self-adaptive request queue module manages UE rendering resources based on a session token pool, distributes tokens through preemptive scheduling, and sends the tokens to a server; in combination with a GPU frame rate feedback dynamic adjustment strategy, queuing through a Promise asynchronous queue when the request fails; the front-end fusing and backoff retry module monitors the connection failure rate, when the connection failure rate reaches a threshold value, fusing is triggered, retry is delayed by adopting an exponential backoff strategy, and a semi-open state tends to recover; and the local cache collaboration module intercepts the request through Service Worker, caches the latest video clip, pushes the local content when the request is delayed, and ensures the freshness through ETag verification. The method and the device are used for efficiently managing loading and interaction of multiple users on real-time rendering pictures at a browser end.
Owner:浪潮智慧城市科技有限公司

Chip verification method and device, equipment, medium and program product

The invention provides a chip verification method and device, equipment, a medium and a program product, and the method comprises the steps: receiving a first operation instruction of a to-be-verified module and a second operation instruction of a reference model in parallel, the reference model being a function model used for simulating the expected behavior of the to-be-verified module; searching a first instruction record matched with the first operation instruction from a request queue of the reference model, and searching a second instruction record matched with the second operation instruction from a request queue of the to-be-verified module; and determining a verification result of the to-be-verified module according to a comparison result of the first operation instruction and the first instruction record and / or a comparison result of the second operation instruction and the second instruction record. According to the invention, a structured and reusable universal chip verification method is realized, so that a tedious customized scoreboard design is replaced, and the efficiency and accuracy of verification work are greatly improved.
Owner:SHANGHAI BIREN TECH CO LTD

Model reasoning scheduling method and system, electronic equipment and storage medium

The invention provides a model reasoning scheduling method and system, an electronic device and a storage medium, the system comprises a plurality of pre-filling server nodes and a plurality of decoding server nodes, and the method comprises the following steps: responding to received user request information, and combining a pre-filling scheduling model based on deep reinforcement learning, obtaining a target pre-filling server node selected from the plurality of pre-filling server nodes and a first priority; sending a pre-filling request to a request queue corresponding to the first priority in the target pre-filling server node, and enabling the target pre-filling server node to process the pre-filling request; after the pre-filling request is processed, obtaining a target decoding server node selected from a plurality of decoding server nodes and a second priority by combining a decoding scheduling model based on deep reinforcement learning; sending a decoding request to a request queue corresponding to the second priority in the target decoding server node, and enabling the target decoding server node to process the decoding request; and reasoning efficiency is improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Fast load of dashboards using data caching

Systems and methods disclosed herein for improving data retrieval for dashboards using data caching. The system may implement a cache memory to store results for executed queries and / or anticipated queries, allowing for data to be retrieved asynchronously, and may implement a request queue to distribute the queries among processing nodes.
Owner:SERVICENOW INC

Timeout processing for aggregation messaging in an integration environment

A method includes: receiving an initial request message in a request queue of the messaging system; receiving an aggregation reply message in a reply queue of the messaging system, wherein the aggregation reply message is received from an integration system that processes the initial request message, and wherein the aggregation reply message includes an aggregation identifier associated with the initial request message; in response to receiving the aggregation reply message, starting a timeout process; monitoring the timeout process; holding one or more messages that are in the reply queue and that include the aggregation identifier until the timeout process expires or an expected number of responses has been received; and making available the one or more messages that are in the reply queue and that include the aggregation identifier based on the timeout process expiring or the expected number of responses having been received.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Data replay method, data processing unit, network interface card, device, and storage medium

The present application relates to a data replay method, a data processing unit, a network interface card, a device, and a storage medium. The method comprises: when a storage stack is restarted, acquiring a replay-pending request queue from a persistent memory, the replay-pending request queue being a queue obtained by mapping a transmission buffer area into the persistent memory, and the transmission buffer area being used for receiving a request from a host side; on the basis of the replay-pending request queue, acquiring a replay-pending request from the persistent memory; and resending the replay-pending request to a target acceleration component directed to a storage device, so as to instruct the storage device to store target data which resides in the persistent memory and corresponds to the replay-pending request. In the method, a request queue, a request, and data corresponding to the request are stored in the persistent memory in advance, and if the storage stack is restarted, the replay-pending request is recovered from the persistent memory, eliminating the need for DMA and encryption / decryption techniques, and shortening an interruption duration of a storage data path, thereby improving data storage efficiency on the host side.
Owner:NANJING JAGUAR MICROSYSTEMS CO LTD +1

Method and device for releasing prefetch request

The invention provides a method and a device for releasing prefetch requests. The method comprises the following steps: acquiring an earliest prefetch request from a prefetch request queue as a prefetch request to be processed; if it is judged that the to-be-processed prefetch request is hit in a virtual-real address translation cache and hit in a target cache, or the to-be-processed prefetch request is hit in the virtual-real address translation cache and hit in an address missing state tracking register; and if so, releasing the to-be-processed prefetch request and the related prefetch request from the prefetch request queue. The device is used for executing the method. According to the prefetching request releasing method and device provided by the embodiment of the invention, the extra power consumption of data prefetching is reduced.
Owner:HYGON INFORMATION TECH CO LTD

Network-on-chip system based on cache consistency

The invention discloses an on-chip network system based on cache consistency, belongs to the technical field of computer network communication, and aims to solve the technical problems of global consistency flow overload, on-chip network scheduling delay, directory node load imbalance and too high core resource occupation of a traditional router architecture. Comprising a multi-core cluster, distributed directory nodes, a multi-channel network-on-chip and a hardware acceleration module, cores are connected through a bus, and local cache consistency is achieved based on an MSI protocol; a directory node divides a physical address space into a plurality of fragments based on an address hash fragmentation mechanism, each directory node maintains a sparse directory table, and when the length of a request queue of any directory node exceeds a set threshold value, a dynamic load balancing mechanism is triggered; the multi-channel network-on-chip provides a control channel, a data channel and a configuration channel, and each channel performs message scheduling based on a dynamic priority mechanism; and the hardware acceleration module is used for processing three types of consistent transactions.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Asynchronous communication system, method and equipment based on Netty and medium

The invention provides an asynchronous communication system, method and device based on Netty and a medium, and belongs to the technical field of communication. The system comprises: an asynchronous network communication module, which is used for establishing asynchronous non-blocking network communication connection between a client and a server by using a Netty asynchronous non-blocking I / O model; the memory management module is used for carrying out allocation and recovery management on the memory in the communication process through a memory pool and object pool technology; the protocol adaptation module is used for adapting to different protocols through a self-defined codec and carrying out encoding and decoding processing on data of different protocols; and the back pressure control module is used for introducing a machine learning model, predicting the processing capability of the server and the request trend of the client, dynamically adjusting a semaphore threshold value or the length of a request queue, and realizing adaptive flow control. According to the invention, the communication efficiency is improved, the memory management is optimized, the protocol adaptation is flexible, and the adaptive flow control is realized, so that the performance and service adaptability of a communication system are enhanced.
Owner:INSPUR ARTIFICIAL INTELLIGENCE RES INST CO LTD SHANDONG CHINA

Cache, cache management method and electronic device

Cache, cache management method, and electronic device. The cache includes: multiple cache lines, a first read request queue and a second read request queue; a first read request queue is configured for storing and sending a first read request to a memory controller; the first read request is configured for requesting data from memory and storing the data in the memory controller; the number of the first read requests stored in the first read request queue is greater than that of the multiple cache lines; the second read request queue is configured for storing and sending a second read request to memory controller; and the second read request corresponds one-to-one with the first read request, and the second read request is used to request data corresponding to the first read request from the memory controller when a cache line corresponding to the first read request is idle
Owner:VERISILICON MICROELECTRONICS (SHANGHAI) CO LTD

Method and system for balancing dynamic wear of solid state disk based on data popularity

The invention relates to a solid state disk dynamic wear balancing method and system based on data popularity, and the method comprises the steps: tracking the access characteristics of each data page in a solid state disk through a page-level counter, and enabling the access characteristics to comprise the write operation frequency, the modification period and the survival time; the data pages are dynamically divided into high-heat data, medium-heat data or cold data based on a preset threshold value, and corresponding heat layering labels are marked; collecting accumulated writing times and I / O request queue length of each channel of the solid state disk in real time, and calculating wear imbalance degree and load imbalance degree; performing data exchange according to a preset exchange mode priority based on the popularity hierarchical label, the wear imbalance degree and the load imbalance degree; and updating the logic mapping table and the channel state parameters, carrying out statistics on the balance income and the system overhead, and dynamically correcting the tolerance coefficient based on the current load type according to the income-overhead ratio.
Owner:SHENZHEN PENGRUNNING TECH CO LTD

An e-book borrowing management system based on OPAC two-way connection

This invention provides an e-book lending management system based on OPAC bidirectional linkage, belonging to the field of library information technology. It includes: a hardware acceleration layer that constructs a dynamic resource distribution matrix and calculates the matrix norm in real time using an FPGA coprocessor, combined with an ASIC anti-collision controller to implement request queue scanning every 100ms and a redundant copy generation algorithm; a system service layer that deploys a dynamic priority scheduling engine, achieving intelligent resource allocation based on cross-campus collaboration coefficients and edge-cloud collaboration strategies, while recording borrowing operations and performing automated copyright revenue sharing audits through blockchain notarization services; and a user interaction layer that provides a bidirectional association search interface and a priority borrowing channel based on credit scoring. The system improves resource scheduling efficiency in high-concurrency scenarios through hardware-accelerated matrix norm calculation, an elastic copy allocation mechanism, and an anti-collision algorithm embedded in ASICs. Combined with geolocation scoring and blockchain technology, it achieves cross-campus resource collaboration and accurate allocation of copyright revenue.
Owner:SHANDONG CHINESE EDUCATION IND DEVELOPMENT CO LTD

Kubernetes-based large model reasoning optimization method, system and equipment

The invention provides a Kubernetes-based large model reasoning optimization method, system and equipment, and the method comprises the steps: constructing a resource scheduling assembly for a large model, and monitoring a real-time resource state at a current moment; constructing a batch processing agent component for the large model, performing feature analysis on the received reasoning request queue, and determining request features corresponding to the reasoning requests in the reasoning request queue; the real-time resource state is evaluated according to the request features, and the batch processing size is adjusted based on the evaluation result to generate the optimal batch processing length; batching the reasoning requests in the reasoning request queue according to the optimal batch processing length, and packaging each batch of reasoning requests into batch data; the resource scheduling component is utilized to distribute batch data to Pod in the K8s cluster for reasoning, a large model reasoning result is obtained, the quantity of the reasoning Pod is automatically adjusted according to the flow and the resource load required by reasoning, a GPU with poor performance is prevented from becoming a system bottleneck, and the responsiveness and the stability of service are improved.
Owner:广域铭岛数字科技有限公司 +1

Designated driving service request processing method and device, computer equipment and storage medium

The invention relates to a designated driver service request processing method and device, computer equipment and a storage medium. The method comprises the following steps: in response to a received current request sent by a designated driver service request queue, determining a target request type according to the current request; the designated driving service request queue comprises a new request sub-queue, a detection retry request sub-queue and a step retry request sub-queue; the target request type is the request type of the current request; the request type comprises a new request, a detection retry request or a step retry request; in response to the fact that the target request type is a new request or a stepped retry request, obtaining a historical request processing success rate; the historical request processing success rate is the processing success rate of each historical request in the previous preset first time window; and in response to the situation that the historical request processing success rate is smaller than a first success rate threshold, intercepting the current request, updating the target request type as a detection retry request, and transferring and storing the current request to a detection retry request subqueue. By adopting the method, the efficiency can be improved.
Owner:BEIJING LONGJU YIXING TECH CO LTD

Management method and device of dynamic multi-thread access memory

The invention discloses a dynamic multi-thread memory access management method, and relates to the technical field of computers, in particular to a dynamic multi-thread memory access management method and device.The method comprises the following steps that a plurality of memory access requests outside a processor and / or inside the processor are obtained, and a first request queue is formed; wherein the memory access request comprises a thread identifier, an access type and a target memory address; dynamically sequencing the memory access requests in the request queue according to a preset priority rule, and calculating the memory access request with the highest priority; wherein the priority rule comprises at least one of a thread priority rule, a request source priority rule and a memory address priority rule; obtaining and executing the memory access request with the highest priority; the memory resource allocation efficiency can be effectively improved, the system adaptability and flexibility are enhanced, the hardware implementation complexity is simplified, and the response delay is reduced.
Owner:SUZHOU HONGXIN INTEGRATED CIRCUIT CO LTD

System, method and device for improving reasoning efficiency of large language model and medium

The invention relates to the technical field of artificial intelligence, and relates to a system, method and device for improving the reasoning efficiency of a large language model, and a medium. The system comprises a request scheduling module, an assembly line control module, a prompt processing work pool, a token generation work pool and a shared memory management module. The method comprises the following steps: initializing a system; external reasoning requests are received and stored in a to-be-processed request queue; parallel execution of the prompt processing task and the token generation task is realized through asynchronous task driving; dynamically adjusting task scheduling priorities based on real-time hardware indexes; and the token generation working unit returns a generated text to the client after completing the generation of one request, and releases a corresponding KV Cache space through the shared memory management module. According to the method, the problem of hardware idleness caused by resource demand mismatching in the reasoning process is solved, so that the comprehensive utilization rate of hardware and the total throughput of the system are remarkably improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Task assignment system and computing chip

The invention relates to the technical field of chip micro-architecture, and provides a task assignment system and a computing chip wherein the system comprises a central controller and a plurality of computing units; the central controller comprises a queue manager, and the queue manager is used for managing task queues in the central controller; identifiers of a plurality of to-be-processed task blocks are stored in the task queue; the computing unit is used for sending a task request to the queue manager under the condition that any task block is completed; the queue manager is used for responding to the task request, acquiring an identifier of a to-be-processed target task block from the managed task queue, and returning the identifier of the target task block to the computing unit corresponding to the task request; and the computing unit is also used for receiving the identifier of the target task block and performing task execution based on the identifier of the target task block, so that full utilization of chip computing resources is realized, a'trailing effect 'that other computing units are forced to be idle and wait due to long time consumption of individual tasks is avoided, and the overall operation performance is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Server-free function deployment method for space-air-ground integrated network

The embodiment of the invention provides a server-free function deployment method for an air-space-ground integrated network, which is applied to the technical field of edge computing, and comprises the following steps: constructing a three-layer network model of the air-space-ground integrated network, and determining information of nodes and links thereof; establishing a function dependency graph of a server-free application and generating a task request queue; modeling a server-free function placement problem in the space-air-ground integrated network as a weighted sum of minimum request execution cost and end-to-end delay, and designing a constraint condition and a target function; establishing a reward function required by reinforcement learning according to the target function; and obtaining an optimal server-free function placement strategy through a deep reinforcement learning algorithm based on graph perception embedding. According to the method, the end-to-end execution cost and time delay are reduced, adaptive optimization can be carried out for the dynamic nature of the network and the heterogeneity of node resources, and it is ensured that a delay sensitive task obtains stable and high-quality services in a complex scene.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Mass inference and data flow optimization system for MOE architecture-oriented large model

The present application relates to the technical field of project management, in particular to a large model batch inference and data flow optimization system for MOE architecture. The present application sets a collaborative architecture of a request access module, an environment perception module, an expert routing engine, a resource scheduling module and a dynamic optimization control module, uses the request access module to extract the text length and the subject type to provide the basis for accurate routing, collects the GPU display memory, the I / O bandwidth and the request queue depth in real time through the environment perception module, comprehensively monitors the system load, activates the expert sub-network according to the request characteristics through the expert routing engine to avoid invalid calculation, manages the weight loading and resource allocation through the resource scheduling module to reduce the I / O bottleneck, and finally intelligently triggers the optimization strategy based on the routing conflict factor through the dynamic optimization control module to solve the problems of large inference delay fluctuation and unbalanced resource utilization mentioned in the background technology, and realize stable low-delay response and resource collaborative optimization in a high-concurrency scenario.
Owner:VIRTAI TECH BEIJING CO LTD

System for carrying out self-adaptive current limiting on token number and QPS called by large model

PendingCN121479508AData setAdaptive management
The invention relates to the technical field of large model service management, and discloses a system for carrying out self-adaptive current limiting on a token number and a QPS called by a large model. A dynamic monitoring module of the system is responsible for collecting a request flow data set comprising a token consumption sequence, a request frequency sequence and a response delay sequence; a traffic feature extraction module performs multi-dimensional feature analysis on the data set to generate a traffic feature matrix containing token consumption rate, request frequency fluctuation coefficient and delay sensitivity index; the adaptive current-limiting decision module dynamically matches a load balancing strategy based on the matrix, and generates a current-limiting control parameter set containing a token quota threshold and a QPS upper limit threshold; the real-time regulation and control module dynamically adjusts the request queue according to the request queue to generate a new request scheduling sequence; and the feedback optimization module monitors the execution effect and iteratively updates the load balancing strategy. According to the invention, through real-time perception and dynamic adjustment, refined and adaptive management of large model service resources is realized.
Owner:HANGZHOU JIHEXIN TECHNOLOGY CO LTD

Managing lock state data of datasets dispersedly stored across a database system

A method includes obtaining, by a lock state management module of a leader computing node of a plurality of computing nodes, a first lock request from a first client. The first lock request corresponds to a first data access request of a first dataset. The first data set is dispersedly stored across a set of computing nodes of the plurality of computing nodes. The method further includes identifying a first pending lock request queue of a plurality of pending lock request queues of shared lock state data based on first lock metadata of the first lock request. The shared lock state data is shared by the leader computing node with other computing nodes of the plurality of computing nodes. The method further includes adding the first lock request to the first pending lock request queue in accordance with a prioritization ordering, determining to grant the first lock request based on a set of grant conditions to produce a granted first lock request of a first set of granted lock requests of the shared lock state data, and removing the first granted lock request from the first set of granted lock requests upon a removal indication.
Owner:OCIENT HOLDINGS LLC

Request processing method and device in distributed system, equipment, medium and product

The invention discloses a request processing method and device in a distributed system, equipment, a medium and a product, and relates to the field of distributed systems. The method comprises the following steps: sending a target IO request; inserting the target IO request into a request queue based on the priority information of the target IO request; processing at least one IO request in the request queue in sequence; wherein the at least one IO request is ranked in the request queue according to the priority from high to low. The IO requests which can cause the real-time influence on the training tasks are processed preferentially, and the IO requests which do not influence the training real-time performance are processed in parallel while the training tasks are ensured to be continuously carried out. The waiting time for the computing nodes to carry out the training tasks is shortened, and the overall resource utilization rate and the task execution efficiency of the distributed system are improved.
Owner:MOORE THREADS TECH CO LTD

Selectively bypassing an external cache of a storage system based on saturation monitoring relating to the external cache

Systems and methods for selectively bypassing an external cache (EC) of a storage system are provided. In one example, when the EC backing storage device is saturated, reads bypass the EC and are completed via a redundant array of independent disks (RAID) subsystem of the storage system. One or more performance metrics for the EC backing storage device may be monitored to predict one or more saturation thresholds (e.g., in terms of latency and / or throughput). Based on this monitoring, tuning may be performed to drive utilization of the EC into a “knee region” of a performance (or response) curve of the EC backing storage device. For example, the depth of one or more request queues at the front-end of an EC lookup may be manipulated based on a current measure of saturation relating to the EC backing storage device to limit the number of in-flight reads pending for the EC.
Owner:NETAPP INC

Task dispatch system and computing chip

The application relates to the technical field of chip micro-architecture, and provides a task dispatching system and a computing chip, wherein the system comprises a central controller and a plurality of computing units; the central controller comprises a queue manager, the queue manager is used for managing a task queue in the central controller; the task queue stores the identities of a plurality of to-be-processed task blocks; the computing unit is used for sending a task request to the queue manager when any task block is completed; the queue manager is used for responding to the task request, obtaining the identity of a target task block to be processed from the managed task queue, and returning the identity of the target task block to the computing unit corresponding to the task request; and the computing unit is further used for receiving the identity of the target task block and performing a task based on the identity of the target task block, so that the computing resources of the chip are fully utilized, the "tail effect" that other computing units are forced to idle and wait due to long time consumption of individual tasks is avoided, and the overall operation performance is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Unified protocol based file system kernel service and high concurrency scheduling method

PendingCN122309467AFile systemTimestamping
This invention relates to the field of computer file system technology, and in particular to a file system kernel service and high-concurrency scheduling method based on a unified protocol. The method includes the following steps: Step 1: Starting a receiving thread and a file operation thread on the server side, the file operation thread including a lightweight thread and a heavyweight thread; Step 2: Listening to the network port through the receiving thread to receive JSON-formatted requests sent by clients, the requests containing timestamps and operation types; operation types include copy, move, create directory, delete, list, and restore; Step 3: Storing the received requests in a receiving queue; Step 4: Retrieving requests from the receiving queue and distributing them according to the operation type; Step 5: Executing the operation; This invention exposes file operation capabilities through independent processes and a unified JSON protocol. The six types of operation semantics, as well as the semantics of abort and recoverable deletion, are all uniformly defined at the protocol layer. The dual-type threads and request queue improve throughput and stability under multiple clients and large-scale tasks.
Owner:ALL THINGS SEARCH (GUANGZHOU) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

System and method for network configuration

A system of a server is associated with channels, a plurality of client devices subscribed to the channels, and the server includes: one or more processors; and memory storing instructions that, when executed by the one or more processors, cause the system to perform operations comprising: receiving a system configuration change request from a client device, wherein the system configuration change request comprises a first system configuration file and channel information of a first channel; obtaining, based on the channel information, a second system configuration file that is currently deployed on client devices subscribed to the first channel; displaying, on a graphic user interface, a file comparison result between the first system configuration file and the second system configuration file; and in response to the file comparison result being verified, storing the first system configuration file in a request queue for the client devices to poll and deploy.
Owner:PALANTIR TECHNOLOGIES INC

Data processing method and apparatus

Provided in the embodiments of the present disclosure are a data processing method and apparatus. The method comprises: acquiring a data processing request queue of a cloud storage node, wherein the data processing request queue comprises a specified data processing request and a non-specified data processing request, and the specified data processing request is a data processing request which requires encryption or decryption processing to be performed on carried data; performing cyclic processing on the data processing requests in the data processing request queue, which comprises: performing encryption or decryption processing on data carried in the specified data processing request, giving a response by using data obtained after the encryption or decryption processing, and giving a response by using data carried in the non-specified data processing request; and when it is detected that the accumulated data volume involved in the encryption or decryption processing during the current cyclic processing is greater than a preset data volume, stopping processing the specified data processing request during the current cyclic processing.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Cache device, operation method, electronic equipment and artificial intelligence processor

The embodiment of the invention provides a cache device, an operation method, electronic equipment and an artificial intelligence processor. The cache device comprises a receiving module, a scheduling module and an execution module. The receiving module is configured to receive a plurality of requests; the scheduling module comprises a plurality of waiting queues, the scheduling module is configured to store the plurality of requests into the plurality of waiting queues respectively and correspondingly based on the request types of the plurality of requests, and the plurality of waiting queues comprise a read-write request queue and a calculation request queue; the execution module comprises a plurality of execution assembly lines, the execution assembly lines are respectively coupled with corresponding waiting queues in the plurality of waiting queues, the execution assembly lines further comprise a read-write assembly line and a calculation assembly line, the read-write assembly line is configured to execute read-write requests transmitted from the read-write request queue, and the calculation assembly line is configured to calculate the read-write requests transmitted from the read-write request queue; the compute pipeline is configured to execute compute requests transmitted from the compute request queue. The cache device can improve the calculation efficiency of the universal graphics processor.
Owner:SHANGHAI BIREN TECH CO LTD

Data read / write method, system, and apparatus, computing device, and storage medium

PCT designated stageWO2026144711A1Computer hardwareEngineering
Embodiments of the present disclosure provide a data read / write method, system, and apparatus, a computing device, and a storage medium. The data read / write method comprises: receiving, by means of an asynchronous I / O layer, a data read / write request sent from a user space, and storing the data read / write request in a request queue; when the asynchronous I / O layer determines that the data read / write request is of a pass-through type, sending the data read / write request to a driver layer; and parsing the data read / write request by means of the driver layer, and sending request information of the data read / write request to a target processing end device for read / write processing. The data read / write request sent from the user space is received by means of the asynchronous I / O layer, and is sent to the driver layer after the data read / write request is determined to be of the pass-through type, thereby reducing context switching and data copying overhead, and improving the processing efficiency of I / O operations. The request is parsed at the driver layer and sent to the processing end device for read / write processing, thereby optimizing the scheduling of I / O tasks, reducing disk addressing time and I / O latency, and improving overall performance and system stability.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD +1