Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2059 results about "Shared memory" patented technology

In computer science, shared memory is memory that may be simultaneously accessed by multiple programs with an intent to provide communication among them or avoid redundant copies. Shared memory is an efficient means of passing data between programs. Depending on context, programs may run on a single processor or on multiple separate processors.

Efficient remote pointer sharing for enhanced access to key-value stores

A method to share remote DMA (RDMA) pointers to a key-value store among a plurality of clients. The method allocates a shared memory and accesses the key-value store with a key from a client and receives an information from the key-value store. The method further generates a RDMA pointer from the information, maps the key to a location in the shared memory, and generates a RDMA pointer record at the location. The method further stores the RDMA pointer and the key in the RDMA pointer record and shares the RDMA pointer record among the plurality of clients.
Owner:IBM CORP

Real-time interactive digital human system supporting high concurrency and implementation method thereof

The invention discloses a real-time interactive digital human system supporting high concurrency and an implementation method thereof, and relates to the technical field of digital human interaction.The system comprises a model instance pool module used for loading the weight of a deep learning model to a shared memory area through the memory mapping technology; the multi-thread scheduling module is used for scheduling user requests by adopting a lock-free queue and a dynamic priority algorithm; the asynchronous pipeline processing module is composed of decoupled micro-services, and the modules are connected in series through asynchronous message queues; the client SDK is used for dynamically switching a rendering mode according to terminal hardware performance and network conditions; the elastic capacity expansion and contraction module is used for monitoring resource loads in real time and automatically adjusting the number of service instances; and the audio and video synchronization calibration module is used for ensuring that the synchronization error of the audio and video frames is lower than a preset threshold value through a timestamp alignment algorithm. According to the scheme, core challenges in a high-concurrency digital human interaction scene can be systematically solved, and innovative support is provided for large-scale real-time application.
Owner:LIANGSHENG DIGITAL ARTIFICIAL INTELLIGENCE (SHENZHEN) CO LTD

Method and apparatus for efficient access to multidimensional data structures and / or other large data blocks

A parallel processing unit comprises a plurality of processors each being coupled to a memory access hardware circuitry. Each memory access hardware circuitry is configured to receive, from the coupled processor, a memory access request specifying a coordinate of a multidimensional data structure, wherein the memory access hardware circuit is one of a plurality of memory access circuitry each coupled to a respective one of the processors; and, in response to the memory access request, translate the coordinate of the multidimensional data structure into plural memory addresses for the multidimensional data structure and using the plural memory addresses, asynchronously transfer at least a portion of the multidimensional data structure for processing by at least the coupled processor. The memory locations may be in the shared memory of the coupled processor and / or an external memory.
Owner:NVIDIA CORP

Cross-container application fusion switching method of swan gap system

The invention discloses a cross-container application fusion switching method of a swan monk system, which comprises the following steps: a system service layer deploys a container application management service, an application framework layer realizes a proxy application manager, when a container application is started, a container side allocates a shared memory and configures authority, the container manager collects metadata to initiate registration, and the application framework layer realizes a proxy application manager; the container application management service converts a memory handle into a texture handle, allocates a unique identifier and triggers a proxy application manager to generate a proxy application, and the proxy application initializes a Vulkan rendering environment; the container side renders an application interface to a shared memory, the proxy application imports texture and constructs a lightweight Vulkan rendering pipeline, and scaling sampling is carried out to generate a thumbnail of the container application; and when the container application exits, the shared memory is released, the container application management service cleans the shared memory reference, triggers and destroys the proxy application, and recycles the texture resources through the reference counter, so that seamless fusion, low-delay switching and efficient resource utilization of the container application in a native application thumbnail form are realized.
Owner:北京麟卓信息科技有限公司

Android container rendering optimization method based on cross-domain hard real-time Fen synchronization

The invention discloses an android container rendering optimization method based on cross-domain hard real-time Fen synchronization, which comprises the following steps: by taking a swan-mong system as a host and an android system as a container, creating a hash table, a synchronous thread and a shared memory corresponding to a GPU core when the host is started, acquiring VSync cycle registration callback, transmitting shared memory FD to the container, and finishing shared memory mapping and alignment by the container. Registering a GPU queue to complete callback; when the Android application is started, a container obtains a queue and a physical address of a rendering buffer area, creates a Fen and binds the Fen to the queue, after the queue is submitted, metadata is written into a shared memory to inform a host, after the host receives the metadata, nodes are created and stored in a hash table, and an overtime timer is registered; after the GPU completes the command queue, the container calls back an update state and a verification value to notify the host, and after the host is verified to be valid, the corresponding hash table is updated, the timer is reset, and asynchronous screen loading is triggered; and the host executes buffer area synthesis and submission of the display equipment to finish on-screen, so that the stability of the rendering frame rate is improved, and the reliability of cross-domain synchronization is ensured.
Owner:北京麟卓信息科技有限公司

SDR-oriented heterogeneous task scheduling and transmission system and method

The invention discloses an SDR-oriented heterogeneous task scheduling and transmission system and method.The system comprises a heterogeneous platform composed of an ARM and an FPGA, a dynamic scheduling decision engine is arranged in the ARM and receives a task feature vector and a platform state parameter set as input, and the dynamic scheduling decision engine evaluates income and cost of task allocation to the ARM or the FPGA for execution to make a task allocation decision; the data generated in the task execution process or the execution completion result are interacted and synchronized between the ARM and the FPGA through the heterogeneous shared memory area HSM. The method realizes reduction of signal processing pipeline delay, improvement of data throughput and optimization of system power consumption, and is suitable for communication, electronic countermeasure and other scenes needing high-performance real-time signal processing.
Owner:XIAN LIUWEI PLATINUM ELECTRONICS CO LTD

Sublimation hardware-based video multi-target intelligent detection method and system

The invention discloses a video multi-target intelligent detection method and system based on mercuric chloride hardware. A hardware decoding module decodes an input video stream in real time, generates video frames and stores the video frames in a shared memory queue. And inputting the video frame into the YOLOv5 model converted by the mercuric chloride OMG tool, pre-loading a plurality of model instances into a memory by using an ACL interface of the mercuric chloride NPU, calling different model instances through a polling scheduling mechanism, and outputting structured data comprising a multi-target detection frame, a category label and confidence. And carrying out non-maximum suppression processing on the reasoning result, and judging whether the target is in an alarm monitoring area or not by adopting a central point detection method. According to the invention, real-time multi-target intelligent detection of the input video stream can be realized. And the decoded video frames are stored in a shared memory queue, so that efficient data transmission and processing are realized. The multi-model parallel reasoning improves the detection precision, and is suitable for the application scene of real-time video multi-target intelligent detection.
Owner:CHENGDU SIWEI INTERACTIVE TECH CO LTD

Multi-processor core communication method and device, electronic equipment and storage medium

The invention relates to the field of satellite navigation anti-interference and the technical field of computers, and discloses a multi-processor core communication method and device, electronic equipment and a storage medium, and the method comprises the steps that a second processor core writes target information into a shared memory, and obtains a first idle channel number; the second processor obtains a first interrupt signal based on an interrupt trigger register corresponding to the first idle channel number; the first processor core reads target information in a shared memory according to the first interrupt signal; the first processor core processes the target information to obtain a processing result, and writes the processing result into the shared memory to obtain a second idle channel number; the first processor core obtains a second interrupt signal based on an interrupt trigger register corresponding to the second idle channel number; and the second processor core reads the processing result in the shared memory according to the second interrupt signal, and writes the processing result into the cache region. The multi-processor core communication method is suitable for different operating systems and is high in transportability.
Owner:CHANGSHA HAIGE BEIDOU INFORMATION TECH CO LTD

Container application on-screen method based on texture full-process optimization

The invention discloses a container application on-screen method based on texture full-process optimization, which comprises the following steps: taking a swan-gap system as a host and an Android system as a container, distributing a plurality of first buffers by the host, screening out a first texture format by the container, writing original textures of the format generated by rendering into frame buffers, identifying a dirty area of the rendered textures, and displaying the original textures in the frame buffers; packaging the texture data of the region, storing the packaged texture data into a shared memory, transmitting a file descriptor FD to a host, binding an idle first buffer setting priority, a state, a generation timestamp and a texture data pointer at the same time, and obtaining the texture data by the host through the FD; the host establishes a fixed thread task pool, executes multiple tasks in parallel to obtain on-screen texture data, and then updates the state of the first buffer and a texture data pointer; and finally, the host queries the first buffer with the processed state, reads the texture data after sorting according to the priorities and the generation timestamps to finish on-screen, and resets the state of the first buffer to be idle, so that the on-screen delay is effectively reduced, and the actual frame rate of the Android application is improved.
Owner:北京麟卓信息科技有限公司

Information processing method and device, electronic equipment and storage medium

The invention discloses an information processing method and device, electronic equipment and a storage medium, and is applied to the field of information processing. The information processing method comprises the steps that a first utilization rate is determined in combination with a data flow between a memory and a shared memory when a processor executes a general matrix multiplication operator, and the first utilization rate is used for indicating performance evaluation of the general matrix multiplication operator at a processor level; determining a second utilization rate in combination with a data stream between a shared memory and a register file in a single computing unit when the processor executes the general matrix multiplication operator, the second utilization rate being used for indicating performance evaluation of the general matrix multiplication operator at the computing unit level; and based on the smaller one of the first utilization rate and the second utilization rate, determining the theoretical performance of the processor when executing the universal matrix multiplication operator. At present, only chip-level performance evaluation is considered in performance evaluation of a GEMM operator, and double-level-dimension theoretical performance evaluation is provided, so that the evaluation result is more accurate.
Owner:SHANGHAI BIREN TECH CO LTD

Method for sharing file system by multiple hosts, product, equipment and storage medium

The invention discloses a method, a product and equipment for sharing a file system by multiple hosts and a storage medium, and relates to the technical field of computers, which comprises the following steps: configuring a shared memory device as a character device, directly mapping to a user mode virtual address space, bypassing a kernel page cache, and directly reading and writing a physical address of the shared memory. A shared memory is divided into a super block area, a metadata area, a log area and a data area during formatting, all hosts perform unified operation through the shared metadata area, stand-alone cache interference is avoided, a hidden metadata file is created to map a non-data area during mounting, a log item is scanned to reconstruct a directory file structure, and the mounting efficiency is improved. All hosts generate a directory file structure completely consistent with the shared memory through log replay, the problem of data consistency caused by a cache strategy in a traditional multi-host file system is solved, and the effects that multiple hosts share a unified memory view, operation is real-time and synchronous, and data access delay is reduced are achieved.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Cross-host memory sharing method, system, equipment and medium

The invention discloses a cross-host memory sharing method, system and device and a medium, the memory sharing system comprises a master node and a plurality of slave nodes, the memory sharing method is applied to the master node, and the method comprises the steps that a memory application request of a target slave node is received; applying for a target memory in a target memory sharing area according to a memory size requirement corresponding to the request, feeding back memory information of the target memory to a target slave node, so that the target slave node judges a corresponding memory type according to the memory information, and if a process corresponding to the target slave node executes a memory mapping method, sending the target memory to the target slave node. If yes, the target memory can be mapped to the proceeding virtual address space according to the target type and the offset corresponding to the target memory, so that the process can use the distributed target memory. Therefore, it can be guaranteed that when one node writes the memory, other nodes cannot synchronously write the memory, and then the data consistency of the shared memory accessed by multiple hosts is guaranteed.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Convolution operation method and device, electronic equipment and storage medium

The invention relates to the technical field of artificial intelligence chips, and provides a convolution operation method and device, electronic equipment and a storage medium, and the method comprises the steps: traversing convolution kernel elements, and determining a current to-be-loaded data block based on the shape of a target data block and the coordinates of the traversed current convolution kernel element block; based on the current to-be-loaded data block, determining a to-be-covered data block and a newly-added data block, and loading the newly-added data block into the shared memory; based on the initial position, reading the input data block from the shared memory, applying the input data block and the current convolution kernel element block, performing matrix multiply-accumulate operation, and accumulating an operation result to an output result; and determining a convolution operation result based on an output result obtained by accumulation after traversal is completed. According to the method and the device, only the newly added data blocks are loaded into the shared memory, so that repeated data loading can be avoided, the data volume loaded each time is reduced, and the memory bandwidth is saved.
Owner:SHANGHAI BIREN TECH CO LTD

Building decoration curtain wall construction progress intelligent monitoring method and system

The invention discloses an intelligent monitoring method and system for the construction progress of a building decoration curtain wall. The method comprises the steps that multi-source construction data of a construction site are collected through Internet of Things equipment and a high-definition camera; constructing a multi-dimensional tensor model and extracting data association features by applying a local reversible mapping algorithm to generate a standardized feature data set; performing parallel computing through shared memory hierarchical process mapping by utilizing the standardized feature data set, and outputting a project progress state evaluation result; constructing a decision diagram based on the project progress state evaluation result, and generating a construction progress deviation report; predicting a construction progress trend by adopting a sample optimal private regression algorithm and generating an optimization scheme; and visually displaying the construction progress prediction report and carrying out early warning. According to the invention, the problems of low data integration efficiency, delayed decision response and weak multi-dimensional data association analysis capability of traditional curtain wall construction monitoring are solved, and intelligent and efficient monitoring of the construction progress of the building decoration curtain wall is realized.
Owner:CHINA CONSTRUCTION SCIENCE & TECHNOLOGY DEVELOPMENT CO LTD

DCU-based high-performance sparse stiffness matrix vector multiplication method

The invention provides a DCU-based high-performance sparse stiffness matrix vector multiplication method, which comprises the following steps of: according to a sparse stiffness matrix, dividing a non-zero element into a plurality of calculation unit blocks by rows, pre-loading non-zero element data to an L1 shared memory or a register file through an on-chip shared memory controller of the DCU, a high-bandwidth crossbar switch of the DCU is used for realizing data copying and transmission; starting multi-row fusion execution for a short row of which the row non-zero element is lower than a DCU single-instruction multi-data width threshold value; constructing a wavefront scheduler based on a DCU asynchronous computing engine: binding an independent instruction cache region for each wavefront, and loading a multiply-add operation instruction set in advance through a prefetch instruction queue; a calculation unit state register is established, and when a wavefront scheduler ready signal is triggered, a scalar unit of the DCU is activated to execute calculation; and realizing cross-thread block reduction by adopting a DCU atomic operation accelerator. According to the method, the calculation throughput and the memory bandwidth utilization rate of large-scale structural mechanics stiffness matrix vector multiplication are effectively improved.
Owner:HENAN POLYTECHNIC

Network message processing method and apparatus, and computer device and storage medium

The present application belongs to the technical field of data processing and relates to a network message processing method and apparatus, and a computer device and a storage medium. The method comprises: on the basis of a received target network message, acquiring a target memory block from a pre-constructed shared memory pool, so as to generate a message memory; on the basis of a network-interface-card driver, performing first identification processing on the target network message, so as to obtain a target network message descriptor; on the basis of the target network message descriptor, performing second identification processing on the message memory, so as to obtain a first socket buffer; and performing parsing processing on the first socket buffer by means of a network protocol stack, and on the basis of a parsing processing result, reading network message data from a virtual memory corresponding to the message memory, so as to complete the processing of the network message. In the present application, on the basis of a constructed shared memory pool, zero-copy network packet reception can be realized during network message processing, thereby avoiding performance loss caused by the memory copy of messages.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Network card data local preprocessing system fused with edge computing

The invention discloses a network card data local preprocessing system fused with edge computing, and relates to the technical field of edge computing and artificial intelligence collaborative optimization. Comprising an edge computing unit, a hierarchical collaborative architecture, a model hot switching and generative fragmentation module, an intention recognition and adaptive scheduling module, a delay energy consumption optimization scheduling module, a CXL zero-copy sharing module, an edge computing unit integrated processor, an FPGA or ASIC and a neuromorphic computing unit. According to the method, an FPGA, an ASIC and a neuromorphic computing unit are integrated in an intelligent network card, microsecond-level dynamic connection reconfiguration and adaptive generative model fragmentation execution are realized through a reconfigurable Mesh interconnection matrix, an attention layer and a feed-forward layer of a Transform class model are fragmented and allocated to different computing units for parallel execution, and cross-card streamlined processing is realized in cooperation with a zero-copy shared memory. And the intention recognition module is deeply coupled with the model hot switching module, so that dynamic model switching and fragmentation strategy optimization based on service priorities and system loads are realized.
Owner:ZHUHAI SHININGDA TECH CO LTD

Miniature data transaction device and intelligent transaction system thereof

The invention relates to the field of edge computing and Internet of Things, and discloses a miniature data transaction device which comprises an edge computing unit used for executing a centralized or decentralized matchmaking algorithm and a transaction protocol; the multi-dimensional detection module comprises a compliance detection unit, a safety detection unit, a quality evaluation unit and a scene applicability evaluation unit; the hardware accelerator is used for accelerating a detection task and encryption operation; the communication module supports multi-protocol data transmission; the storage unit is used for caching transaction data and intelligent contract codes; the power management unit is used for controlling the total power consumption of the equipment, and the edge computing unit comprises a hierarchical trusted execution environment architecture consisting of a safe enclave and a general computing unit. According to the method, a hardware isolation architecture of the secure enclave and the general computing unit is adopted, and the dynamic key encryption shared memory technology is combined, so that the effects of privacy computing and efficient matching parallel execution are achieved.
Owner:CHENGDU PATZHILIHU DIGITAL TECHNOLOGY CO LTD

Communication method for user program and virtual machine on microkernel Hypervisor

The invention discloses a method for communication between a user program and a virtual machine on a microkernel Hypervisor, the microkernel Hypervisor is provided with two shared memory areas, the shared memory area 1 is accessed by the user program and a root service Rootserver, the shared memory area 2 is accessed by the Rootserver and the virtual machine, the user program writes communication request data with the virtual machine into the shared memory area 1, and the user program writes communication request data with the virtual machine into the shared memory area 2. The method comprises the following steps that a VMM sub-thread is used as a shared memory area 1, a Rootserver is notified through inter-process communication, the Rootserver reads communication request data from the shared memory area 1, the communication request data is written into a shared memory area 2 after being analyzed by the VMM sub-thread, a system calls a syscale to transmit a communication request to a kernel, the kernel injects virtual interrupt into a virtual machine, and the virtual machine sends the communication request to the Rootserver. And an interrupt processing program of the virtual machine processes the communication request and writes a processing result into the shared memory area 2, then the processing result is returned to the Hypervisor through the Hypercall, and a Rootserver of the Hypervisor feeds back the processing result to a user program through the IPC. The method is designed for the microkernel Hypervisor environment, and the overall performance and efficiency of the embedded virtualization system are improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Method for improving calculation speed of model based on mercuric chloride AI processor

The invention relates to a method for improving the calculation speed of a model based on a mercuric chloride AI processor. The method comprises the following steps: deploying the mercuric chloride AI processor and a deep learning model in a server, and carrying out adaptation and optimization on the mercuric chloride AI processor; before the data in the first buffer area is read, predicting and preloading the data to be processed, and loading the data from the global memory to the second buffer area in advance; a double-buffer mechanism is arranged in the mercuration AI processor, interrupt and event trigger points are set, the operation of loading data to a next buffer area is immediately started when a specified calculation stage is finished, and the time sequence of data circulation is accurately controlled; and a control parameter is automatically adjusted based on monitoring data of the real-time monitoring module, the calculation process is decomposed into a plurality of stages, different buffer areas are allocated for each stage, and access to the shared memory is optimized by setting a specified cache replacement strategy and the size of a cache line. According to the process, the calculation speed of the model in the deep learning field is improved, the resource utilization rate is improved, and the defect of manual adjustment and optimization is overcome.
Owner:四川华鲲振宇智能科技有限责任公司

Method for performing model reasoning by adopting GPU (Graphic Processing Unit), electronic equipment, computer readable storage medium and computer software product

The embodiment of the invention discloses a method for performing model reasoning by adopting a GPU (Graphics Processing Unit), electronic equipment, a computer readable storage medium and a computer software product. The method comprises the following steps: pre-loading and preheating a plurality of model instances in the GPU; setting a shared memory shared by at least two threads, and setting a state machine for the shared memory; the state machine can comprise a writable state and a non-writable state; when the state machine is in a writable state, writing a video frame into the shared memory, and setting the state machine to be in a non-writable state; calling the model instance loaded and preheated in the GPU to perform reasoning based on the video frames in the shared memory; and after the reasoning is completed, setting the state machine to be in a writable state. According to the embodiment of the invention, high efficiency and stability of model reasoning are realized, and calculation overhead and system complexity are reduced.
Owner:TURBULENCE (HANGZHOU) SOFTWARE ENGINEERING CO LTD +1

Attention mechanism calculation method and device, storage medium and product

The invention discloses an attention mechanism calculation method and device, a storage medium and a product, and the method comprises the steps: carrying out matrix multiplication operation through employing a query matrix block of a first register block and a key matrix block of a shared memory, obtaining a first product matrix block, and writing the first product matrix block into a second register block; performing exponential operation by using the first product matrix blocks to obtain sub-matrix blocks, writing the sub-matrix blocks into a second register block in a covering manner, and writing the sub-matrix blocks into a third register block in the form of a target precision type; performing matrix multiplication operation by using the sub-matrix blocks of the third register block and the value matrix blocks of the shared memory to obtain second product matrix blocks, and writing the second product matrix blocks into a second register block; performing softmax operation by using the second product matrix blocks to obtain attention result matrix blocks, and writing the attention result matrix blocks into a fourth register block; and writing the attention result matrix of the fourth register group into the shared memory in blocks. According to the embodiment of the invention, overflow of the register can be avoided, and the utilization of hardware resources is maximized.
Owner:SHANGHAI BIREN TECH CO LTD

Attention calculation implementation method and device, medium, equipment and product

The invention discloses an attention calculation implementation method and device, a medium, equipment and a product, and the method comprises the steps: loading the current to-be-calculated ith query block from a shared memory to a first register group of a consumer thread group, and carrying out the internal splitting calculation of the query block and a key block, so as to obtain corresponding attention score blocks; sequentially obtaining MK attention score blocks, and executing attention fusion calculation of mixing precision with the corresponding value blocks to obtain an attention output block corresponding to the ith query block; and after the attention output blocks are logically divided into N2 batches, a specified register group for storing target precision type data in the fusion calculation process is multiplexed and executed according to batches, target precision type conversion is carried out, and an output result after conversion of each batch is written back to a shared memory. According to the method, the existing hardware resources can be efficiently utilized to improve the calculation performance, and the method is particularly suitable for large-size query block and key block scenes.
Owner:SHANGHAI BIREN TECH CO LTD

Model scheduling method and device based on multi-level cache, equipment and medium

The invention discloses a model scheduling method and device based on multi-level cache, equipment and a medium, and relates to the field of artificial intelligence, and the method comprises the steps: in a model deployment process, updating the current model popularity of a corresponding to-be-reasoned model in a preset cache architecture based on an obtained model access request; wherein cache layers of the preset cache architecture are respectively a video memory cache layer, a process memory cache layer, a shared memory cache layer and a persistent cache layer from top to bottom; determining a target cache layer of the to-be-reasoned model in a preset cache architecture, and when the target cache layer is a video memory cache layer, triggering a preset reasoning operation on the to-be-reasoned model; and when the target cache layer is a non-video memory cache layer, scheduling the to-be-reasoned model to the video memory cache layer based on each current model heat in other cache layers located above the target cache layer, so as to trigger a preset reasoning operation on the to-be-reasoned model in the video memory cache layer. Therefore, efficient utilization and intelligent management of model resources can be realized.
Owner:HANGZHOU SHIQU INFORMATION TECH CO LTD

Metadata access method and apparatus, device, storage medium, and program product

A metadata access method, apparatus, and computer-readable storage medium for efficient metadata retrieval through cache management. The method receives metadata query requests including target index information from processes and performs matching operations on a global cache file containing records with index information and slot identifiers. Each cache description array corresponds to memory blocks caching metadata. Upon successful matching, the target cache description array is accessed using the target slot identifier. The data state of target metadata is determined from the cache description array, and target address information indicating the location of the target memory block in shared memory is obtained and returned to the requesting process, enabling efficient shared memory-based metadata access.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

MQTT message transmission optimization method and system

The invention discloses an MQTT message transmission optimization method and system. The method comprises the steps of analyzing a theme, extracting a device type, a data feature and a geographic position triple, calculating a hash value, and mapping a device to a specified Broker fragment cluster node to generate a fragment mapping table; processing the equipment data in the edge domain in the fragment mapping table through an edge calculation layer, removing invalid data according to a preset rule, merging the equipment data in the same fragment node, and embedding a fragment node ID for an aggregation message generated after merging; identifying a fragment node ID and routing to a target fragment node, positioning a corresponding shared memory pool, and writing the message into the shared memory pool; and responding to a direct access request of the client, so that the client directly accesses the data from the shared memory pool through the user mode network stack. According to the method, the problem of uneven load is effectively solved, multiple times of state switching in the data transmission process is avoided, and the MQTT message transmission efficiency is improved while the transmission cost is reduced.
Owner:GUANGZHOU SIYUN DATA TECH CO LTD

Protocol-Aware Provisioning of Resource over CXL Fabrics

Dynamic provisioning of resources in a datacenter improves the utilization efficiency of compute, memory, storage, and network resources, while maintaining flexibility to meet changing demands. Embodiments herein disclose protocol-aware provisioning of resources over Compute Express Link (CXL) fabrics, enabling multi-protocol pooling and disaggregation of resources, including memory and workload-specific accelerators such as GPUs and DSAs. In some embodiments, a Resource Provisioning Unit (RPU) facilitates intent-based protocol translations and mappings between address spaces, potentially enabling the creation of large-scale compute-memory fabrics that may utilize both coherent and non-coherent Non-Transparent Bridging (NTB) between multiple protocols in a single system, serving as the underlying infrastructure for executing workloads such as Large-Language Models (LLMs) and Deep Learning Recommendation Models (DLRMs). Some embodiments also optimize low-latency communication between processes running on different nodes by enabling host-to-host memory provisioning for libraries such as OpenMP, Pthreads, or CUDA, which can utilize shared memory.
Owner:HYATT GAYA OPAL MS +1

Access method, device and equipment and computer readable storage medium

The invention discloses an access method, device and equipment and a computer readable storage medium, which are applied to the technical field of computers, and comprise the following steps: loading a consistency label list corresponding to an authorized access consistency domain; when an access request is initiated to a target consistency domain in the shared memory, performing permission verification on the access request through the consistency label list and the target consistency domain; when the permission verification is passed, determining whether an access conflict exists or not through a local access state cache and the type of the access request; if the access conflict exists, a controller arbitration engine is triggered, and the access request is executed according to an arbitration result; and if the access conflict does not exist, executing the access request. According to the method, the consistency label list and the access state cache are formulated for each device, the consistency problem existing when multiple devices access the shared memory in an existing system is solved, and efficient, low-conflict and flexible consistency guarantee can be achieved under the scene that multiple devices concurrently access the shared memory.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Multi-core log management architecture and method based on power gap operating system and computer program product

The invention provides a multi-core log management architecture and method based on a power gap operating system and a computer program product, which are suitable for the technical field of computers, and comprise a master core, a slave core and a shared memory area, the slave core further comprises a log generation module, a buffer module and a log uploading module; the main core is responsible for operating a power gap operating system; the slave core is responsible for running a naked running program or a real-time operating system; the log uploading module writes target log data in the buffer module into a shared memory area when the data cached in the buffer module reaches a preset condition; under the condition that the master core monitors that the slave core uploads the target log data to the shared memory area in the shared memory area, the target log data in the shared memory area is read, and integration, classification and storage operation is conducted on the target log data in sequence. Therefore, the efficiency problem of log collection, storage and management in the multi-core heterogeneous architecture can be solved, the characteristics of the master core and the slave core are fully utilized, and the log processing capacity of the system is improved.
Owner:CYG SUNRI CO LTD +1

Multi-intersection traffic signal cooperative control method driven by cross attention neural network

The invention discloses a multi-intersection traffic signal cooperative control method driven by a cross attention neural network, and the method employs a local cooperative Transform architecture, integrates a decision converter and a shared memory mechanism, and achieves the efficient modeling of a space-time dependence relation of a multi-intersection traffic state. The method comprises the following steps: firstly, through a memory head module, extracting a hidden state of each agent in a sequence modeling process, and updating global shared memory for supporting information interaction and strategy collaboration among multiple agents; and then, a cross attention module is adopted to carry out cross calculation on the local representation and the shared memory of each agent, so that dynamic perception and efficient modeling of the global state of the traffic system are realized. A backbone network of the model is based on Transform, and the understanding ability of time and space traffic characteristics is enhanced through position coding, self-attention and cross attention mechanisms. In the fine tuning stage, only the inserted Adapter module and the output layer are subjected to parameter updating.
Owner:NANJING TECH UNIV