Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

964 results about "Batch processing" patented technology

Computerized batch processing is the running of "jobs that can run without end user interaction, or can be scheduled to run as resources permit."

Enterprise data dynamic integrated management system based on lightweight

PendingCN121301619ABiological modelsOther databases indexingManufacture execution systemBatch processing
The invention relates to the technical field of enterprise data management, in particular to a lightweight-based enterprise data dynamic integrated management system, which is characterized in that an acquisition module is used for deploying edge computing nodes, receiving multi-source heterogeneous information streams from manufacturing execution systems and equipment logs, dynamically analyzing and standardizing the information streams, and adding metadata tags; uploading is carried out in a batch processing mode; the map construction module is used for constructing a semiconductor blood relationship map by taking the standardized key information as a blood relationship clue; a graph database is used for efficient storage, and a RESTful API interface is configured to support batch import, so that data storage and relevance expression are more flexible and efficient; the prediction module performs reasoning on the atlas by adopting a graph neural network to generate predictive risk distribution, and a correlation analysis set generated by the prediction module is stored back to the atlas in a structured manner; by introducing a multi-thread concurrent write-in and lock mechanism, the atlas supports complex combination query based on a Cypher query language, and supports multi-level and traceability query.
Owner:NANJING SPEED DISTRIBUTION INFORMATION TECHNOLOGY CO LTD

TR component gold wire bonding process parameter prediction method based on multilayer perceptron neural network

The invention discloses a TR assembly gold wire bonding process parameter prediction method based on a multilayer perceptron neural network, and belongs to the technical field of microwave device intelligent manufacturing. According to the method, an intelligent mapping model of gold wire bonding geometric parameters and radio frequency performance is constructed by fusing a multi-layer perceptron neural network and parameterized electromagnetic simulation. The method specifically comprises the following steps: generating 45 groups of samples in a process parameter space by adopting Latin hypercube sampling; obtaining an S parameter data set through batch processing electromagnetic simulation; box-Cox conversion and normalization preprocessing are carried out on the data; the method comprises the following steps: constructing an MLP neural network model of a 3-32-16-2 structure, and determining hyper-parameters by using Bayesian optimization; and after training is completed, rapid reverse mapping from target performance to process parameters is realized. According to the method, the number of traditional tests is reduced from more than 200 to 45, the predicted root-mean-square error of S21 is smaller than or equal to 0.12 dB, the determination coefficient is larger than or equal to 0.96, and the parameter backstepping time lt is obtained; according to the method, full-process automation from simulation, training, optimization to production and issuing is realized, and the development efficiency of the TR component is remarkably improved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Model Controller Framework for Automated Model Deployment & Monitoring

The invention provides a system and method for managing the lifecycle of machine learning models, from development to deployment and ongoing operation, across various environments including on-premises, cloud, and hybrid infrastructures. The system features a model build platform for data processing, feature generation, model development, training, and hyperparameter tuning. A model analytics engine extracts metadata, performs complexity analysis, and generates configuration files specifying environment settings and resource needs. A secure model repository enables version-controlled storage, while a deployment platform retrieves, validates, and deploys models in containerized environments like OpenShift or Kubernetes. The platform dynamically allocates resources, supports real-time and batch scoring, and monitors model performance with guardrails. Customizable agents provide real-time feedback and automated optimization, and the system can securely decommission models while maintaining detailed lifecycle records. The invention enhances the efficiency, security, and scalability of machine learning operations with continuous performance improvement and compliance automation.
Owner:BANK OF AMERICA CORP

Large model dynamic batch processing method based on sequence splicing

The invention discloses a large model dynamic batch processing method based on sequence splicing. The method comprises the following steps: receiving input sequences of a plurality of users in a batch; converting the input sequence into a corresponding token sequence; carrying out heterogeneous splicing on all token sequences along a sequence length dimension to form a unified joint token; the spliced joint tokens pass through a normalization layer, standardization operation is executed on the joint tokens, and data distribution is unified; carrying out linear projection on the joint token through a shared linear transformation layer to generate a joint query vector, a joint key vector and a joint value vector, and splitting the joint query vector, the joint key vector and the joint value vector into sub-vector groups corresponding to each user; executing multi-head attention calculation to obtain attention output of the user; and carrying out linear transformation on the attention output, and inputting a transformed result into a shared MLP to carry out nonlinear feature extraction and enhancement so as to obtain a final output corresponding to each user. According to the method, the problems of efficiency bottleneck and resource consumption when a large model processes mass data are effectively solved.
Owner:VISIOCO (SUZHOU) TECHNOLOGY CO LTD

AI big data real-time processing and analysis method

The invention discloses an AI big data real-time processing and analysis method, and solves the problems of insufficient real-time performance and resource waste in traditional processing. The method comprises the steps of collecting heterogeneous data from multiple sources, dividing priorities through feature vector construction and a dynamic evaluation model, and shunting to an edge rapid processing channel, an edge-cloud collaboration channel and a cloud batch processing channel. The edge node preprocesses the high / medium priority data and compresses an analysis result in a layered manner; and the cloud receives the compression result, the middle-priority residual data and the low-priority data, a unified view is established through fusion, and the edge analysis model parameters are iteratively optimized in real time. And finally feeding back the edge preliminary analysis and the cloud depth result to the terminal. According to the method, through dynamic distribution, cooperative processing, differential compression and model iteration, real-time response of high-priority data, efficient resource allocation and continuous improvement of analysis precision are achieved, and the method is suitable for multi-scene heterogeneous data processing.
Owner:XIAMEN MANLIN INFORMATION TECHNOLOGY CO LTD

Knowledge full life circle management system and construction method

The invention relates to the technical field of knowledge management systems, and discloses a knowledge full-life-cycle management system and a construction method, and the system comprises a technical platform and multi-modal perception layer, a calculation and arrangement engine layer, a unified API gateway layer, a multi-modal processing module layer, a knowledge graph layer, a cognitive reasoning layer and an application layer. Real-time collection and batch processing of multi-source heterogeneous data are achieved through a standardized SDK / API; spark / Flink is combined with Kubernetes to complete the ETL (Extract Transform Load) and resource scheduling of the multi-modal data; constructing an entity-relationship-attribute knowledge graph through cross-modal feature fusion; intelligent decision support is realized based on rule reasoning, graph calculation and GNN; the application layer provides intelligent questioning and answering, decision support and personalized recommendation services; the technical problems of knowledge islands, sharing barriers, knowledge statics and the like in the prior art are solved.
Owner:CHONGQING VISION INFORMATION IND GRP CO LTD

Electric power marketing business abnormity real-time detection method and system based on stream-oriented computing

The invention relates to an electric power marketing business abnormity real-time detection method and system based on stream-oriented computation, and belongs to the technical field of electric power system optimizing.The method comprises the steps that data snapshots are extracted from an electric power marketing business system, difference comparison is conducted on the data snapshots and historical snapshots of an intermediate library, and standardized increment events are generated and stored; capturing an incremental event in real time through a data change capturing tool and pushing the incremental event to a message queue; a streaming computation engine consumes the event stream, sequentially performs data cleaning, association with a static dimension table and sliding window statistical feature calculation, and constructs a feature vector; and performing parallel analysis and weighted fusion on the feature vectors based on a business rule base and an online machine learning model to generate a comprehensive risk score, and outputting an abnormal event when the score exceeds a threshold value. According to the method, the problems of exception identification lagging and complex work order process in a traditional batch processing mode are solved, the crossing of the business risk from hour-level detection to minute-level real-time perception is realized, and the timeliness and accuracy of power marketing risk management and control are improved.
Owner:FUJIAN ELECTRIC POWER CO LTD XIAMEN ELECTRIC POWER SUPPLY CO +1

Container-based parallel computing system

A container-based parallel computing system for executing high-performance computing (HPC) applications. The system leverages container technology to package the applications executed at the nodes in a cluster. To load and execute a job in the parallel computing system, containers are deployed in a cluster that include all the application resources and configuration information that the particular HPC application needs to execute. An event-driven batch scheduler may be used to dynamically allocate resources for executing multi-node jobs in the container-based parallel computing system, handling the coordination of resource allocation for the customer. The scheduler insures that jobs begin executing as fast as possible, and handles failure conditions such as partial scaling. Virtual network interfaces are attached to the containers that allow the containers to connect to and communicate with other containers in the cluster directly through the network interfaces of host machines using IP addresses provided by the virtual network interfaces.
Owner:AMAZON TECH INC

Intelligent manufacturing real-time decision-making method and system based on digital twinning

The invention discloses an intelligent manufacturing real-time decision-making method and system based on digital twinning, and relates to the technical field of intelligent manufacturing. Comprising the following steps: S1, collecting multi-source operation data in a physical production system in real time; s2, dynamically constructing and updating a digital twinborn model of the physical production system based on the multi-source operation data acquired in real time; and S3, based on the digital twin model, monitoring the operation state of the physical production system in real time. According to the method, the multi-source operation data of the physical production system are collected in real time, and the digital twin model is dynamically updated, so that no lag deviation between the model and the physical system is ensured; when a dynamic disturbance event is identified, event-driven real-time simulation and predictive analysis are started immediately, delay caused by traditional batch processing is avoided, and disturbance influence can be accurately evaluated in combination with historical data and expert experience; and then an optimal decision scheme is generated and screened in real time through a multi-objective optimization algorithm.
Owner:QINGDAO WONGOING INFORMATION TECH CO LTD

Multi-source fusion maintenance work order association early warning method and system

The invention relates to the technical field of information processing and artificial intelligence, and discloses a multi-source fusion maintenance work order association early warning method and system. The method is used for solving the problem that the accuracy and timeliness of early warning are insufficient due to comprehensive modeling of multi-modal data in a traditional method. The method comprises the steps that firstly, a work order, a log event, a monitoring index and asset information are accessed, and after time caliber and field caliber unification and quality labeling are completed, the work order, the log event, the monitoring index and the asset information are written into a real-time data lake and a batch processing data warehouse respectively; normalizing the work order text, extracting a key identifier, and generating a feature index bound with the work order number; aligning the log and the index data according to a work order time window to form an event fragment and an index fragment, and associating an asset node; and constructing a dependency graph based on an asset ledger and a deployment link, forming a work order candidate pair, carrying out classification and hierarchical pushing, and recording feedback and version information for tracing.
Owner:SI-TECH INFORMATION TECH CO LTD

Multi-level data storage and mixed query method and system, storage medium and equipment

The invention relates to the technical field of big data storage, and discloses a multi-level data storage and mixed query method and system, a storage medium and equipment, and the method comprises the steps: collecting original buried point data, carrying out real-time stream processing, storing the processed real-time data in a hot storage layer, carrying out timing offline batch processing on historical buried point data, and storing the processed real-time data in a hot storage layer; converting into an optimized storage format and storing in an object storage system; an extended metadata management system is constructed, and structure information, partition information and access frequency information of data in the object storage system are automatically extracted and managed; executing query, and reducing the data scanning amount through a query optimization strategy; automatically migrating data among different storage layers through a dynamic cold and hot data layering mechanism; a unified query interface is provided, a query request is distributed to a corresponding real-time processing system or an offline query system based on an intelligent routing strategy, and a query result is cached and returned, so that high-performance and low-cost storage and second-level query of data are realized.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Real-time target detection method and system based on RTSP flow and NPU collaborative optimization

A real-time target detection method and system based on RTSP flow and NPU collaborative optimization belong to the technical field of computer vision and artificial intelligence, and are characterized by comprising the following steps: adopting dynamic memory optimization management of a hybrid pipeline architecture, and performing single-time continuous copying through a CPU (Central Processing Unit); through deep integration of innovative technologies such as DMA direct transmission NPU continuous memory pool management, dynamic batch processing scheduling, hybrid assembly line processing, intelligent equipment load balancing, parallel preprocessing optimization and vectorization post-processing, RTSP flow collaborative optimization and the like, the NPU utilization rate is improved to 85% or above, the overall average FPS is improved by 200% or above, the assembly line parallelism degree achieves three times of performance gain, and the production efficiency is greatly improved. The data transmission delay is reduced by 80%, the system stability is remarkably improved, performance degradation is avoided after long-time operation, and the method is suitable for various scenes such as edge calculation and cloud reasoning.
Owner:XIAN KEYWAY TECH

Network fault self-healing and prediction maintenance method based on artificial intelligence

The invention relates to the technical field of network fault maintenance, in particular to a network fault self-healing and prediction maintenance method based on artificial intelligence, and the method comprises the steps: S1, constructing a multi-source data real-time collection framework; s2, deploying a lightweight A I model at an edge node; s3, introducing an interpretable AI technology; s4, constructing a causal reasoning module; s5, designing a dynamic self-healing strategy library; s6, establishing an online model learning mechanism; s7, developing a simulation verification environment; and S8, realizing a man-machine cooperative operation and maintenance workflow. According to the scheme, data is subjected to streaming preprocessing and intelligent dimension reduction at the source, the transmission load is greatly reduced, the processing efficiency is improved, real-time, near-real-time and batch processing tasks are further distinguished through the edge side parallel assembly line technology, it is ensured that key indexes are preferentially processed, and compression and acceleration are achieved on the model level through knowledge distillation, quantification and pruning technologies.
Owner:WUXI YUANSHUCHENG TECHNOLOGY CO LTD

Multi-tenant visual large model reasoning resource dynamic allocation and isolation method

The invention provides a multi-tenant visual large model reasoning resource dynamic allocation and isolation method. An integrated system of a multi-tenant environment, visual large model reasoning, dynamic resource allocation and isolation guarantee is constructed. The system controls resources such as model copy number, video memory quota, batch size, queue weight, bandwidth and the like at tenant level fine granularity, and adjusts priority and quota based on closed-loop feedback of real-time indexes (such as queuing length, delay and throughput). Interference suppression among tasks is realized through priority grading, a hard / soft isolation strategy and a tenant-model copy mapping mechanism in combination with GPU MIG, video memory partitioning, queue scheduling and other technologies. Aiming at the characteristics of high video memory, large input and the like of a visual large model, model loading, batch processing and video memory multiplexing strategies are optimized, and differential resource allocation of heterogeneous models is supported. According to the overall scheme, on the premise that the service quality and isolation are guaranteed, the GPU resource utilization rate is remarkably increased, and the operation cost is reduced.
Owner:CHINA COAL TECH & ENG GRP CHONGQING RES INST CO LTD

Multi-modal document data processing method and system oriented to large language model training

ActiveCN121093293ANeural learning methodsBatch processingCharacter (computing)
The invention discloses a multi-modal document data processing method and system for large language model training, and the method comprises the steps: receiving a plurality of original documents in various formats, extracting the structure information of each original document, and recognizing a text region and an image region of each original document based on the structure information; performing optical character recognition on the text region and the image region by adopting a parallel OCR (Optical Character Recognition) engine based on GPU (Graphics Processing Unit) acceleration and heterogeneous calculation to generate recognition text data of the corresponding original document; performing multi-dimensional quality evaluation and cleaning on the recognition text data of each original document, and outputting normalized text data; and storing the standardized text data into a distributed knowledge base according to a predefined structure, and performing copyright and compliance test on the standardized text data. By adopting a parallel OCR recognition engine based on GPU acceleration and heterogeneous calculation, efficient and high-precision batch processing of multi-modal documents is realized, and the processing speed, the recognition accuracy and the data quality are improved.
Owner:HANGZHOU BINGTE TECH

Data sharing and storage model based on double-layer block chain

The invention provides an Internet of Vehicles data sharing and storage model based on a double-layer block chain, and aims to solve the problems of low data interaction efficiency, poor expansibility, insufficient security and the like in the traditional Internet of Vehicles. According to the model, road side units are grouped geographically, an optimized Raft protocol is adopted in each group to realize rapid consensus, and an improved PBFT protocol is adopted among the groups to ensure global consistency. Self-adaptive leader election, batch processing and assembly line mechanisms are introduced into the Raft protocol, so that the throughput and response speed of the system are improved; a reputation value-based VRF main node election mechanism and a BLS aggregation signature technology are introduced into the PBFT protocol, so that the communication complexity is reduced, and the system security and fault-tolerant capability are enhanced. According to the method, the fault-tolerant rate and the message complexity of the system are analyzed theoretically, the advantages of the system in the aspects of throughput, time delay and fault-tolerant performance are verified through simulation experiments, and the method is suitable for large-scale and high-concurrency car networking application scenes.
Owner:BEIJING TECH & BUSINESS UNIV

Layered self-adaptive full block pre-filling scheduling method and system for large language model reasoning

The invention discloses a hierarchical self-adaptive full block pre-filling scheduling method and system for large language model reasoning, and the method comprises the steps: carrying out the hierarchical portrait analysis of a to-be-served model, and dividing the to-be-served model into partitions with different calculation characteristics according to the calculation intensity and memory access characteristics of each layer; then, making a layering and partitioning strategy based on a partitioning result, allocating a larger partitioning size to a calculation-intensive partition, allocating a smaller partitioning size to a memory bandwidth-intensive partition, and generating a layering and partitioning mapping table; and finally, when the online scheduling is executed, querying the mapping table according to the request processing progress to determine the block target size, and jointly forming a batch processing unit by the decoding task and the pre-filled block with the heterogeneous size under the constraint of the iteration time budget to be executed. According to the method, accurate matching of calculation and bandwidth resources is achieved, the system throughput can be effectively improved, tail delay and fluctuation thereof can be remarkably reduced, bubbles under pipeline parallelism are reduced, and the method is suitable for various attention mechanisms and distributed reasoning architectures.
Owner:ZHEJIANG LAB

Event processing method, system and device based on Internet of Things device and medium

The invention provides an event processing method, system and equipment based on Internet of Things equipment and a medium, and belongs to the technical field of Internet of Things. The method comprises the following steps: receiving an Internet of Things equipment message through an MQTT Broker, identifying an event type according to a theme, and distributing the event type to a corresponding thread pool; concurrently analyzing the JSON load in each thread pool, and converting the JSON load into a structured data object; caching the data in a batch processing queue in a classified manner, and triggering batch processing based on a queue state; extracting batch data, generating an optimized SQL statement, executing writing, and processing exceptions; screening data based on a routing rule and asynchronously forwarding the data to message middleware; meanwhile, processing parameters are dynamically adjusted according to the equipment state data and the system load. According to the invention, efficient processing of Internet of Things equipment events is realized, and the system throughput, the resource utilization rate and the data processing real-time performance are improved.
Owner:SHANDONG INSPUR ULTRA HD INTELLIGENT TECH CO LTD

Fire-fighting spatio-temporal trajectory data processing and off-line warehouse counting construction method based on stream batch fusion architecture

The invention discloses a fire-fighting spatial-temporal trajectory data processing and off-line data warehouse construction method based on a stream batch fusion architecture, and relates to the field of digital fire fighting, and the core steps of the method are as follows: after multi-source data are collected by an edge gateway, high throughput distribution and storage are realized by Kafka; real-time optimization such as track cross repair and offset correction is completed through Flink streaming calculation; performing offline batch processing on the historical data by the Spark to generate an analysis result; generating an optimal rescue path through a distributed Dijkstra algorithm in combination with road topology, real-time traffic and fire fighting truck parameters; the method comprises the following steps: realizing mixed storage by adopting ElasticSearch secondary index and HBase total storage; based on hot and cold data hierarchical management, intelligent applications such as fire-fighting commanding and dispatching, battle comment and review and the like are supported. Practice verifies that the method can shorten the fire control time of the main urban area by 60%, significantly improves the fire emergency response capability, and is suitable for large-scale fire track data scenes.
Owner:ANHUI TELECOMM PLANNING & DESIGNING

High-concurrency large language model high-speed reasoning deployment method

The invention relates to the technical field of large language models, in particular to a high-concurrency large language model high-speed reasoning deployment method, which comprises the following steps of: extracting a weight coefficient corresponding to a user identity identifier, an emergency degree numerical value mapped by a request type and a resource occupation value converted by an estimated calculation amount according to the user identity, the request type and the estimated calculation amount; according to the method, the dynamic priority is generated by performing comprehensive operation on the user identity, the request type and the estimated calculation amount, so that differentiated services for reasoning tasks are realized, interactive requests with high urgency degree can bypass batch processing tasks with long time consumption, the response delay and delay jitter of key services are reduced, and the user experience is improved. Meanwhile, according to the complexity of the request text and the geographic coordinates of the user, the task is intelligently routed to the model node with the most suitable scale in the distributed network, so that the wide area network transmission delay is greatly reduced through edge processing, and the computing power waste caused by using a super-large-scale model to process the simple task is also avoided.
Owner:GUANGXI SHUZHI PUBLISHING MEDIA CO LTD

KV-cache streaming for improved performance and fault tolerance in generative model serving

A method of serving a generative transformer model includes determining a batch size to use in processing inference requests and allocating at least one prompt pipeline and at least on token pipeline to the generative transformer model to process the batch of inference requests. The number of prompt pipelines and the number of token pipelines, and the depths of the pipelines are determined based on the batch size, an average prompt length, a cache requirement per stage, and a memory footprint of model weights for the generative model per stage using a resource allocator component of the model serving system. Cache streaming is used to stream prompt cache from prompt pipelines to token pipelines to generate tokens. Cache streaming involves gather-copy operations which may be performed using compute kernels.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Batch processing job automatic debugging method, device and equipment and storage medium

The invention discloses a batch processing job automatic debugging method and device, equipment and a storage medium, and relates to the technical field of data processing and debugging verification, and the method comprises the following steps: obtaining a debugging copy; replacing a target table in the debugging copy, and determining debugging copy information; deploying a corresponding batch processing script based on the debugging copy information by using the allocated temporary account necessary permission, and determining a target batch processing operation script; and performing batch processing job debugging based on the target batch processing job script to generate a batch processing job debugging result. According to the method, the target table in the debugging copy is replaced, the batch processing script is deployed according to the distributed temporary account minimum permission, batch processing operation debugging is carried out on the target batch processing operation script, the batch processing operation debugging result is generated, and the temporary account only having the minimum necessary permission is dynamically generated for each time of debugging. And a special temporary table is automatically created as a safe output environment, so that authority distribution according to needs and operation sandbox isolation are realized.
Owner:CHINA MERCHANTS BANK

Dynamic batching for inference system for transformer-based generation tasks

An inference system applies a machine-learning transformer model to a batch of requests with variable input length or variable target length or variable internal sate length by selectively batching a subset of operations in the transformer model but processing requests in the batch individually for a subset of operations in the transformer model. In one embodiment, the operation to be processed individually is an attention operation of an encoder or a decoder of the transformer model. By selective batching, the inference system can allow batching operations to be performed for a batch of requests with variable input or target length or internal state length to utilize the parallel computation capabilities of hardware accelerators while preventing unnecessary computations that occur for workarounds that restrain the data of a batch of requests to a same length.
Owner:FRIENDLIAI

Partitioning and hierarchical caching-based track point visual rendering method adaptive to localization

The invention discloses a localization-adaptive track point visual rendering method based on partitioning and hierarchical caching, and the method comprises the steps: S1, inputting a track data stream, carrying out the reading according to the source type of the data stream, and carrying out the data partitioning and LOD processing of the track point data; s2, performing data analysis and rendering on the data; s3, carrying out batch processing on GPU vertexes, according to the state of the circular buffer, if the vertexes are writable, incrementally writing vertex data, and if Wrap is needed, carrying out loopback writing and updating a pointer; and S4, finally, the descriptors and the styles are updated, rendering increment drawing is carried out on the track points, and visual display is carried out. The method supports real-time rendering and playback of ten-million-level track points, has visual and instantiated rendering of window dynamic loading, space-time partitioning, hierarchical caching and WebWorker decoupling calculation, and improves the space-time data rendering capability of the domestic autonomous controllable field by ten thousand times; and the user experience is greatly improved by a non-perceptual interaction and transition smoothing algorithm.
Owner:NANJING HONGSONG INFORMATION TECH CO LTD

Dynamic batch processing and delay optimization method for deep learning model reasoning service

The invention discloses a dynamic batch processing and delay optimization method for a deep learning model reasoning service, and the method comprises the steps: constructing a system architecture which comprises a request classification module, a dynamic batch processing scheduling module, a delay prediction and optimization module, and a reasoning execution module; accurate classification of reasoning requests, intelligent scheduling of dynamic batch processing and effective prediction and optimization of reasoning delay are realized. Wherein the dynamic batch processing scheduling module adopts a self-adaptive batch processing strategy based on reinforcement learning, and dynamically adjusts the batch processing size in combination with the priority, the type and the delay demand of a request; and the delay prediction and optimization module predicts reasoning delay by using a deep learning model, and performs optimization through resource dynamic allocation and model parameter adjustment. Experimental results show that the batch processing efficiency can be remarkably improved and the reasoning delay can be reduced on the premise of ensuring the reasoning precision.
Owner:GUANGDONG KAITUO DIGITAL INNOVATION TECHNOLOGY CO LTD

RDMA queue pair multiplexing method based on shared memory

An RDMA queue pair multiplexing method based on a shared memory efficiently maps a large number of logical queue pairs (LQPs) to a small number of physical queue pairs (PQPs), and combines shared memory zero-copy communication and O (1) active queue management, so that ten thousand-level LQPs are supported by a small number of PQPs, the overhead of a single-connection memory is reduced by about 60%-64%, and context jitter of a network interface card (NIC) is avoided; the connection establishment time delay and the short flow end-to-end time delay are greatly reduced; through O (1) active tracking and batch processing, CPU idling and system calling overhead are greatly reduced, and high throughput and expandability are achieved; the method can be implemented in a pure user mode, a kernel or NIC does not need to be modified, and the method can be immediately deployed in an existing RDMA environment.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Decentralized high-performance multicast method and system based on Gossip and RDMA

The invention discloses a decentralized high-performance multicast method and system based on Gossip and RDMA, and relates to the field of computer network communication. The method has better reliability, realizes zero-loss transmission of data packets through message de-duplication, timeout retransmission, ACK / NACK convergence and a dynamic topology adaptation mechanism, ensures data consistency, and solves the problem of insufficient reliability of traditional UDP multicast; the method has the advantages that the real-time performance is high, the delay is low, the throughput is high, the zero copy and kernel bypass characteristics of the RDMA and the batch processing mechanism of the MPMC are combined, the average delay is lower than 10 microseconds, the throughput in a large data packet scene reaches 374GB / s, and the low-delay requirements of high-frequency transaction, high-performance calculation and the like are met; the method has high expansibility, is based on a decentralized architecture of an improved Gossip protocol, does not need core node maintenance topology, supports dynamic joining or quitting of nodes, can adapt to large-scale cluster expansion, and solves the expansibility bottleneck of centralized multicast.
Owner:CHONGQING UNIV OF TECH

Distributed read-write method and device for large model training

The invention relates to the technical field of computer data storage and distributed systems, in particular to a distributed read-write method for large model training, which comprises the following steps of: constructing a hierarchical structure of local cache, distributed shared cache and persistent storage, and maintaining a global cache table copy at each client to register file blocks and cache positions. And during writing, firstly writing in a local cache and copying to other client nodes, and submitting persistent storage after registration is completed. During reading, the file blocks are obtained from the local cache and the distributed shared cache in sequence, if the file blocks are not hit, the file blocks are obtained from the persistent storage, meanwhile, the effective file blocks are written back to the local, and table items are supplemented. Unified visibility is achieved through position mapping and version constraint, nearby reading and fault recovery are achieved, back-end pressure and access delay are reduced, and the method is suitable for large-scale training and batch processing scenes.
Owner:CHONGQING ZHONGKE YUNCONG TECH CO LTD +1

Intelligent data management method and system

The invention provides an intelligent data management method and system, and relates to the technical field of data processing. Standardized packaging and automatic execution of the data governance instruction are achieved, and the problem that a traditional storage medium cannot dynamically adapt to the multi-source data processing requirement is solved. Through the decoupling design of the pre-compilation instruction and the runtime environment, the same storage medium can adapt to data management nodes of different hardware architectures, and the deployment complexity during system upgrading is effectively reduced. Furthermore, based on an instruction execution mechanism of a streaming computing engine, the real-time performance of data quality evaluation and early warning triggering is ensured, and decision lag caused by batch processing delay is avoided. Through the on-demand loading characteristic of the modular instruction set, the function expansion requirement of a specific business scene is supported while the core governance function is ensured.
Owner:WUHAN ENYI INTERNET TECH CO LTD

Differentially private stochastic gradient descent using optimized correlation matrices

Methods, systems, and apparatuses, including computer programs encoded on computer storage media, for training a neural network using differentially private stochastic gradient descent (DP-SGD) with correlated noise. In some implementations, a system determines a correlation matrix by performing an optimization to improve a utility metric that is dependent on a noise multiplier. The system can then train the neural network using the optimized correlation matrix and a privacy-amplifying batching scheme to produce a neural network with high utility while satisfying a data security criterion.
Owner:GDM HOLDING LLC