Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

99 results about "High performance computation" patented technology

High-Performance Computing. High Performance Computing (HPC) is the IT practice of aggregating computing power to deliver more performance than a typical computer can provide.

A multi-node server network topology and energy consumption-aware task scheduling method

The application discloses a kind of multi-node server network topology and energy consumption perception's task scheduling method, belong to data center and high performance computing technical field.It includes: constructing the weighted directed graph topology model including server node and switch node, the dynamic energy consumption weight function of link is defined;Establish the three-dimensional energy consumption and thermal interference evaluation model of node, introduce space thermal interference factor;Analysis the characteristics of task to be scheduled;Based on energy consumption potential field screening candidate node;Build the global energy consumption cost function including computing energy consumption, network transmission energy consumption and thermal coupling penalty energy consumption;According to the principle of minimum total energy consumption, select target node to execute scheduling and update topology state.The application solves the problem of splitting computing and network energy consumption, ignoring the space coupling effect of heat dissipation in the prior art by the fusion of network topology perception and multi-dimensional energy consumption model, and realizes the optimization of global energy consumption of data center.

Task-driven defogging perception method and system

The application provides a task-driven defogging perception method and system, and belongs to the technical field of cooperative perception of Internet of Vehicles. The method comprises the following steps: collecting a first foggy image stream and a second foggy image stream; obtaining a first processed image by intercepting and processing the first foggy image stream; obtaining a second processed image by intercepting and processing the second foggy image stream; respectively performing geometric alignment on the first processed image and the second processed image to obtain a first aligned image and a second aligned image; performing defogging and image enhancement processing on the first aligned image and extracting a first feature; performing defogging and image enhancement processing on the second aligned image and extracting a second feature; respectively repairing the first feature and the second feature by using a diffusion model; fusing the repaired first feature and the second feature, and determining a final perception result based on the fused feature. The application improves the robustness of automatic driving and balances the low power consumption of AI intelligent glasses and the high performance computing demand of a vehicle-mounted system.
Owner:NORTHEASTERN UNIV CHINA

An accelerator control system based on a RISC-V processor

ActiveCN224553771ULoad instructionPerformance computing
The utility model relates to an accelerator control system based on RISC V processor, include: RISC V processor, including vector extension interface and scalar extension interface, can load instruction and decode into vector extension instruction and / or scalar extension instruction after sending to multiple accelerator groups, multiple accelerator groups are used for executing the instruction, instruction distribution unit is connected with the RISC V processor, is used for distributing the instruction and operand of the RISC V processor to each accelerator group, and group selector is connected with the instruction distribution unit and each accelerator group respectively to select specific accelerator group execution instruction. Through the setting of vector extension interface and scalar extension interface, the instruction and operand can be efficiently distributed to multiple acceleration cores, so that the RISC V processor can manage multiple accelerator cores, and thus the demand of high-performance computing can be met.
Owner:WUXI MICRONANO CORE ELECTRONIC TECH CO LTD

A meteorological data parallel loading method and system based on sliding window data locality

This invention discloses a method and system for parallel loading of meteorological data based on sliding window data locality, relating to the fields of artificial intelligence and high-performance computing. The method includes: allocating continuous data file segments to each parallel process based on a global data file path list, sliding window length, the number of time steps contained in each data file, and the total number of parallel processes; determining the data file path where the sample is located and its time step offset within the file; constructing an in-memory cache pool based on an ordered dictionary for each parallel process, setting an upper limit for the cache pool size; reading the time step data from the corresponding data file in the cache pool according to the sliding window length, concatenating them to form the sequence data required for model training, and outputting it to a computing device for model training. This invention achieves efficient and scalable parallel data loading by constructing a data locality-aware caching mechanism and an adaptive multi-process partitioning strategy.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

A hardware acceleration system and method supporting matrix and tensor computation

PendingCN122433821AComputer hardwareData stream
The application discloses a kind of hardware acceleration system and method supporting matrix and tensor calculation, belong to high-performance computing and artificial intelligence chip design field.The system is based on the computing core array of annular interconnection, integrated uniform scheduling and variable unit USTU and configurable data flow controller CDC.USTU is responsible for the real-time analysis conversion of high-level tensor operation to optimized matrix multiplication GEMM subtask and completes uniform fine-grained scheduling;CDC dynamically configures the data flow mode of computing core array, storage access mode and inter-core communication logic according to the characteristics of subtask.The application realizes the adaptive mapping of neural network operator to bottom hardware data flow, can efficiently execute various tensor calculations, significantly improves hardware utilization, energy efficiency ratio and application flexibility.
Owner:SHANGHAI QUSU CHAOWEI TECHNOLOGY CO LTD

Causal ordering with multicast communication for scalable decentralized systems

Approaches are described for a decentralized computing platform that enables dynamic execution of distributed services through a hybrid on-chain and off-chain architecture. The platform consists of a network of nodes, each running a virtual machine to host and execute service instances. A sequencer component records, sequences, and notarizes function calls to maintain a consistent service-level state across the network. Computation takes place off-chain on nodes that participate dynamically in specific services based on resource availability. A persistent virtual memory structure manages both service-level and instance-level state. The service-level state is synchronized across nodes participating in the same service, while the instance-level state is isolated to individual nodes for independent operations. The memory system is non-volatile, retaining state across function calls and enabling direct execution from memory. A universal procedure call (UPC) framework facilitates decentralized communication through multicast-enabled operations with customizable fault tolerance and supports applications from finance to high-performance computing.
Owner:LYQUOR LABS INC

A cross-node heterogeneous memory pooling method and system based on a CXL protocol

The application discloses a cross-node heterogeneous memory pooling method and system based on a CXL protocol, and belongs to the technical field of server architecture and memory resource management. The method builds a communication link between a computing node and a heterogeneous memory resource pool through a CXL exchange unit, deploys an FPGA-accelerated intelligent memory controller and optimizes an access path; divides the storage functions of DRAM, CXL memory and HBM, establishes a unified management mechanism to realize resource centralized scheduling; relies on a machine learning model to distinguish different access frequency data and migrate to the corresponding memory, dynamically adjusts the memory bandwidth and capacity; constructs a key data redundancy image, expands an atomic operation protocol to guarantee concurrent access consistency, and establishes a node fault monitoring and recovery mechanism. The system adapts to the above method and integrates hardware and software modules to work cooperatively. The scheme realizes efficient sharing and dynamic management of cross-node memory resources, improves resource utilization and access efficiency, enhances system stability and expandability, and meets the memory requirements of high-performance computing scenarios.
Owner:四川华鲲振宇智能科技有限责任公司

A robust high resolution beamforming device with narrow mainlobe and low sidelobe effect

ActiveCN224328235UWave based measurement systemsHigh resolution imagingSoftware system
The utility model discloses a kind of narrow main lobe and low side lobe effect's robust high-resolution beamforming device, including host computer rack, the top of host computer rack is fixedly installed with host computer body, the side of host computer body is equipped with processing equipment, the top of mounting plate is fixedly installed with Jetson Xavier NX development board kit, a kind of narrow main lobe and low side lobe effect's robust high-resolution beamforming device of the utility model, by setting Jetson Xavier NX development board kit, to improve active acoustic imaging resolution capability, on the basis of establishing active detection space-time two-dimensional convolution model, joint azimuth and space two-dimensional deconvolution processing means, realize in angle and time two dimensions on the resolution capability and interference suppression capability of active detection sonar to target are improved;By setting host computer body and processing equipment, array detection software system platform based on CPU and GPU collaborative architecture is constructed, to meet the demand of high-performance computing of complex algorithm for high-resolution imaging sonar beamforming.
Owner:HARBIN ENG UNIV

Deformable convolution acceleration method, electronic device, storage medium and program product

This invention relates to the field of artificial intelligence technology, providing a deformable convolution acceleration method, electronic device, storage medium, and program product, comprising: acquiring an input feature map, an offset tensor, and an original weight tensor; performing interpolation sampling on the input feature map based on the offset tensor, concatenating and unfolding sampling results belonging to the same spatial location along the channel dimension to generate a continuously stored intermediate tensor; performing shape transformation on the original weight tensor to obtain a target weight tensor with a spatial dimension of one-to-one and a channel dimension matching the intermediate tensor; and performing a one-to-one convolution operation based on the intermediate tensor and the target weight tensor to obtain an output feature map. This invention, through the collaborative design of continuous data arrangement and weight shape rearrangement, transforms discrete data computation into an efficient one-to-one convolution operation that supports continuous data input, thereby fully utilizing high-performance computing units and improving the utilization rate of hardware computing resources and operator execution performance.
Owner:SHANGHAI BIREN TECH CO LTD

A single-computer multi-CXL card-based pooled RAMDisk system

PendingCN122365536AResource poolAlgorithm
This invention discloses a pooled RAMDisk system based on a single machine with multiple CXL cards. By reconstructing the underlying architecture, this invention integrates parallel expansion of multiple CXL cards on a single machine, and deep fusion of local main memory and non-volatile storage into a unified pooled RAMDisk resource pool, completely breaking the limitations of existing technologies that focus on "single feature optimization." Simultaneously, the system leverages end-to-end CXL direct connection channels and security management mechanisms to ensure both nanosecond-level ultra-low latency and multi-node sharing while maintaining security. Furthermore, it utilizes underlying persistence and multi-card redundancy to resolve the contradiction between high-speed access and data volatility, and, in conjunction with dynamic collaborative scheduling, eliminates performance friction and resource waste in multi-card environments. Ultimately, it achieves a perfect synergy of five characteristics: high speed, large capacity, non-volatility, sharing, and high resource utilization, comprehensively filling the storage technology gap in high-performance computing scenarios. Therefore, this invention is highly suitable for large-scale application and promotion.
Owner:SHENZHEN JINGTIE COMPUTING SYSTEM TECHNOLOGY CO LTD

A high-performance computing cluster resource scheduling method and system based on TR-DQN

The application discloses a kind of high-performance computing cluster resource scheduling method and system based on TR-DQN, first user submits task request, all requests enter waiting queue and wait for scheduling;Then the priority of submitting task is calculated, and the waiting queue is reordered;Then the node information and task information of cluster are collected and processed, and the processed data is input into TR-DQN model for scheduling;Finally, after task scheduling is completed, it enters corresponding node operation.TR-DQN model combines the characteristics of high-performance computing cluster scheduling into deep reinforcement learning, and introduces two-level neural network structure, the first neural network is used to select tasks for immediate execution or reserved execution, and the second neural network is used to select tasks for backfilling, which can improve the resource utilization of the cluster, reduce the waiting time of the task, and quickly adapt to changes in the cluster load environment. In addition, it can also minimize the problem of work starvation in the cluster.
Owner:WUHAN UNIV

A reinforcement learning fault-tolerant training system oriented towards human feedback

This invention belongs to the field of artificial intelligence technology, specifically a reinforcement learning fault-tolerant training system oriented towards human feedback. The system comprises a decoupled recovery engine, a mirror checkpoint engine, and a generation phase checkpoint engine; it supports component-level hot recovery to shorten restart latency after a failure; it utilizes data redundancy between the training and inference engines based on the RLHF paradigm to achieve low-overhead mirror checkpointing; and it sets semantic checkpoints in the generation phase to limit recomputation to sub-iteration ranges, thereby reducing recovery overhead and improving resource utilization. Experimental results show that this invention can reduce interruption time caused by cluster failures, training failures, and suspensions in reinforcement learning jobs based on human feedback, improving cluster training efficiency, and is suitable for large-scale RLHF training under high-performance computing resources.
Owner:FUDAN UNIVERSITY

METHOD AND SYSTEM FOR FACILITATING CABIN MONITORING IN A VEHICLE

This document discloses a method and a system for facilitating cabin monitoring in a vehicle (200). The method comprises receiving a first feed containing spike-train data of one or more events (219) from an event-based camera (103) and recognizing the one or more events (219) as critical or non-critical by analyzing the spike-train data. After recognizing the one or more events (219) as critical, the method generates a wake-up signal to activate a frame-based camera (105) of a high-performance computing (HPC) chip (107) to receive a second feed and transforms one or more spike-train features corresponding to the spike-train data. The second feed comprises one or more image features (119).Finally, the procedure includes performing a fusion of the one or more transformed spike-train features with the one or more image features (119) from the second feed to verify the detection of the critical event.
Owner:MERCEDES BENZ GROUP AG

Efficient inference method for large neural network model based on multi-gpgpu

The application belongs to the technical field of artificial intelligence and high-performance computing, and specifically relates to a neural network large model efficient inference method based on multiple GPGPUs. The method aims to solve the problems of large communication overhead between multiple processors, uneven load, low resource utilization, and high data transmission delay. Through static analysis and mixed-granularity partitioning of the model computation graph, the computation task is divided into multiple subgraphs; combined with heterogeneous resource perception and dynamic mapping strategy, the subgraphs are allocated to the optimal GPGPU based on a weighted cost function; a global pipeline scheduling plan is constructed using communication topology perception to maximize computation and communication overlap; data is loaded in advance through the host-side hierarchical cache and asynchronous prefetch mechanism to hide transmission delay; and multi-stream concurrent execution and event-based lightweight synchronization are adopted on each GPGPU to reduce the waiting overhead. The application can significantly reduce inference delay, improve throughput and hardware utilization, and has good adaptability and scalability.
Owner:BEIJING TOPMOO TECH

Method and apparatus for accelerating exact probability simulation in combinational equivalence checking

This invention discloses an accelerated method and apparatus for precise probabilistic simulation in combinatorial equivalence checking, relating to the fields of electronic design automation (EDA) and high-performance computing. The method includes: parsing the input circuit combinational logic netlist to form an EPS operator graph and corresponding data distribution configuration, transmitting this data to a GPU; starting the GPU to distribute the EPS operator graph to a processing unit array for EPS calculation based on the data distribution configuration using a data flow-driven approach; and collecting the running status information of each round of calculation and feeding it back to the CPU for dynamic optimization of the next round of calculation. This method can transform circuit calculation tasks into data flow-driven EPS operator graphs and execute them on the GPU, efficiently utilizing the GPU's massively parallel computing capabilities to overcome performance bottlenecks. The pipelined task advancement based on superblocks and the dynamic optimization mechanism based on performance reports can adapt to the computational characteristics of different circuit structures, balance the load, optimize resource allocation, and significantly shorten the verification cycle of combinatorial equivalence checking.
Owner:BEIJING NEW ENERGY VEHICLE TECH INNOVATION CENT CO LTD

A joint alignment and quantification method for large-scale mass spectrometry cohort

The application discloses a large-scale mass spectrometry queue-oriented joint alignment and quantification method, relates to the cross field of bioinformatics, artificial intelligence and high-performance computing, and outputs a quantitative feature matrix through an end-to-end deep learning network; wherein, in a unified back propagation process, the end-to-end deep learning network synchronously optimizes the retention time alignment, the cross-batch intensity drift correction and the quantitative regression based on a joint loss function, the application regards the retention time alignment, the cross-batch intensity drift correction and the quantitative regression as a whole joint operation instead of linearly performing the operations one by one, and overcomes the defects of the step-by-step strategy in the prior art.
Owner:SHANGHAI DEV CENT OF COMP SOFTWARE TECH

A communication efficient memory fusion method and system for super-memory graph neural network training and a medium

The application belongs to the technical field of artificial intelligence, graph data processing and high-performance computing, and discloses a communication-efficient memory fusion method and system for super-memory graph neural network training and a medium. A task-specific memory management scheme, zero-copy CPU-GPU transmission, a layout-aware GPU intercommunication pipeline and an out-neighbor gradient aggregation mechanism are used to significantly reduce the overall data communication overhead while ensuring training accuracy. In the fusion memory-aware graph partitioning, a same subgraph is maintained by the storage layout of multiple GPUs. For P task partitioning characteristics and T task partitioning vertex dimensions, the number of graph partitions required by the system is reduced, thereby reducing neighbor replication, and the overall CPU-GPU communication overhead of the system is successfully reduced by 50% to 64%. The back propagation of the out-neighbor gradient aggregation of the application changes the back propagation into an out-neighbor aggregation process in the preprocessing stage.
Owner:NORTHEASTERN UNIV CHINA

A system and method for evaluating colorectal cancer screening strategies based on high-performance computing.

This invention discloses a high-performance computing-based system and method for evaluating colorectal cancer screening strategies, belonging to the field of medical information processing technology. It includes: a dual-pathway disease model module supporting both adenoma-carcinoma pathways and serrated adenoma pathways, modeled using age-dependent normal distribution progression time; an individualized risk scoring module integrating risk factors such as family history, inflammatory bowel disease, obesity, and diabetes; a screening strategy simulation module supporting combined screening strategies using different tools, employing compliance parameters specific to the Chinese urban population; a high-performance computing module using Numba JIT compilation technology for acceleration, multi-process parallel processing for further acceleration, and data structure optimization such as integer age to reduce memory usage; and a health economics evaluation module calculating indicators such as QALY, LYG, and ICER. This invention achieves significant performance improvements while maintaining model accuracy, greatly reducing simulation time for large-scale populations.
Owner:SHENZHEN NANSHAN DISTRICT CHRONIC DISEASE CONTROL CENT (SHENZHEN NANSHAN DISTRICT MENTAL HEALTH CENT)

Code data conversion method, system and electronic device

PendingCN122285015AData transformationLogisim
This application discloses a code data conversion method, system, and electronic device, relating to the field of high-performance computing technology. The code data conversion method of this application uses a pre-trained model fine-tuned based on different training sets to convert code data of different forms, improving the accuracy of code conversion. Furthermore, this application provides code correctness verification logic; in the event of verification failure, a conversion rollback operation is added, enabling automatic execution of code conversion and improving efficiency. Therefore, this application can solve the technical problems of low conversion accuracy and inefficiency, achieving the technical effect of improving both the efficiency and accuracy of code conversion.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

A method and system for dividing GPU space resources based on persistent threads

The application discloses a GPU space resource partitioning method and system based on a persistent thread, and belongs to the field of high-performance computing and computer system structure.The method extracts original PTX code from a binary module by dynamically intercepting a CUDA interface at a driver layer; a static compiling conversion technology is used to recombine an instruction stream, construct a persistent thread structure, and innovatively introduce a "kernel fusion" mechanism in the persistent thread; through instruction body multiplication and virtual index remapping, the static demand of a single physical thread block for registers and shared memory is artificially doubled, so that the single physical thread block reaches the upper limit of hardware resources of a streaming multiprocessor; finally, a kernel starting interface is intercepted, and the original kernel is transparently replaced by the converted kernel.The application effectively solves the defect that traditional persistent threads cannot realize strict physical exclusive on a streaming multiprocessor due to unsaturated resource occupation, and realizes fine-grained and deterministic partitioning and isolation of GPU space resources without modifying user source code.
Owner:XI AN JIAOTONG UNIV

A method and apparatus for selecting recomputation timing during large model training

This invention discloses a method and apparatus for selecting recomputation timing during large model training, belonging to the field of distributed training optimization technology for large models at the intersection of artificial intelligence and high-performance computing. The method includes: dynamically predicting the impact of pipeline bubbles and memory usage by real-time monitoring of computation, communication tasks, and memory usage during training iterations, using a pre-trained machine learning model, and determining the timing for inserting recomputation operations. This method takes large model training state data as input, aims to minimize iteration latency, and intelligently outputs the optimal recomputation trigger timing. The apparatus includes state monitoring, intelligent decision-making, feature engineering processing, and execution modules to implement this method. This invention upgrades static rules to intelligent decision-making, enabling more precise utilization of computation and communication gaps, adaptively balancing computation, communication, and memory resources, thereby significantly improving the throughput and efficiency of large model training.
Owner:ZHEJIANG LAB

Energy efficiency optimization oriented intelligent computing center computing and resource collaborative scheduling method, system, device and medium

The application discloses a method and system for computing and resource collaborative scheduling of a wisdom calculation center facing energy efficiency optimization, belongs to the technical field of data centers and high-performance computing, and comprises the following steps: establishing a multi-level resource and energy efficiency perception model of the wisdom calculation center to acquire real-time perception data; constructing a full-stack energy efficiency evaluation and prediction model based on the real-time perception data; solving an optimal collaborative scheduling strategy through a joint optimization engine based on the full-stack energy efficiency evaluation and prediction model, and outputting a prediction value according to the collaborative scheduling strategy; executing the collaborative scheduling strategy, monitoring actual energy consumption and task performance data after execution of the collaborative scheduling strategy, taking the difference between the monitoring result and the prediction value as a feedback signal, and inputting the feedback signal into the full-stack energy efficiency evaluation and prediction model to realize online self-adaptation and closed-loop optimization of the model. The application realizes optimization of the overall energy efficiency ratio of the wisdom calculation center by constructing a unified energy efficiency model and intelligent collaborative decision-making.
Owner:LIAOYANG POWER SUPPLY COMPANY OF STATE GRID LIAONING ELECTRIC POWER SUPPLY +2

System and method for using infiniband routing algorithms for ethernet fabrics in a high performance computing environment

Systems and methods for using InfiniBand routing algorithms for Ethernet fabrics in a high performance computing environment. The method can provide, at a computer comprising one or more microprocessors, a plurality of switches, a plurality of hosts, a topology provider (TP) module, a routing engine (RE) module, and a switch initializer (SI) module. The method can perform a discovery sweep, by the TP, of the plurality of hosts and the plurality of switches and assigns an address to each of the plurality of hosts and the plurality of switches. The method can calculate, by the routing engine, a routing map, based upon a routing scheme, for the plurality of hosts and the plurality of switches, the routing map comprising a plurality of forwarding tables. The method can configure, each of the plurality of switches with a forwarding table of the plurality of forwarding tables calculated by the routing engine.
Owner:ORACLE INT CORP

Control system and method for hybrid HPC-QPU orchestration

PCT designated stageWO2026150286A1Performance computingOrchestration (computing)
Disclosed are a computer-implemented method and a computer-implemented system for orchestrating hybrid quantum-classical execution between high-performance computing (HPC) infrastructure and one or more quantum processors (QPU). An orchestration module breaks down a work flow into subtasks and, for each subtask, compiles multiresource HPC and QPU processing telemetrics, normalises same to a canonical vector and evaluates proposed plans under hard budget constraints (time, executions, cost and / or depth). The system automatically selects an execution resource (HPC or QPU) and, when executed in an QPU, applies tiered error mitigation and verifies the result using an acceptance gate. In the event of rejection, the system executes a deterministic recovery policy with retry, execution resource switching and / or backup routing to HPC. An auditable record of decisions and metrics is generated for traceability and feedback.
Owner:CARRANZA VILLALOBOS CARLOS MIGUEL

Single integrated substrate package for high performance computing

PCT designated stageWO2026142391A1Memory chipComputer architecture
The present invention relates to a single integrated substrate package for high performance computing (HPC), in which a plurality of integrated circuits (ICs) such as a memory, a signal processing device, and an optical engine chip can be packaged as one chip. The single integrated substrate package is a package on which a memory chip, a logic semiconductor integrated circuit chip, and an optical engine are mounted, the single integrated substrate package comprising: an electrical-routing component for inputting and outputting an electrical signal for the package; an optical-routing component for processing input and output of an optical signal for the optical engine; a mold body surrounding the electrical-routing component and the optical-routing component; and a first redistribution layer having a wiring line for performing interconnection between mounted components, wherein the optical engine includes a photonic integrated circuit (IC) and an electronic integrated circuit (IC) molded inside the mold body.
Owner:LIPAC CO LTD

Utilization of high performance computing for artifical intelligence driven vehicle customization

An artificial intelligence (AI) driven customization system for a vehicle includes an input device configured to receive a customization request from a user, the customization request relating to a feature of the vehicle and a high performance computing (HPC) controller configured to access a large language model (LLM), analyze the customization request using the LLM, based on the analysis of the customization request, obtain a software code for the customization request, and execute the obtained software code to preview a customization of the vehicle feature to the user.
Owner:FCA US LLC

Safety unstacking visual physical reasoning method and system based on segmentation guided transformer

PendingCN122366658AAlgorithmRobotic arm
This invention discloses a visual-physical reasoning method and system for safe depalletizing based on segmentation-guided Transformer. The method constructs a 4-channel input using RGB images and target masks, extracts visual features using an improved ResNet-50, and feeds them into a dual-branch Transformer inference network that includes global stability prediction and affected area mask prediction. This enables stacking safety prediction and pixel-level localization of unstable areas after removing candidate boxes. Simulated counterfactual annotation and a composite loss function are used for joint training to improve the model's generalization ability and robustness of physical reasoning. The system includes an RGB-D camera, a six-DOF robotic arm, an end effector suction cup, and a high-performance computing platform, enabling real-time instance segmentation, physical consequence reasoning, and safe grasping decisions. This invention upgrades traditional passive geometric perception to active physical causal reasoning, effectively preventing depalletizing collapse and ensuring the safety of automated operations. It is suitable for unstructured and chaotic stacking and depalletizing scenarios such as warehousing and logistics.
Owner:ZHIYOUWUJIE (SHENZHEN) INTELLIGENT TECHNOLOGY CO LTD

A resource pool-based nccl migration method

PendingCN122332113AResource poolData pack
The application discloses an NCCL migration method based on a resource pool, relates to the technical field of high-performance computing communication optimization, and comprises the following steps: obtaining thread identification records, synchronization timestamp markers and thread state information from the resource pool by scanning the multithread running environment of the NCCL communication program based on the resource pool at a source end; generating a thread state snapshot in combination with data packet transmission progress, constructing a data transmission state file containing transmission interruption recovery points and synchronization signal priorities, and determining an initialized synchronization marker data set; transmitting the thread state snapshot and a data check code value to a target end according to the initialized synchronization marker data set, decoding the synchronization marker by using an inheritance marker analysis rule, generating a control module state cache of the target end, and obtaining a preliminary synchronization marker inheritance result; and the NCCL migration method based on the resource pool significantly improves the migration efficiency, data link continuity and overall running stability of the NCCL communication program in a high-load scene.
Owner:VIRTAI TECH BEIJING CO LTD

A vehicle dynamics model incremental updating method based on vehicle-cloud cooperation

The application belongs to the intelligent networked vehicle control and vehicle networking technical field, and discloses a vehicle dynamics model incremental updating method based on vehicle-cloud cooperation, comprising the following steps: S1. constructing a vehicle-cloud cooperation closed-loop updating framework based on an information-physical fusion system, wherein the vehicle-cloud cooperation closed-loop updating framework comprises a vehicle-side edge computing layer and a cloud-side high-performance computing layer; S2. in the vehicle-side edge computing layer, establishing a high-value difficult example sample intelligent screening mechanism based on "surprise degree"; S3. in the cloud-side high-performance computing layer, establishing an incremental learning model updating mechanism based on mixed experience playback and physical constraints; S4. establishing a "shadow mode" model verification and parameter distribution process. The method can realize the full life cycle high-precision maintenance of the dynamics model by establishing a bidirectional closed-loop link connecting the vehicle-side entity and the cloud-side digital twin, can effectively overcome the catastrophic forgetting defect of deep learning, and can significantly reduce the prediction error of the new working condition on the premise of ensuring the high precision of the original working condition.
Owner:CHONGQING UNIV