Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

72 results about "Extreme scale computing" patented technology

The purpose of the Joint Laboratory for Extreme Scale Computing (JLESC) is to be an international, virtual organization whose goal is to enhance the ability of member organizations and investigators to make the bridge between Petascale and Extreme computing. The founding partners of the JLESC are INRIA and UIUC.

Electromagnetic particle adaptive simulation method and system for non-uniform and conformal grids

The invention discloses a non-uniform and conformal grid-oriented electromagnetic particle adaptive simulation method and system, and relates to the technical field of electromagnetic particle simulation. The method comprises the following steps: performing non-uniform right-angle mesh generation based on a charged particle average free path, and identifying conformal mesh units; carrying out initialization by adopting macro particles; executing electromagnetic field propulsion, field interpolation, particle propulsion, charge current backfilling and collision simulation in each time step; and accelerated calculation is realized through a CPU and GPU hybrid parallel architecture. Through the method, the numerical simulation precision can be improved, the adaptability of complex geometry is enhanced, the coupling stability of the electromagnetic field and the particles is ensured, the information interaction between the field and the particles is improved, and the large-scale calculation efficiency is improved.
Owner:深圳十沣科技有限公司

GPUBox hardware decoupling system based on Retimer card and PCIeSwitch chip

The invention discloses a GPU Box hardware decoupling system based on a Retimer card and a PCIe Switch chip, and belongs to the technical field of computer hardware architecture and high-speed interconnection. According to the system, a Retimer card and a PCIe Switch chip are integrated in an independent GPU Box, and a decoupling link of a CPU server and a GPU acceleration card is constructed; the Retimer card realizes 30-meter long-distance PCIe signal transmission and breaks through physical distance limitation; the PCIe Switch chip pools GPU resources through a dynamic routing and MRIOV technology, supports flexible allocation of computing power by multiple servers, and realizes Peer-to-Peer direct connection communication between GPUs. Aiming at a large model reasoning scene, the system optimizes KV cache bandwidth allocation and video memory and memory cooperative scheduling, so that the 100B parameter model reasoning throughput is greatly improved; and meanwhile, the usability of the system is greatly improved through fault isolation and hot plug design. According to the method, the problems of physical binding of the CPU and the GPU, limited transmission distance, rigid resource allocation and the like in a traditional architecture are solved, and the method is suitable for large-scale AI calculation and distributed GPU cluster deployment.
Owner:HEFEI FENGZHIYI SEMICON CO LTD

Method for protecting parameters of a network model based on multiple keys

The application discloses a network model parameter protection method based on a multi-layer key, relates to the technical field of data protection, and calculates a low-entropy page upper limit at a single moment based on device parameters and a maximum layer scale of a model, and determines a page budget balance in combination with a remaining count of a previous round; in response to a hierarchical decryption request, if the balance is sufficient, a certain amount is deducted and a one-time session key is derived in combination with a device fingerprint; then, the assigned physical page is subjected to no-information-residue verification, after the verification, the parameter ciphertext is decrypted using the session key and written, and the current low-entropy page count is increased in real time; after the page is used up, forced overwriting and clearing verification are performed, after the verification, the page budget balance is backfilled and the low-entropy page count is reduced, and the one-time session key is destroyed; finally, a digital signature is generated and stored as evidence. The application establishes a budget gating and recycling mechanism, realizes dynamic limitation of the amount of plaintext in the runtime memory, and solves the exposure risk problem caused by full-decryption of model parameters of an edge device.
Owner:BEIJING DAOPOWER

Fault suppression type high-reliability CNN quantitative perception training method and system

The invention discloses a fault suppression type high-reliability CNN quantitative perception training method and system, and the method comprises the steps: calculating the occurrence probability of an expected single event upset event according to the type and scale of a CNN network layer; according to the expected single event upset event occurrence probability, random fault generation is carried out in the quantitative perception training simulation integer inference process; extracting a model fault propagation path, analyzing a fault propagation influence mechanism, performing a model fault-tolerant capability node test, finding out a sensitive node of which the model precision oscillation exceeds a preset threshold value, and performing targeted fault generation and training on the sensitive node; according to the model structure, layer-by-layer fault iteration training is carried out, convergence results are obtained, the training result of each convergence is used as a pre-training parameter of next simulation fault training until no sensitive node appears, and the quantification process of the CNN is completed. According to the method, the reliability and robustness of the model in a high-radiation environment are remarkably improved.
Owner:GUANGZHOU UNIVERSITY

Server heat dissipation device and method, electronic equipment and storage medium

The invention discloses a server heat dissipation device and method, electronic equipment and a storage medium, and relates to the field of financial science and technology or other related technical fields, and the device comprises a cold plate which is internally provided with a micro-channel flow channel, and liquid cooling heat dissipation is carried out through the micro-channel flow channel; the fin assemblies are distributed in the edge area of the cold plate or the non-flow-channel covering area in the cold plate and used for conducting air cooling heat dissipation; the standby fan group is used for performing heat exchange by utilizing the surface of the fin assembly in an air cooling mode; and the intelligent control system comprises a flow sensor and a temperature sensor, and is used for monitoring the inlet flow of the cold plate and the temperature of the heating element to obtain a monitoring result, and performing intelligent heat dissipation adjustment based on the monitoring result in combination with at least one of the cold plate, the fin assembly and the standby fan group. According to the method and the device, the technical problems that server heat dissipation depends on single liquid cooling heat dissipation, the heat dissipation effect is poor and the emergency capacity is poor in a large-scale calculation scene in the prior art are solved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Distributed computing and NPU scheduling system based on underground coal mine embedded equipment

The invention discloses a distributed computing and NPU scheduling system based on underground coal mine embedded devices. The distributed computing and NPU scheduling system comprises an AI scheduling server host which is used for monitoring the computing state of each embedded device in a computing cluster in real time, networking each NPU module arranged on the embedded device underground and scheduling the computing task of the embedded device; the computing cluster is used for uploading a computing state to the AI scheduling server host and receiving task allocation sent by the AI scheduling server host; and the NPU module group is used for completing an AI reasoning task. According to the distributed computing and NPU scheduling system based on the underground coal mine embedded devices, a distributed computing architecture and an autonomous scheduling mechanism are adopted, so that the computing power of each embedded device can be fully utilized, and the superposition of the overall computing power of the system is realized in the face of large-scale computing tasks.
Owner:TIANJIN HUANING ELECTRONICS

Calculation method and application of underwater random bubble curtain pressure value simulation model

The invention discloses a calculation method and application of an underwater random bubble curtain pressure value simulation model. The method comprises the following steps: firstly, establishing a bubble generation model based on random characteristics, and realizing random distribution of bubble positions and sizes through an MATLAB program and a randperm function; then, a fluid and bubble dynamics bidirectional coupling model is constructed, and pressure field association is established by keeping grid sizes consistent; generating a grid by adopting an SALE method, and combining the bubble model, the target plate model and the fluid model through open and import commands; and finally, carrying out numerical calculation by adopting an LS-DYNA solver, iteratively adjusting parameters based on an optimization algorithm, and outputting a pressure value simulation result. According to the method, full-process automatic modeling is realized, the efficiency is improved by more than 50%, million unit-level large-scale calculation is supported, the model precision is high, the universality is high, and the method can be widely applied to simulation analysis and device optimization of underwater noise reduction, antifouling and oxygenation scenes.
Owner:JIANGNAN IND GRP CO LTD

A large-scale computing interactive all-in-one machine for multi-modal data

The utility model discloses a large -scale calculation interactive integrated machine of multimode data, including base, the top front end fixedly connected with display case of base, the rear end fixedly connected with a plurality of groups of wire conduit of display case, the rear end fixedly connected with terminal block of wire conduit, the top fixedly connected with terminal box of terminal block, be equipped with a plurality of groups of wire storage cavity that go up and down and pass through in terminal box, the top still fixedly connected with a plurality of groups of terminal post of terminal block, and terminal post all are arranged in the inside of wire storage cavity, and the outer wall top of terminal post all fixedly connected with spring wire, and the outer wall of spring wire all fixedly connected with protective sheath, and the top of spring wire all fixedly connected with joint, through setting the weight block, when it slides down, can drag spring wire to wire storage cavity inside, not only save staff's time, improve the use efficiency of whole integrated machine simultaneously, through setting the protective sheath, can protect spring wire to prevent spring wire breakage.
Owner:TIANJIN BENLAI TECHNOLOGY DEVELOPMENT CO LTD

Computing apparatus, memory module group, computing device and computing device cluster

The present application relates to the technical field of chip packaging, and provides a computing apparatus, a memory module group, a computing device, and a computing device cluster. The computing apparatus comprises at least one memory module group, a processor, and a mainboard. Each memory module in the memory module group is formed by stacking memory chips having a bit width of at least 8 bits, the size is small, the stacking process is simple, each memory module is accessed through at least one memory channel, and the density of the memory channels in unit area is large, so that the memory module group in the computing apparatus can provide large bandwidth to meet the requirements of AI or large-scale computing scenarios on bandwidth, without adopting a HBM having high process difficulty and high costs, thereby reducing engineering difficulty and saving costs. In addition, the processor and the memory modules communicate on the basis of a DRAM interface, and compared with an HBM interface, there is less delay, resulting in lower power consumption generated by the delay, thereby improving the data read-write efficiency, and significantly improving the random access efficiency.
Owner:HUAWEI TECH CO LTD

Hardware-friendly Transform column balance pruning model compression and efficient deployment method

The invention discloses a hardware-friendly Transform column balance pruning model compression and efficient deployment method. A model compression algorithm, a lightweight parameter storage format, an operation data buffer, a systolic array operation block, a vector operation unit, a nonlinear operator unit, a data flow controller and a DMA unit are included. A model compression algorithm and an efficient deployment architecture are explored according to Transform network software and hardware collaborative reasoning requirements: in a software level, the scale calculation complexity of model parameter quantities is reduced through a fine-grained column balance structured pruning strategy, and parameters are stored in a single-instruction multi-data-stream format and the parameter storage efficiency is optimized through mask code storage; according to the hardware level, an edge computing-oriented Transform special accelerator architecture is designed, so that the architecture can support column balance structured pruning characteristics and a lightweight parameter storage scheme in an original manner. According to the Transform model compression and efficient deployment method, the parameter sparsity after structured pruning is fully utilized, so that the parameter storage pressure of a hardware architecture is reduced, the complex balance of an arithmetic unit is ensured, the operation efficiency of an accelerator is improved, and load balance and efficient reasoning during software and hardware collaborative optimization are realized; the method is widely applicable to efficient deployment scenes of Transform models for edge calculation.
Owner:BEIJING UNIV OF TECH

Method and system for automatically selecting monitoring indexes of large-scale computer system

The invention discloses an automatic selection method and system for monitoring indexes of a large-scale computer system, and the automatic selection method for out-of-band monitoring indexes comprises the following steps: aiming at a time sequence data sequence of M in-band monitoring indexes and N out-of-band monitoring indexes from the same computer node, quantifying the dynamic association strength between the out-of-band monitoring indexes and the in-band monitoring indexes, and selecting a plurality of out-of-band monitoring indexes with relatively high dynamic association strength with the in-band monitoring indexes to obtain a key out-of-band monitoring index set; and for the key out-of-band monitoring index set, carrying out space-time correlation analysis of in-band index anomaly and out-of-band hardware alarm by extracting in-band anomaly, and further optimizing the key out-of-band monitoring index to obtain a finally selected out-of-band index. The invention aims to select an out-of-band monitoring index with strong system operation indication in an automatic and unsupervised manner by associating the in-band operation state of the system, thereby reducing the problems of large cardinal number of an acquisition sensor, large data transmission pressure and large processing delay.
Owner:NAT UNIV OF DEFENSE TECH

Method and apparatus for sequential continuity testing and assembly of wafer-scale silicon circuit boards

PCT designated stageWO2026139941A1Probe cardWafering
A manufacturing and test methodology for Zetta-scale computing engines is disclosed. The system employs a reticle-sized MEMS probe card to sequentially verify the electrical continuity of passive interconnects on a Wafer-Scale Silicon Circuit Board (WSSCB) prior to component assembly. By stepping across the wafer and testing one module at a time, the apparatus generates a defect map of the substrate's high-density wiring without the need for full-wafer probing or active powering. Following verification, Known Good Die (KGD) stacks—comprising Logic and HBM—are attached to the valid sites of the WSSCB via micro-bonding. This "step-and- repeat" verification strategy ensures high manufacturing yield for the complex all-silicon domain assembly.
Owner:SILVEBROOK KIA

A runtime hang detection and localization method and apparatus for parallel programs

The application provides a runtime suspension detection and positioning method and device for parallel programs, and applies to the technical field of large-scale computing. The method comprises the following steps: constructing a program counter value queue of each target execution unit at each sampling time; recording a stagnation program counter value of a target execution unit in a stagnation state; clustering the stagnation program counter value of the target execution unit to determine a suspended target execution unit; based on a machine instruction corresponding to the stagnation program counter value of the suspended target execution unit, obtaining source file and line number information corresponding to the stagnation program counter value through debugging information; and when the machine code corresponding to the stagnation program counter value of the suspended target execution unit is a call instruction, obtaining function call information. In this way, non-intrusive and low-load monitoring of parallel programs is realized, normal synchronization waiting and abnormal suspension processes are effectively distinguished, and problems are positioned to source code lines.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

Computing device, memory module group, computing equipment and computing equipment cluster

The invention provides a computing device, a memory module group, computing equipment and a computing equipment cluster, and belongs to the technical field of chip packaging. The computing device comprises at least one memory module group, a processor and a mainboard. Each memory module in the memory module group is formed by stacking the memory chips with the bit width of at least 8 bits, the size is small, the stacking process is simple, each memory module is accessed through at least one memory channel, and the density of the memory channels in unit area is large, so that the memory module group in the computing device can provide a large bandwidth, and the computing device is suitable for the computing device. According to the method, the requirements of AI or large-scale computing scenes on bandwidth are met, an HBM with high process difficulty and high cost does not need to be adopted, the engineering difficulty can be reduced, and the cost can be saved; besides, the processor and the memory module communicate based on the DRAM interface, and compared with an HBM interface, the delay is smaller, so that the power consumption generated by the delay is smaller, the data read-write efficiency is improved, and the random access efficiency is obviously improved.
Owner:HUAWEI TECH CO LTD

Multiply-accumulator based on unary calculation

The invention discloses a multiply-accumulator based on unary calculation, which is characterized by comprising a sequence generation unit, a signal processing unit, a signal processing unit, a signal processing unit, a signal processing unit and a signal processing unit, wherein the sequence generation unit is used for converting a fixed-point decimal between 0 / -1 and 1 into a unary sequence; the unary calculation multiply-accumulate unit is fused with the sequence generation unit; and the probability estimation unit is used for converting a unary calculation result obtained by the unary calculation multiply-accumulate unit back to the fixed-point decimal. The method has the following beneficial effects: a unary sequence generation unit with lower hardware overhead is provided, and the problem that the energy efficiency of unary calculation is reduced in a high-precision scene is solved; a novel low-delay unary calculation multiply-accumulate unit which is high in energy efficiency and can be terminated in advance is provided, and the problem that the energy efficiency of unary calculation is reduced under large-scale calculation is solved.
Owner:SHANGHAI TECH UNIV

A vehicle-network interaction supply and demand coordination mode deduction method

This invention relates to the field of electric vehicle charging and discharging technology, and more particularly to a method, system, and computer-readable storage medium for extrapolating a vehicle-to-grid (V2G) interaction supply-demand coordination model. The method first collects data and determines model input parameters based on local policies, electricity pricing mechanisms, and technological development trends. Then, it constructs a user utility function, eliminates differences in indicator dimensions through normalization, and calculates the probability of user selection for each charging mode, thus constructing a demand-side mode selection model. Based on user selection probabilities and facility operation indicators, it establishes a model for calculating facility expansion scale, and uses a dual-objective optimization framework to balance maximizing user satisfaction with minimizing investment costs to determine the annual number of new facilities. The number of new facilities is fed back to the user model to update facility accessibility parameters until the supply and demand status meets preset convergence conditions. This invention significantly improves the scientific nature and prediction accuracy of V2G interaction planning through dynamic closed-loop modeling of supply and demand coordination and a dual-objective optimization strategy.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +2

An abnormality positioning method based on a cloud management platform and a cloud management platform

PendingCN122309110AComprehensive considerationImprove the efficiency of abnormal locationNode clusteringExtreme scale computing
This application discloses an anomaly localization method and a cloud management platform based on a cloud management platform, which can improve the efficiency of anomaly localization to a certain extent. The method includes: the cloud management platform can create a large-scale computing node cluster for tenants, which can contain multiple computing nodes. If these multiple computing nodes fail to execute a job task, the cloud management platform can first obtain the logs generated by these multiple computing nodes during the execution of the job task, and generate log space characteristics and log time characteristics of these multiple computing nodes based on the logs. Then, the cloud management platform can evaluate these multiple computing nodes based on these log space characteristics and these log time characteristics to obtain evaluation values ​​for these multiple computing nodes. Subsequently, the cloud management platform can identify computing nodes whose evaluation values ​​are greater than or equal to a preset threshold as computing nodes that have experienced an anomaly.
Owner:HUAWEI TECH CO LTD

Storage particle, storage controller, storage chip, storage device and equipment

The invention discloses a storage particle, a storage controller, a storage chip, a storage device and equipment, relates to the technical field of computers, and is used for reducing the requirement of a model parameter matrix for memory capacity in a large-scale calculation process and solving the problem that the large-scale calculation performance is influenced by the storage medium interface bus rate. The storage device comprises a storage controller used for sending a plurality of model parameter sub-matrixes included in a model parameter matrix to a plurality of storage particles so as to store the plurality of model parameter sub-matrixes in the plurality of storage particles; the plurality of storage particles are used for concurrently calculating intermediate calculation results of the input data matrix and at least one model parameter sub-matrix stored in the storage particles, and sending the intermediate calculation results obtained by the storage particles to the storage controller; and the storage controller is also used for calculating the calculation results of the input data matrix and the model parameter matrix according to the plurality of intermediate calculation results correspondingly obtained by the plurality of storage particles.
Owner:HUAWEI TECH CO LTD

Scheduling method for cloud edge cooperative computing task of smart steel mill

The invention discloses a scheduling method for cloud edge cooperative computing tasks of a smart steel mill. The scheduling method comprises the steps of collecting computing tasks of a whole production link of the smart steel mill; obtaining the total number of CPU cores and the total number of memories in the cloud side server cluster; a K-Means algorithm is adopted to divide calculation tasks into cloud-assisted edge calculation and edge-assisted cloud calculation; constructing a constraint model, and constraining the communication time delay and the communication cost of the calculation task by taking minimization of the maximum communication time delay and the total communication cost of the task as a target; establishing two types of computing task agents, training the two types of computing task agents, and optimizing communication time delay and communication cost; and determining a distribution scheme of the calculation tasks according to the trained two types of calculation task intelligent agents. By adopting the method, dynamic scheduling optimization of large-scale computing tasks in the cloud edge collaborative manufacturing system is realized, the computing efficiency of the cloud edge server and the flexibility of the manufacturing process are effectively improved, and the communication burden of the cloud edge end is reduced.
Owner:NORTHEASTERN UNIV CHINA

Method and device for creating intelligent decision data set based on log data, electronic equipment, storage medium and program product

The invention provides a method and device for creating an intelligent decision data set based on log data, and relates to the technical field of cloud computing, cluster management, big data processing and artificial intelligence, and the method comprises the steps: collecting original log data from a server and a cluster, and achieving the collection and preprocessing of the log data; key information is extracted from the preprocessed log data, and log event analysis and structured processing are realized based on a self-attention mechanism; high-quality features are extracted and constructed from structured log data, and feature engineering and data enhancement are achieved; and combining the enhanced feature set with a predefined decision target to generate an intelligent decision data set for model training and evaluation, thereby realizing construction and evaluation of the intelligent decision data set. According to the method, high-quality data support is provided for the intelligent decision model, the system resource demand, the task execution time and the failure risk can be effectively predicted, the resource allocation efficiency is greatly optimized, and a solid foundation is provided for large-scale computing system resource management and scheduling.
Owner:BGP INC CHINA NAT PETROLEUM CORP +1

Secure multi-party computing method and system in memory limited environment

The invention provides a secure multi-party computing method and system in a memory limited environment. The method comprises the following steps: receiving a computing task of a user; dividing the calculation task into a plurality of data blocks according to the scale of the calculation task, the calculation operation and the available memory of each calculation node, and generating a plurality of calculation sub-tasks according to the dependency relationship of the calculation task; distributing the sub-calculation tasks to different calculation nodes, and carrying out parallel calculation; after calculation of each sub-calculation task is completed, a newly generated intermediate result secret share is dumped to an external memory, each participant performs hash on respective share files, and a Merkel tree containing all hash values is jointly constructed; and after calculation of all the sub-calculation tasks is completed, collaborating each participant to reconstruct the secret share of the final result block, recovering a plaintext result and returning the result to the user. According to the method, the privacy calculation of the large-scale matrix can be safely and reliably executed in a memory limited environment.
Owner:DAREWAY SOFTWARE

Hardware fault repairing method and device, medium and program product

PendingCN121880061AImprove fault operation and maintenance effectsResolve artificial dependenciesBiological modelsNon-redundant fault processingComputer hardwareLinguistic model
The invention discloses a hardware fault repairing method and device, a medium and a program product, and relates to the technical field of automatic operation and maintenance of large-scale computing clusters. The hardware fault repairing method comprises the steps of obtaining a preliminary repairing result of a target state machine; when the preliminary restoration result is that restoration cannot be carried out, acquiring preliminary restoration alarm data; and collecting alarm context information according to the preliminary repair alarm data, and generating repair plan data according to the alarm context information and a large language model. The technical scheme provided by the embodiment of the invention can be suitable for operation and maintenance requirements of complex hardware faults of a large-scale cluster, and the dependence on manpower is broken.
Owner:SHANGHAI TAIZE SEMICONDUCTOR CO LTD

Three-dimensional map construction method and device based on multi-machine collaboration

The invention provides a three-dimensional map construction method and device based on multi-machine collaboration, and the method comprises the steps: determining sampling region division information according to the target region information and preset sampling region area information when the target region information is received, and determining the deployment position of a ground terminal according to the target region information, a sampling advancing route corresponding to the sampling unmanned aerial vehicle group is generated according to the sampling area division information; according to the sampling area division information, the deployment position and the charging efficiency of the transmission unmanned aerial vehicle, determining the carrying number of the transmission unmanned aerial vehicle and a sampling area corresponding to each transmission unmanned aerial vehicle; and generating a first flight plan corresponding to each transmission unmanned aerial vehicle and a second flight plan corresponding to the sampling unmanned aerial vehicle group according to the sampling advancing route and the sampling area corresponding to each transmission unmanned aerial vehicle. The method does not need to wait for the return flight of the sampling unmanned aerial vehicle group and transmit all the map data to the ground terminal for large-scale calculation, and can save a lot of time and terminal calculation resources.
Owner:SHENZHEN TIANJING YUHONG TECHNOLOGY CO LTD

Data processing method based on reinforcement learning, electronic device and readable medium

ActiveCN116776097Befficient schedulingMeet training needsBiological modelsData transportEngineering
Embodiments of the present disclosure disclose a data processing method based on reinforcement learning, an electronic device and a readable medium. The method is applied to an execution node, a policy node and a training node included in a reinforcement learning system, and a specific implementation of the method comprises: performing environment simulation by the execution node to obtain an observation result; performing policy updating by the training node according to a training sample and a model to be trained, wherein the training sample is transmitted from the execution node to the training node through a sample transmission stream; the policy node performs policy inference according to the observation result and updated policy information, wherein in a non-in-line inference mode, the execution node and the policy node perform data transmission on the observation result and an action to be executed through an inference transmission stream; and the execution node executes the action to be executed to generate an updated observation result. The implementation realizes effective scheduling of large-scale computer hardware resources, meets the training requirements for complex reinforcement learning tasks, and improves the training speed.
Owner:SHANGHAI QI ZHI INSTITUTE

Electric energy substitution optimization method and system

The invention relates to the technical field of electric energy substitution optimization, and discloses an electric energy substitution optimization method and system. The method comprises the following steps: processing an earlier-stage multi-view through a first feature selection model to obtain an initial low-dimensional feature and an incidence matrix so as to initialize feature selection model parameters corresponding to a later-stage multi-view, and constructing a third feature selection model in combination with a multi-constraint condition; target low-dimensional features are obtained through a minimum total loss solving model, and finally an electric energy substitution optimization strategy is determined by combining with a historical load data input prediction model. According to the method, the data dimension is remarkably reduced through feature selection and dimension reduction processing, and computing resource consumption is reduced; the calculation efficiency is improved through staged model parameter initialization and constraint condition optimization. The technical problems that in the prior art, when large-scale power grid data is processed, calculation resource consumption is too large, calculation efficiency is low, and particularly in the face of complex power grid data and large-scale calculation, the real-time requirement is difficult to meet in the prior art are solved.
Owner:ZHANJIANG POWER SUPPLY BUREAU OF GUANGDONG POWER GRID CO LTD

A vehicle computing task scheduling method based on vehicle-edge-cloud collaboration

The application discloses a vehicle computing task scheduling method based on vehicle-edge-cloud cooperation and belongs to the technical field of intelligent transportation systems and Internet of Vehicles. The method is aimed at the problems of limited computing resources of intelligent networked vehicles, high task delay and large energy consumption. An improved greedy algorithm is used to optimize the task execution order, and local greedy selection is performed according to the data volume, computation volume and maximum tolerable delay of the task to output an optimal platform allocation label. Then, a feedforward neural network is trained using the labels. The trained neural network can quickly predict an optimal platform vehicle-edge-cloud cooperative allocation scheme according to the input task characteristics in a real-time scenario. Through the fusion of the algorithm and the neural network, the application realizes efficient online real-time scheduling, significantly improves the task scheduling efficiency and system reliability, has excellent real-time performance and scalability, effectively reduces the task execution delay and energy consumption, and is suitable for efficient scheduling and resource allocation of large-scale computing tasks in the Internet of Vehicles.
Owner:KUNMING UNIV OF SCI & TECH

Maximum and minimum radar echo GPU (Graphics Processing Unit) calculation acceleration method based on protocol in thread

The invention discloses a maximum and minimum radar echo GPU (Graphics Processing Unit) calculation acceleration method based on a protocol in a thread, which is based on a GPU calculation platform and adopts the protocol in the thread, and comprises the following steps of: firstly, opening up a shared memory of the thread by the GPU, carrying out initialization and assignment operation, executing the protocol for radar echo data of a plurality of threads, and then storing the radar echo data into the shared memory; secondly, executing protocols on shared memories in groups according to the GPU block groups, and storing results into local memories; and finally, executing inter-group protocols by adopting atomic operation, and storing the inter-group protocols into a global memory. According to the method, aiming at the problem of large-scale calculation caused under the condition of specific high sampling rate in the radar data, the loading operation of the thread is converted into the protocol loading operation, so that variables stored into a shared memory are reduced by half, and meanwhile, the problem of great waste of calculation resources is avoided.
Owner:NO 8511 RES INST OF CASIC

Hierarchical configuration management and automatic synchronization method and device for efficient molecular dynamics calculation

The embodiment of the invention provides a hierarchical configuration management and automatic synchronization method and device for efficient molecular dynamics calculation, and the method comprises the steps: obtaining configuration information from a superior node of a current node when the current node is started for each node in a molecular dynamics calculation cluster three-level architecture, establishing a version number, a change log and an event subscription relationship in the configuration information; when the current node configuration information is changed, generating a change event and a version number, recording a change log according to the change event, and spreading the change event through an event bus; regularly determining whether the configuration version numbers of the current node and the superior node are consistent or not; if not, sending a synchronization request to the superior node; according to the configuration version number of the current node, incremental data is extracted from the superior node, so that the current node applies incremental configuration to change the configuration information of the current node, and the three-level layered architecture can realize configuration management of a large-scale computing cluster and support differentiated configuration requirements.
Owner:SHANGHAI SMARTLOGIC TECHNOLOGY LTD

Model deployment method and system in resource-constrained environments

This invention provides a model deployment method and system for resource-constrained environments. By jointly introducing a tensor library structure and a matrix pairing mechanism, the parameter size is significantly reduced while maintaining the model's expressive power. The constructed compressed model only reconstructs and replaces the fully connected layers on the original network structure, without changing the overall forward propagation logic. The training process is simple and stable, and requires no additional large-scale computing resources; model training can be completed on ordinary GPUs or high-performance personal computers. The construction of the tensor library and the matrix pairing operation have low computational overhead, involving only standard tensor shrinkage operations, and the overall inference speed remains basically unchanged. The method of this invention maintains high prediction accuracy under different network sizes and compression ratios, and has good versatility and scalability, providing an efficient compression solution for deploying traditional neural networks and large models in resource-constrained environments.
Owner:SHANGHAI JIAOTONG UNIV

Multi-jet synchronous printing system based on a distributed architecture

PendingCN122331848AData packPersonalization
This invention relates to the field of high-precision industrial printing technology and discloses a multi-printerhead synchronous printing system based on a distributed architecture, including a workstation, an FPGA processing board, and print head modules. This system separates non-real-time large-scale computing from real-time high-precision control through a distributed architecture. When increasing the number of print heads, only additional computing tasks need to be added to the workstation, and logic units and I / O interfaces are expanded on the FPGA. The workstation parses the image data of each print job, completes dynamic load segmentation, establishes a segmentation index table, and then packages the data and embeds timestamps, ensuring that the FPGA processing board has sufficient time to complete decompression and caching, and also ensuring precise start of the printing operation at the specified time point. The distributed parsing has strong scalability, and the FPGA processing board constructs a global synchronization mechanism to ensure precise alignment of data from different channels in the time dimension. Waveform queries are performed on data packets to generate personalized drive waveforms adapted to each piezoelectric print head, resulting in excellent multi-printer collaboration.
Owner:WEINAN ZHENCHENG TECH CO LTD