Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2282 results about "Concurrent computation" patented technology

Concurrent computing is a form of computing in which several computations are executed during overlapping time periods— concurrently —instead of sequentially (one completing before the next starts). This is a property of a system—this may be an individual program, a computer, or a network —and there is a separate execution point or "thread of control" for each computation ("process").

Distributed real-time monitoring and early warning system for temperature field of smelting furnace

The invention discloses a distributed real-time monitoring and early warning system for a temperature field of a smelting furnace, and relates to the technical field of industrial process intelligent monitoring. The problems of accumulated measurement errors and non-stationary hotspot escape reconstruction hysteresis caused by static emissivity setting in an existing system are solved. Collecting multiband radiation intensity and voltage signals through time domain alignment of the multispectral sensor array and the thermocouple array; iterating emissivity parameters in real time by adopting a dynamic ash body spectrum ratio algorithm in combination with flue gas absorption characteristics; fusing non-contact and contact temperature measurement data based on weighted Kalman filtering and complementary filtering; constructing a space-time variable covariance function to carry out non-stationary Kriging interpolation; dynamically optimizing the local grid resolution by combining an adaptive grid module; the processing flow is accelerated through the parallel computing module; early warning is triggered based on abnormal probability judgment and is fed back to emissivity correction and grid optimization; according to the invention, the monitoring precision and real-time performance of the temperature field are obviously improved, and the risks of false alarm, missing alarm and equipment melting loss are effectively inhibited.
Owner:XICHUAN BEIJING JINYANG VANADIUM IND CO LTD

Automatic driving system based on real-time holographic modeling and dynamic shielding compensation and vehicle

The invention discloses an automatic driving system based on real-time holographic modeling and dynamic shielding compensation and a vehicle, and aims to solve the problem of complicated scene perception. The system comprises a real-time holographic environment modeling module for constructing a dynamic 3D environment model; the dynamic shielding compensation module outputs compensated environment data; the sensing data fusion module is used for integrating multi-source data; the high-precision positioning and mapping module is used for generating a centimeter-level environment map; the edge computing node is used for cooperatively processing fusion data; the AI self-adaptive decision making system generates a driving instruction, and the decision making efficiency is optimized and improved through a satellite auxiliary decision making unit; the V2Xcommunication module supports low-delay and high-security data interaction; the execution control unit is used for accurately controlling the vehicle; an FPGA + GPU parallel computing framework is adopted, and decision making and control are supported after modeling and compensation data are fused. According to the invention, optimal decision making and high-precision real-time control in a complex environment can be realized, and development of an automatic driving technology is promoted.
Owner:邵伟奇

Heterogeneous computing thread block optimal scheduling method and system based on dynamic topology mapping

The invention belongs to the field of parallel computing architecture optimization, and relates to a matrix multiplication acceleration method and system based on dynamic computing resource mapping, and the method comprises the steps: constructing a dynamic topology model driven by tensor dimension features, and generating a thread block distribution mode according to matrix parameters and GPU hardware information; constructing a multi-dimensional resource scheduling strategy library, dynamically selecting an optimal thread block distribution strategy from the multi-dimensional resource scheduling strategy library, and generating a binding relationship between the thread blocks and the data blocks; calculating collaborative access logic of thread blocks and storage hierarchies based on block parameters and dynamic mapping function optimization; distributed calculation is carried out, calculation and data transmission are parallelized through pipelining and a double-buffering mechanism, and result aggregation across calculation units is completed synchronously through atomic operation and a barrier. According to the method, discontinuous memory access conflicts can be effectively reduced, the execution efficiency of the calculation instruction and the utilization rate of the cache space are improved, the parallel calculation process of accelerating and optimizing the general matrix multiplication is realized, and the data processing efficiency is improved.
Owner:SOUTH CHINA UNIV OF TECH

Multi-platform concurrent transmission edge computing data acquisition method

The invention specifically relates to the technical field of edge computing, and discloses a multi-platform concurrent transmission edge computing data acquisition method, which comprises the following steps: S1, concurrent access of multi-platform equipment: establishing a mapping relationship between equipment identity and a communication link; s2, data concurrent collection and preprocessing: obtaining a preprocessed data stream; s3, concurrent transmission priority scheduling: allocating a transmission bandwidth for each data stream; s4, edge node parallel computing: distributing the subtasks to a plurality of computing nodes; s5, cloud collaboration and global optimization: generating a global optimization strategy; s6, carrying out full-link monitoring and dynamic regulation and control; concurrent access of heterogeneous equipment is realized by generating data streams in a uniform format, multi-platform compatibility is effectively enhanced, bandwidth is dynamically allocated to the data streams according to priority weight indexes by calculating the priority weight indexes of the data streams, data transmission efficiency is improved, and the data transmission efficiency is improved by allocating sub-tasks to a plurality of computing nodes. And the resource utilization rate of the edge computing gateway is improved.
Owner:NINGXIA JINRUNCHANG TECH CO LTD

Full-layout defect rapid detection system based on pattern matching and classification

The invention provides a full-layout defect rapid detection system based on pattern matching and classification, which relates to the technical field of semiconductor manufacturing and comprises a data acquisition and multi-source fusion module, a self-adaptive preprocessing module, a dynamic pattern matching module, a multi-source feature fusion module and a closed-loop optimization and output module, the dynamic mode matching module dynamically generates a defect template from real-time data through a clustering algorithm (such as DBSCAN) and a generative adversarial network (GAN), the speed and accuracy of defect detection are remarkably improved in combination with reinforcement learning filtering and region focusing technologies of the self-adaptive preprocessing module, a detection strategy can be adaptively adjusted according to the real-time data, and the defect detection accuracy is improved. According to the method, a high-risk area is preferentially matched, and GPU parallel computing acceleration processing is performed, so that the detection time is shortened to 50% or below of that of a traditional method, deformation or fuzzy defects can be effectively identified, the false alarm rate is reduced by about 30%, and high-quality input is provided for subsequent matching and classification.
Owner:上海芯无双仿真科技有限公司

Data production and application method based on index management

The invention discloses a data production and application method based on index management, and the method comprises the steps: receiving an index query request containing a target index identifier, a dimension constraint condition and a time range parameter, carrying out the analysis and semantic verification of the target index identifier based on a global index asset library, and obtaining the index definition information; generating a standardized query statement according to the index definition information, and selecting an optimal data calculation engine to execute query; performing multi-level cache optimization and parallel computing acceleration on the query process to obtain an original data set; processing the original data set in real time according to a preset analysis model to generate structured index data containing trend analysis, anomaly detection or attribution inference results; and filtering and desensitizing the data based on a fine-grained permission control strategy, only returning contents in a permission range and recording an audit log. By means of the method, unified management, efficient query and intelligent analysis of the index data are achieved, the automation level and query performance of data production are improved, and meanwhile data safety and access controllability are guaranteed.
Owner:FUJIAN PUPU INFORMATION TECH CO LTD

Self-adaptive efficient simulation method and system for electric propulsion plasma oscillation

The invention provides a self-adaptive efficient simulation method and system for electric propulsion plasma oscillation, and relates to the technical field of electric propulsion plasma oscillation simulation. According to the method, a high-precision electric propulsion plasma numerical simulation model is adopted to ensure the precision of a simulation result; according to the method, the calculation tasks of the simulation model are decomposed through various adjustment parameters, that is, all the calculation tasks are completed in a parallel calculation mode, and therefore the simulation efficiency can be effectively improved. Moreover, when the simulation model is subjected to simulation solving, the electromagnetic field solving model is solved by adopting a multi-grid method, and the simulation calculation efficiency can be further improved based on the solving mode. Besides, the target neural network model is utilized to process the simulation result so as to update all the adjustment parameters in the process, so that a plurality of adjustment parameters used in the simulation process are continuously and adaptively optimized, and the technical effect of considering the simulation efficiency and precision of the electric propulsion plasma oscillation simulation model is achieved.
Owner:BEIHANG UNIV

Heterogeneous unmanned aerial vehicle cluster intelligent task planning method

The invention relates to the technical field of heterogeneous unmanned aerial vehicle cluster task planning and intelligent scheduling, and discloses a heterogeneous unmanned aerial vehicle cluster intelligent task planning method, which is characterized by comprising task and unmanned aerial vehicle information initialization, task preliminary allocation based on a quantum annealing algorithm, real-time adjustment based on reinforcement learning, and a communication and cooperation mechanism. Executing and monitoring a task; the super-strong parallel computing capability and the quantum tunneling effect of a quantum annealing algorithm are utilized to quickly process a large-scale task allocation problem and find an approximate range of a globally optimal solution, and then an unmanned aerial vehicle cluster continuously learns and adjusts a task execution strategy according to real-time environment feedback in a task execution process by means of reinforcement learning, so that the task execution efficiency is improved. Therefore, the method adapts to dynamically changing scenes.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Integer parallel computing method and device based on distributed storage and computer equipment

The invention belongs to the field of high-performance computing, and relates to an integer parallel computing method and device based on distributed storage and computer equipment, and the method comprises the steps of collecting real-time resource indexes, dynamically identifying fault nodes, triggering task migration, and performing data verification and hard disk fault detection. The weight value of each node is calculated, the nodes are arranged according to the descending order of the weight values, and the nodes with high load capacity are selected to distribute tasks; dynamically distributing a data generation task to a computing node, executing parallel computing, and performing distributed storage on a result; obtaining an operand, converting the operand into a first-order tensor form of a basic operand, serializing tensor data, and sending the serialized tensor data to a parallel computing layer; distributing a search task to a computing node, retrieving storage data in parallel, reading effective data from a storage layer, and combining search results into a partial sum; and summarizing and then outputting. The system has dynamic resource management and fault-tolerant capabilities, and can realize efficient task allocation and load balancing.
Owner:SHENZHEN Y& D ELECTRONICS CO LTD

Knowledge graph reasoning method based on dynamic rule perception memory

The invention discloses a knowledge graph reasoning method based on dynamic rule perception memory. The method aims at solving the problems that an existing neural combination rule learning method is insufficient in expression ability and prone to splitting global semantics and local relation modes. The method comprises the following steps: introducing a lightweight dynamic relationship memory module, and executing self-attention on all relationships in a knowledge graph to capture global semantics; meanwhile, local relation interaction features are extracted through convolution, semantic fusion is carried out through a Transform encoder with multi-head attention and relative position coding, and unified and parallel-computing high-quality combination representation is constructed for each pair of relations in parallel. And meanwhile, a closed high-quality reasoning path is generated in combination with bidirectional breadth-first search. Through global semantic and local interactive collaborative modeling, the extendibility is ensured, the understanding of global semantics is enhanced, and the reasoning accuracy is improved. Global and local relation dependence is fused, an interpretable reasoning path is generated, and reasoning accuracy, expandability and robustness are improved.
Owner:NINGXIA UNIVERSITY +1

Parallel task scheduling algorithm for heterogeneous multi-core processor

The invention relates to the technical field of computer architecture and parallel computing, and discloses a parallel task scheduling algorithm for a heterogeneous multi-core processor, which comprises the steps of task modeling, resource mapping, dynamic load balancing, communication optimization, task scheduling decision and execution monitoring. Task allocation is adjusted in real time through dynamic load balancing, cross-core communication delay is reduced in combination with communication optimization, and an efficient task allocation sequence is generated by using an improved genetic algorithm. According to the method, the resource utilization rate and the task execution efficiency of the heterogeneous multi-core processor in a high-performance computing scene can be improved, meanwhile, the robustness and adaptability of an algorithm are enhanced, and the task allocation problem in a complex computing scene is effectively solved.
Owner:SUZHOU DUXUEKEZHENG INTELLIGENT TECH CO LTD

Fraud website identification early warning method, device and equipment and storage medium thereof

The invention relates to a fraud website identification early warning method, device and equipment and a storage medium thereof. The method comprises the following steps: firstly, acquiring network address information, then extracting feature data from dimensions such as structure, semantics, behavior and vision, accelerating similarity detection by applying a graphics processor parallel computing technology, and generating a composite feature vector containing a domain name similarity score and a registration feature; constructing fraud features in combination with the composite feature vector, mining a data association relationship through a graph neural network, and forming multi-dimensional risk feature data; and finally, processing the risk feature data by using a deep learning model, generating a risk probability, and when the risk probability exceeds a preset threshold, determining that the network address is a high-risk network address and generating an early warning report. By adopting the method, the comprehensiveness and accuracy of fraud website identification can be improved, complex fraud means such as domain name variation and semantic camouflage can be effectively dealt with, the real-time early warning capability for potential risks is enhanced, and efficient technical support is provided for network security protection.
Owner:CHONGQING JIAOTONG UNIV

Numerical control machining defect real-time detection method and system based on AI image recognition and medium

The invention provides a numerical control machining defect real-time detection method and system based on AI image recognition and a medium, and belongs to the technical field of intelligent manufacturing and machine vision crossing. The method comprises the following steps: collecting part surface image data according to a preset frequency, and separating a part from a background through noise reduction, contrast enhancement and a dynamic threshold algorithm of a data preprocessing module; inputting the processed image into a multi-modal feature fused AI detection model, and extracting and fusing multi-modal features to judge the type and position of a defect; and processing the image by using a parallel computing architecture, feeding back a detection result to a numerical control processing control system in real time, and visually displaying the detection result on a monitoring interface. According to the method, real-time and accurate detection and feedback control of numerical control machining defects are realized, and the machining quality and efficiency are effectively improved.
Owner:SHENZHEN HUAZHONG NUMERICAL CONTROL

Information physical security defense method, system and equipment based on distribution network digital twin simulation platform, and medium

The invention discloses an information physical security defense method, system and device based on a distribution network digital twin simulation platform and a medium, and belongs to the technical field of information security and control defense of the power distribution Internet of Things, and the method comprises the steps: constructing a three-dimensional channel for interaction of a physical power grid, an information model and simulation data; on the basis of a three-dimensional channel, a three-state dynamic conversion model among a steady state, a transient state and a recovery state is constructed, and multi-state automatic switching simulation is achieved through sub-region division and parallel computing; a multi-level risk modeling system oriented to various risks is constructed based on simulation results, complex attack scenes are identified, and mapping of risk types and response paths is achieved; security risk assessment and defense strategy generation verification are completed, and intelligent collaboration and continuous learning of cross-domain defense strategies are achieved through collaborative optimization. According to the method, the construction efficiency of a complex risk scene is improved, quantitative analysis of cross-domain risks such as circuit breaker mis-tripping caused by network attacks is realized, and the limitation of traditional single-domain risk analysis is broken through.
Owner:GUIZHOU POWER GRID CO LTD

Energy-saving optimization control method and system for communication base station

The invention relates to the technical field of communication base station energy-saving control, and discloses an energy-saving optimization control method and system for a communication base station. The method comprises the following steps: collecting and preprocessing traffic energy consumption data of a communication base station to obtain standardized characteristics; inputting the features into a double-flow space-time convolutional network, and carrying out parallel calculation through time and space branches to predict a load; performing hierarchical quantitative pruning on the network, and distributing bit width according to the influence degree to obtain a compression structure; and processing the real-time data based on the compression network, and generating an energy-saving control strategy for radio frequency resource configuration and power regulation. Based on accurate load prediction and differentiated resource scheduling strategies, the optimal balance between energy saving and service quality is realized, the problems of passive response and coarse-grained control in a traditional method are effectively solved, and the energy saving effect is maximized on the premise of ensuring communication quality.
Owner:JIAXING YIRUI ELECTRONICS CO LTD

Parallel computing method and system suitable for large-scale data processing

PCT designated stageWO2026007489A1Resource allocationResource poolPathPing
The present application relates to the technical field of large-scale data processing, and particularly relates to a parallel computing method and system suitable for large-scale data processing. The system comprises a task management unit, a distributed load balancing module, an elastic expansion architecture, an intelligent communication optimization module, and a resource monitoring unit, wherein the task management unit divides large-scale data into a plurality of sub-tasks by means of a task decomposer, and distributes the sub-tasks to computing nodes by means of a task scheduler and a priority distributor; the distributed load balancing module achieves global load balancing by means of a load sensing unit, a dynamic adjustment unit and a balance optimization unit; the elastic expansion architecture dynamically adjusts system resources by means of a node manager, a resource pool controller and an expansion decision-making device; the intelligent communication optimization module optimizes inter-node communication by means of a communication path planning unit, a bandwidth distribution unit and a delay compensation unit; and the resource monitoring unit monitors the system performance in real time by means of a performance collector, a state analyzer and an anomaly detector.
Owner:CHONGQING COLLEGE OF FINANCE ECONOMICS

Method and system for evaluating computer hardware performance based on analogue simulation model

The invention provides a computer hardware performance evaluation method and system based on an analogue simulation model, and relates to the technical field of computer performance evaluation. According to the method, hardware component models such as a processor, a memory and I / O are constructed by analyzing a system structure description file, semantic analysis is performed on an instruction stream or a parallel computing graph of a to-be-evaluated application, and a task load model is formed. The model is input into an event-driven scheduling module, simulation is executed on the hardware component model, task-resource mapping is generated, and task time delay, communication traffic and power consumption are recorded. And calculating average response time delay, data bandwidth and unit energy consumption according to the simulation data, and outputting an evaluation report compared with the performance baseline. According to the method, real load-oriented multi-dimensional performance prediction can be realized, and the method can be used for architecture optimization and scheme selection.
Owner:BEIJING ZUNGUAN TECH

Parallel computing method and device, electronic equipment and storage medium

The invention provides a parallel computing method and device, electronic equipment and a storage medium, and relates to the technical field of parallel computing, and the method comprises the steps: carrying out the first protocol operation of a target tensor based on each computing core in each stream processor cluster, and generating a data block containing the computing result of each computing core; writing a data block generated by each stream processor cluster into a shared cache; under the condition that each stream processor cluster completes the first protocol operation, reading a data block written by each stream processor cluster from the shared cache; and executing a second protocol operation on the data block read from the shared cache to generate a calculation result of the target tensor. According to the method and device provided by the invention, the parallel architecture and memory access characteristics of the artificial intelligence chip can be better matched, the unnecessary calculation delay and synchronization overhead of the cross-flow processor cluster in the parallel calculation process are reduced, the bandwidth utilization rate of the shared cache is improved, and the overall performance and calculation efficiency of parallel calculation are remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

Configurable number-theory transformation parallel computing acceleration method and device for post quantum cryptography algorithm

The invention discloses a configurable number-theory transformation parallel computing acceleration method and device for a post quantum cryptographic algorithm, and relates to the technical field of cryptographic algorithms. The number theory transformation is a calculation bottleneck in lattice-based post-quantum cryptography and PQC; a hardware accelerator specially designed for number theory transformation (NTT) is an effective solution for improving the execution speed. At present, a mainstream design neglects the influence of the scale of a calculation array, so that the design of an NTT calculation circuit is poor in flexibility and low in memory utilization rate, and a two-dimensional reconfigurable number theory transformation acceleration circuit is provided; the method is applied to a key generation stage, an encryption stage and a decryption stage of the post-quantum cryptographic algorithm. The NTT reconfigurable computing circuit provided by the invention can keep the utilization rate of hardware resources, and is suitable for mobile terminal equipment with limited resources; the NTT circuit adopts a low-complexity memory mapping scheme, so that the address control logic is greatly simplified, and the hardware overhead is reduced.
Owner:BEIJING INST OF TECH +1

Smart port tallying data processing method and system, medium and product

A tallying data processing method and system for an intelligent port, a medium and a product relate to the field of electric digital data processing, and the method comprises the following steps: obtaining a ship loading and unloading plan, tallying operation records and cost accounting data according to a preset period, and generating standard tallying data; extracting feature data in a pre-time window from the standard tallying data; inputting the feature data into a prediction calculation model, and generating a plurality of groups of loading and unloading plan change prediction scenes according to a historical operation mode; performing parallel calculation on the multiple groups of prediction scenes to obtain scene prediction data; comparing the real-time tallying data acquired at the port site with the scene prediction data, and calculating to obtain a data goodness of fit; selecting the most matched prediction scene as a current execution scene based on the data goodness of fit, and storing other prediction scenes into a standby scene library; and generating a tallying operation instruction and expense settlement data according to the current execution scene. By implementing the application, the efficiency of port multi-link data processing can be improved.
Owner:NANJING ZHONGLI WAILUN TALLY CO LTD

Visualization system and method for urban building group earthquake disaster simulation

The invention relates to the technical field of urban earthquake disaster simulation, and provides a visualization system and method for urban building group earthquake disaster simulation, and the system comprises an offline data preprocessing and block organization module and a block dynamic data management and scheduling module. The GPU acceleration rendering module is used for mapping a global floor index to a vertex to endow the vertex with a unique identifier, and finally submitting the combined model to a GPU, and cooperating with a displacement texture and a user-defined shader to realize real-time rendering of a building form; the GPU parallel displacement picture moving module is deeply coupled to the GPU rendering pipeline and is used for resolving the vertex displacement and the color value in real time and injecting the vertex displacement and the color value into the rendering pipeline; and the efficient interaction module dynamically updates label contents and coordinates to realize three-dimensional interaction. According to the method, massive time sequence displacement data are coded into textures, and parallel calculation is performed by using the GPU vertex shader, so that real-time and smooth animation simulation of earthquake responses of all buildings is realized, and the performance bottleneck of a traditional CPU calculation mode is solved.
Owner:TONGJI UNIV +1

Method, device and equipment for optimizing reasoning operator of large model based on mercuric chloride chip

The technical scheme can be applied to the field of financial science and technology / medical health. The invention discloses a mercuric chloride chip-based large model reasoning operator optimization method, device and equipment, and the method comprises the steps: carrying out the blocking processing of original input data according to the parallel calculation capability of a mercuric chloride chip, and generating the blocking data of an adaptive chip calculation unit; optimizing a memory access path of a matrix multiplication operator by combining a memory hierarchical structure and a calculation core type of a mercuric chloride chip based on the block data, and adjusting a sliding step length and a filling mode of convolution operation; a plurality of operators continuously executed in the large model are fused into a composite operator, the data storage process of the composite operator is optimized, and collaborative execution is achieved by dynamically allocating computing resources; and integrating and decoding block calculation results after collaborative optimization, and adjusting an input data block strategy and operator execution parameters through a verification feedback mechanism to form closed-loop optimization. According to the technical scheme, the reasoning efficiency of a large model can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Automatic online detection method and system for surface smoothness of precise injection mold

The invention provides an automatic online detection method and system for the surface smoothness of a precise injection mold, and relates to the field of precise manufacturing, the method comprises the following steps: S1, synchronously collecting visible light and near-infrared reflection images of the surface of the mold through multispectral coaxial imaging; s2, calculating a dynamic texture parameter based on a multi-directional gradient field weighted variance, and generating a composite reflection parameter in combination with wavelet domain fusion; s3, fusing the features by adopting a dual-path neural network, and comparing the standard library through a modified Stereer ratio model to output the degree of finish; and S4, adjusting the period in real time according to the detected dynamic standard deviation and linking closed-loop control. According to the method, the restriction of single-band characterization is broken through through multispectral data fusion, and the processing speed is increased by utilizing gradient field parallel computing hardware; in combination with a self-adaptive threshold value and a process parameter dynamic correction mechanism, the problems of poor adaptability of a fixed period, high hardware delay and single feature dimension of a traditional method are solved, and the online detection precision and the process control efficiency are remarkably optimized.
Owner:ZHEJIANG JIEZHONG SCI & TECH CO LTD

Gas compressor blisk rupture rotating speed prediction method based on virtual-real fusion

The invention provides a gas compressor blisk rupture rotation speed prediction method based on virtual-real fusion, belongs to the technical field of aero-engines, and particularly relates to an aero-engine gas compressor blisk rupture rotation speed prediction method by establishing a rupture rotation speed-oriented aero-engine gas compressor blisk digital twin model and performing double correction on the model and a failure criterion by using test real-time data. And in the simulation process, a parallel computing technology is used for acceleration, so that a high-precision fracture rotation speed prediction value which aims at a single individual of the aeroengine compressor blisk and accords with physical rules and actual test measurement can be quickly obtained.
Owner:TAIHANG LABORATORY

Attention operator head dimension block calculation method applied to sea light DCU

The invention relates to the field of heterogeneous parallel computing, in particular to an attention operator head dimension block computing method applied to a sea light DCU, which comprises the following steps of: expanding a sequence length block computing range in charge of each thread block in an attention operator and a matrix computing range processed by a single thread working group; a blocking parameter splitqk and a blocking parameter splitv are selected, wherein the blocking parameter splitqk and the blocking parameter splitv are selected; selecting a proper block cutting parameter splitqk for the head dimensions of the query tensor Q and the key tensor K for cutting, accumulating block calculation results, and calculating S block calculation results through a softmax function to obtain P block calculation results; selecting a proper block cutting parameter splitv for the head dimensions of the value tensor V and the output tensor O for cutting, wherein a block calculation result corresponds to a corresponding block of the output tensor O; and finally, segmenting the tensor O according to the splitv, and writing each partitioning result back to the global memory. The method is suitable for a heterogeneous parallel computing system composed of a CPU and a DCU, computing resources of a DCU computing unit can be saved, performance loss caused by resource overflow is reduced, and computing efficiency is improved.
Owner:SOUTH CHINA UNIV OF TECH

Transformer winding state detection method based on temperature characteristics

The invention provides a transformer winding state detection method based on temperature characteristics. The method comprises the following steps: (1) collecting operation data; (2) preprocessing the operation data; (3) taking the normalized transformer load current, the normalized active power, the normalized top oil temperature, the normalized ambient temperature and the normalized oil flow speed as input data, taking the normalized winding hot-spot temperature as output data, and constructing a data set; (4) establishing a VSN-TKAN-GRN deep learning network parallel computing model of the transformer winding hot-spot temperature considering the multi-head attention mechanism; (5) predicting the hot-spot temperature of the transformer winding based on a VSN-TKAN-GRN deep learning network parallel computing model; and (6) calculating the average relative error percentage of the transformer winding hot-spot temperature prediction result and the actual measurement result, and judging the winding state according to the change of the average relative error percentage. According to the method, the state of the transformer winding in operation can be efficiently and accurately judged, so that effective transformer operation and maintenance measures can be taken in time, and large faults are avoided.
Owner:CHINA YANGTZE POWER

Command and control system resource trend prediction method based on fusion of long and short time sequence characteristics

The invention discloses a command and control system resource trend prediction method based on fusion of long and short time sequence characteristics. The method comprises the following steps: acquiring a public power load or similar time sequence monitoring data set, and preprocessing the data in the data set; a deep learning network model based on a TCN-Transformer hybrid model is constructed, a TCN model and a Transformer model are adopted for parallel computing to achieve feature extraction, the TCN model extracts short-term information, the Transformer model extracts long-term features, then fusion features are obtained through a cross attention mechanism and multi-layer perceptron (MLP) weighting, and finally prediction output is generated through full connection layer mapping. Taking data in the training set as input, training the constructed TCN-Transform hybrid model, and continuously optimizing the model until convergence meets a set requirement; and performing prediction by using the trained network model. According to the method, the TCN-Transform hybrid model is constructed, so that local fine-grained features are reserved, the global time trend is effectively captured, and the accuracy of command decision making is improved.
Owner:NANJING UNIV OF SCI & TECH

Large-scale phased array multi-channel partition beam control method and device based on FPGA and storage medium

The invention discloses a large-scale phased array multi-channel partition beam control method and device based on an FPGA and a storage medium. The method comprises the steps that a control instruction of an antenna control unit is acquired; analyzing the control instruction which is verified to be correct, judging an instruction type according to an analysis result, when the instruction type is a beam switching instruction, determining a beam control parameter in the instruction, and calculating a base code of the phased-array antenna according to the beam control parameter; the array elements of the phased-array antenna are divided into a plurality of areas, according to the base codes and the position coordinates of the array elements, the wave control codes of the array elements are calculated in parallel by adopting a multi-channel partitioning method, and the wave control codes are normalized and then written into a beam forming chip in parallel; and after the operation is completed, a corresponding response instruction is sent to the antenna control unit. The method has the advantages of being high in beam switching speed, small in resource consumption and good in universality, an array can form a sparse array in the mode of sending a control instruction, and therefore side lobes are restrained, and power consumption is reduced.
Owner:WUHAN ZHONGYUAN COMM CO LTD

Power distribution area topology identification method and system based on intelligent fusion terminal and correlation coefficient

The invention discloses a power distribution area topology identification method and system based on an intelligent fusion terminal and correlation coefficients, and the method comprises the steps: collecting the multi-source data of a power distribution area through an edge fusion terminal, and carrying out the data preprocessing and feature extraction at a terminal side; carrying out parallel calculation on a voltage fluctuation Pearson's correlation coefficient and a load change Kendall rank correlation coefficient, and constructing a correlation coefficient matrix in different time periods; calculating an adaptive weight based on the load fluctuation entropy, and fusing a dual-mode correlation coefficient; carrying out topology generation by using a graph neural network, and outputting an edge existence probability and a node hierarchy; performing physical constraint optimization on the initial topology by adopting a genetic algorithm; the system comprises an edge fusion terminal cluster, a cloud analysis platform, a topology verification module and a terminal management platform. According to the method, topology recognition precision and dynamic adaptability are remarkably improved, bimodal correlation coefficients and graph neural network space modeling are creatively fused, and a complex topological structure is precisely restored.
Owner:JIANGSU HONGYUAN ELECTRIC

Implementation of hierarchical navigable small world (HNSW) search techniques using NAND memory

To accelerate search speeds for approximate nearest neighbor searches of vector databases, compute-in-memory techniques using NAND memory structures are introduced. For each element of the database, a kernel of its M nearest neighbors is determined. For each vector of the database, both the vector and its kernel are programmed in the arrays of a NAND memory based accelerator card, so that the vectors will be written into the memory arrays both as themselves and also in kernels of vectors for which they are a nearest neighbor. Metadata, associating the locations of the kernel members with the correspond vector is also stored in the memory system. After determining the input's nearest neighbor at one level of search, the metadata is then used to locate that nearest neighbor's nearest neighbors and their distances to the input vector are then computed in parallel in a compute-in-memory vector-vector dot product multiplication.
Owner:SANDISK TECHNOLOGIES LLC