Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1326 results about "Concurrent computation" patented technology

Concurrent computing is a form of computing in which several computations are executed during overlapping time periods— concurrently —instead of sequentially (one completing before the next starts). This is a property of a system—this may be an individual program, a computer, or a network —and there is a separate execution point or "thread of control" for each computation ("process").

Knowledge graph reasoning method based on dynamic rule perception memory

The invention discloses a knowledge graph reasoning method based on dynamic rule perception memory. The method aims at solving the problems that an existing neural combination rule learning method is insufficient in expression ability and prone to splitting global semantics and local relation modes. The method comprises the following steps: introducing a lightweight dynamic relationship memory module, and executing self-attention on all relationships in a knowledge graph to capture global semantics; meanwhile, local relation interaction features are extracted through convolution, semantic fusion is carried out through a Transform encoder with multi-head attention and relative position coding, and unified and parallel-computing high-quality combination representation is constructed for each pair of relations in parallel. And meanwhile, a closed high-quality reasoning path is generated in combination with bidirectional breadth-first search. Through global semantic and local interactive collaborative modeling, the extendibility is ensured, the understanding of global semantics is enhanced, and the reasoning accuracy is improved. Global and local relation dependence is fused, an interpretable reasoning path is generated, and reasoning accuracy, expandability and robustness are improved.
Owner:NINGXIA UNIVERSITY +1

Parallel task scheduling algorithm for heterogeneous multi-core processor

The invention relates to the technical field of computer architecture and parallel computing, and discloses a parallel task scheduling algorithm for a heterogeneous multi-core processor, which comprises the steps of task modeling, resource mapping, dynamic load balancing, communication optimization, task scheduling decision and execution monitoring. Task allocation is adjusted in real time through dynamic load balancing, cross-core communication delay is reduced in combination with communication optimization, and an efficient task allocation sequence is generated by using an improved genetic algorithm. According to the method, the resource utilization rate and the task execution efficiency of the heterogeneous multi-core processor in a high-performance computing scene can be improved, meanwhile, the robustness and adaptability of an algorithm are enhanced, and the task allocation problem in a complex computing scene is effectively solved.
Owner:SUZHOU DUXUEKEZHENG INTELLIGENT TECH CO LTD

Parallel computing method and system suitable for large-scale data processing

PCT designated stageWO2026007489A1Resource allocationResource poolPathPing
The present application relates to the technical field of large-scale data processing, and particularly relates to a parallel computing method and system suitable for large-scale data processing. The system comprises a task management unit, a distributed load balancing module, an elastic expansion architecture, an intelligent communication optimization module, and a resource monitoring unit, wherein the task management unit divides large-scale data into a plurality of sub-tasks by means of a task decomposer, and distributes the sub-tasks to computing nodes by means of a task scheduler and a priority distributor; the distributed load balancing module achieves global load balancing by means of a load sensing unit, a dynamic adjustment unit and a balance optimization unit; the elastic expansion architecture dynamically adjusts system resources by means of a node manager, a resource pool controller and an expansion decision-making device; the intelligent communication optimization module optimizes inter-node communication by means of a communication path planning unit, a bandwidth distribution unit and a delay compensation unit; and the resource monitoring unit monitors the system performance in real time by means of a performance collector, a state analyzer and an anomaly detector.
Owner:CHONGQING COLLEGE OF FINANCE ECONOMICS

Parallel computing method and device, electronic equipment and storage medium

The invention provides a parallel computing method and device, electronic equipment and a storage medium, and relates to the technical field of parallel computing, and the method comprises the steps: carrying out the first protocol operation of a target tensor based on each computing core in each stream processor cluster, and generating a data block containing the computing result of each computing core; writing a data block generated by each stream processor cluster into a shared cache; under the condition that each stream processor cluster completes the first protocol operation, reading a data block written by each stream processor cluster from the shared cache; and executing a second protocol operation on the data block read from the shared cache to generate a calculation result of the target tensor. According to the method and device provided by the invention, the parallel architecture and memory access characteristics of the artificial intelligence chip can be better matched, the unnecessary calculation delay and synchronization overhead of the cross-flow processor cluster in the parallel calculation process are reduced, the bandwidth utilization rate of the shared cache is improved, and the overall performance and calculation efficiency of parallel calculation are remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

Command and control system resource trend prediction method based on fusion of long and short time sequence characteristics

The invention discloses a command and control system resource trend prediction method based on fusion of long and short time sequence characteristics. The method comprises the following steps: acquiring a public power load or similar time sequence monitoring data set, and preprocessing the data in the data set; a deep learning network model based on a TCN-Transformer hybrid model is constructed, a TCN model and a Transformer model are adopted for parallel computing to achieve feature extraction, the TCN model extracts short-term information, the Transformer model extracts long-term features, then fusion features are obtained through a cross attention mechanism and multi-layer perceptron (MLP) weighting, and finally prediction output is generated through full connection layer mapping. Taking data in the training set as input, training the constructed TCN-Transform hybrid model, and continuously optimizing the model until convergence meets a set requirement; and performing prediction by using the trained network model. According to the method, the TCN-Transform hybrid model is constructed, so that local fine-grained features are reserved, the global time trend is effectively captured, and the accuracy of command decision making is improved.
Owner:NANJING UNIV OF SCI & TECH

Power distribution area topology identification method and system based on intelligent fusion terminal and correlation coefficient

The invention discloses a power distribution area topology identification method and system based on an intelligent fusion terminal and correlation coefficients, and the method comprises the steps: collecting the multi-source data of a power distribution area through an edge fusion terminal, and carrying out the data preprocessing and feature extraction at a terminal side; carrying out parallel calculation on a voltage fluctuation Pearson's correlation coefficient and a load change Kendall rank correlation coefficient, and constructing a correlation coefficient matrix in different time periods; calculating an adaptive weight based on the load fluctuation entropy, and fusing a dual-mode correlation coefficient; carrying out topology generation by using a graph neural network, and outputting an edge existence probability and a node hierarchy; performing physical constraint optimization on the initial topology by adopting a genetic algorithm; the system comprises an edge fusion terminal cluster, a cloud analysis platform, a topology verification module and a terminal management platform. According to the method, topology recognition precision and dynamic adaptability are remarkably improved, bimodal correlation coefficients and graph neural network space modeling are creatively fused, and a complex topological structure is precisely restored.
Owner:JIANGSU HONGYUAN ELECTRIC

Implementation of hierarchical navigable small world (HNSW) search techniques using NAND memory

To accelerate search speeds for approximate nearest neighbor searches of vector databases, compute-in-memory techniques using NAND memory structures are introduced. For each element of the database, a kernel of its M nearest neighbors is determined. For each vector of the database, both the vector and its kernel are programmed in the arrays of a NAND memory based accelerator card, so that the vectors will be written into the memory arrays both as themselves and also in kernels of vectors for which they are a nearest neighbor. Metadata, associating the locations of the kernel members with the correspond vector is also stored in the memory system. After determining the input's nearest neighbor at one level of search, the metadata is then used to locate that nearest neighbor's nearest neighbors and their distances to the input vector are then computed in parallel in a compute-in-memory vector-vector dot product multiplication.
Owner:SANDISK TECHNOLOGIES LLC

Dynamic topology mapping and GPU heterogeneous flood coupling model real-time early warning method

The invention discloses a dynamic topology mapping and GPU heterogeneous flood coupling model real-time early warning method, and solves the problems of inflexible coupling of a one-dimensional model and a two-dimensional model, low calculation efficiency, poor dynamic response capability and the like in traditional flood simulation. A dynamic topology mapping mechanism is introduced to respond to emergencies such as dike burst and overflow in real time, the calculation efficiency is remarkably improved by utilizing a GPU heterogeneous parallel architecture, rapid and accurate simulation of the whole process of flood formation, evolution and submerging is achieved, reliable technical support is provided for flood control dispatching and emergency decision making, and the method is suitable for large-scale popularization and application. The method comprises the following steps: constructing a two-dimensional hydrodynamic model based on an unstructured grid, establishing a dynamic topology mapping mechanism and realizing GPU parallel numerical calculation; a two-dimensional model lateral connection coupling mechanism is constructed, dam overtopping flow is calculated in real time, and breach flow dynamic calculation of an improved DAMBRK method is carried out; multi-target parameter calibration, multi-dimensional model verification and one-two-dimensional model GPU collaborative parallel computing are realized.
Owner:JILIN ELECTRIC POWER RES INST LTD +1

Data processing method and device, computer equipment, readable storage medium and program product

The invention relates to a data processing method and device, computer equipment, a computer readable storage medium and a computer program product. After first matrix data is obtained, tensor data to be processed are obtained, and the tensor data are distributed to a plurality of calculation units; calculating a local feature value corresponding to the target calculation unit according to the local tensor data allocated to the target calculation unit; dividing the target calculation unit into data blocks corresponding to the number of the plurality of calculation units; according to the serial number of the data block, mapping the local feature value recorded in each data block to a target storage block of a target memory until the local feature value in each calculation unit is distributed to the corresponding storage block in the target memory, and obtaining a global feature value table; according to the global feature value table, feature value calculation is carried out to determine the target feature value, conflict-free parallel calculation is achieved, the utilization rate of calculation resources is increased, and the data processing efficiency is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Cross-platform video decoding method and device

The invention provides a cross-platform video decoding method and device, computing equipment and a computer readable storage medium, and the method is combined with a factory mode to completely package bottom layer differences of all platform types, and an application layer developer does not need to care about the bottom layer platform differences. The optimal implementation can be automatically selected on different engines and platforms through a uniform API (Application Program Interface); during rendering, a single-frame memory management strategy is used, and only one frame which is currently displayed is reserved in a memory; and in the aspect of performance, the advantages of a multi-thread platform and the parallel computing characteristics of the GPU are fully utilized. According to the embodiment of the invention, through deep fusion of a cross-platform soft decoding architecture, single-frame memory optimization, multi-thread parallel, GPU acceleration and other technologies, perfect balance of cross-platform, high performance and low memory is realized, and the development efficiency of the whole project is improved.
Owner:TUYOO GAMES +3

Large language model reasoning system and method based on multi-chip parallel computing

The invention provides an inference system and method of a large language model based on multi-chip parallel computing, and relates to the technical field of artificial intelligence. The system comprises a pre-calculation module used for processing input instruction information to generate to-be-reasoned data, and the to-be-reasoned data is in a matrix form; the expert parallel module is used for sending the to-be-reasoned data to accelerator chips in the expert parallel module and determining sub-reasoning data processed by the activation expert units corresponding to the accelerator chips respectively, so that the activation expert units carry out calculation based on the corresponding sub-reasoning data and complete parallel calculation result data is determined. The input data is broadcasted to all the accelerator chips, each accelerator chip selects the corresponding input data for calculation according to the set activation expert unit, the same complete calculation result is obtained through global protocol operation among all the accelerator chips, and the overall operation performance and efficiency are improved.
Owner:SHENZHEN CORERAIN TECH CO LTD

Container-based parallel computing system

A container-based parallel computing system for executing high-performance computing (HPC) applications. The system leverages container technology to package the applications executed at the nodes in a cluster. To load and execute a job in the parallel computing system, containers are deployed in a cluster that include all the application resources and configuration information that the particular HPC application needs to execute. An event-driven batch scheduler may be used to dynamically allocate resources for executing multi-node jobs in the container-based parallel computing system, handling the coordination of resource allocation for the customer. The scheduler insures that jobs begin executing as fast as possible, and handles failure conditions such as partial scaling. Virtual network interfaces are attached to the containers that allow the containers to connect to and communicate with other containers in the cluster directly through the network interfaces of host machines using IP addresses provided by the virtual network interfaces.
Owner:AMAZON TECH INC

Three-dimensional scene rendering method and device based on graphics processor, equipment and medium

The invention relates to the technical field of image rendering, and discloses a three-dimensional scene rendering method, device and equipment based on a graphics processor and a medium, the method is applied to the graphics processor, and the method comprises the following steps: receiving instance attribute data sent by a central processing unit; calling a calculation shader to distribute a calculation thread group according to the number of objects in the instance attribute data; calculating a corresponding transformation matrix for each object in parallel through the calculation thread group; adding the instance attribute data and the transformation matrix into an instance buffer area to obtain to-be-rendered instance data of each object; extracting to-be-rendered instance data from the instance buffer area, and performing invisible object removal processing on the to-be-rendered instance data through the calculation thread group to obtain target instance data of a visible object; and rendering the target instance data. According to the invention, the rendering effect and rendering efficiency of the three-dimensional scene are improved.
Owner:CHONGQING WUTONG CAR LINK TECH CO LTD

CFD solution acceleration method and system based on parallel computing and load dynamic balancing

The invention relates to the technical field of calculation, and provides a CFD solution acceleration method and system based on parallel calculation and load dynamic balance. Physical characteristics such as a CFL number and a residual error are quantified into load indexes, so that a load balancing decision and a flow field evolution rule are tightly coupled; by adopting an asynchronous communication overlapping framework deeply coupled with a CFD solver, blocked communication waiting time is converted into effective calculation time, and communication delay is systematically hidden; based on an incremental grid migration strategy of real-time load evaluation, high-load grid blocks in a hotspot process can be intelligently identified and migrated, and differentiation and high-efficiency utilization of computing power are realized.
Owner:SHENZHEN Y& D ELECTRONICS CO LTD

Communication optimization method for topological table persistent storage in efficient parallel computing

The invention provides a communication optimization method for topological table persistent storage in efficient parallel computing, belongs to the technical field of storage communication, and aims to avoid pseudo sharing by aligning memory allocation through cache lines, expand local topological coverage through a topological entropy increment driven prefetching mechanism, and improve the reliability of the topological table persistent storage. Boundary processing is accelerated through a pre-calculation period mapping lookup table and a frequency domain transfer function vector, synchronization overhead is optimized through a concurrency control mechanism perceived by a read-write ratio, targeted cache preloading is achieved through stability and jitter degree two-dimensional evaluation, bandwidth consumption is reduced through an increment synchronization mechanism, and the stability of the cache is improved. The communication template is selected or the communication parameters are generated through adaptive conversion by matching the matching degree decision, and the technical problem that the parallel computing performance is reduced due to the fact that the communication overhead is too large in the topological table persistent storage process is solved.
Owner:青岛国实科技集团有限公司

Group intelligence driven cascade reservoir autonomous negotiation scheduling method

The invention relates to a swarm intelligence-driven cascade reservoir autonomous negotiation scheduling method. The method comprises the following steps: firstly, generating a scheduling basic data set; respectively packaging each reservoir node of the cascade reservoir group into an independent reservoir unit body; each reservoir unit carries out parallel computing and distributed negotiation through a group fusion negotiation strategy based on interaction data of a local scheduling basic data set and a neighborhood reservoir unit, and a preliminary scheduling scheme is generated; the group fusion negotiation strategy is based on the function types of the reservoir units, adopts a distributed negotiation mode of a contract network protocol or a bidding mechanism, takes the minimum transaction cost among the reservoir units as a core objective function, and combines flood control, power generation and ecological multi-objective weight coefficients to determine a preliminary scheduling scheme; and carrying out digital twinborn simulation verification and fine tuning to obtain a final scheme. In case of exception, preferential intra-group coordination is realized, and in case of invalidation, cross-group linkage is realized. The method improves the scheduling efficiency, guarantees the multi-target balance of flood control, power generation and the like, and enhances the anti-risk capability.
Owner:YELLOW RIVER INST OF HYDRAULIC RES YELLOW RIVER CONSERVANCY COMMISSION +1

Instruction-level simulation and performance modeling system for parallel computing architecture

The invention provides an instruction-level simulation and performance modeling system for a parallel computing architecture, and belongs to the technical field of computer architecture and simulation verification, and the system comprises an instruction modeling layer which is used for analyzing and executing an intermediate instruction set defined by the architecture; the scheduling execution layer is used for simulating a multi-thread and multi-core parallel execution process; the storage access layer is used for constructing a hierarchical storage access and bandwidth and delay model; and the performance analysis layer is used for collecting and counting key indexes such as an execution period, an instruction utilization rate and memory access delay, and realizing accurate performance modeling of the parallel architecture. According to the method, the performance bottleneck of the design scheme can be rapidly evaluated in the early stage of architecture design, the simulation speed is high, the module configurability is high, the modeling precision is adjustable, and the method is suitable for functional verification, micro-architecture exploration and compiler performance analysis of parallel computing architectures, accelerator chips, heterogeneous multi-core processors and the like.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Flow scheduling method and device, medium and program product

The invention discloses a flow scheduling method and device, a medium and a program product, and relates to the technical field of heat dissipation, and the method comprises the steps: obtaining a video memory access frequency as a main parameter when a parallel computing function is executed, matching with other auxiliary parameters related to heat consumption, distributing a weight, determining a heat load scoring result, and then adjusting a cold plate micro-channel flow distribution coefficient, and determining an opening value of a flow control valve and sending a flow scheduling instruction to the liquid cooling system so as to dynamically adjust the cooling liquid flow of each cold plate micro-channel. In this way, accurate temperature control and heat dissipation of the video memory area of the computing device can be achieved, energy waste caused by unnecessary circulation of cooling liquid is avoided, the heat dissipation efficiency is remarkably improved, energy consumption is reduced, and the service life of the computing device is prolonged; in addition, the mode can autonomously complete full-flow operation such as video memory heat consumption monitoring, score calculation and flow adjustment, intelligent management of the heat dissipation process is achieved, manual operation errors and cost are reduced, and the stability and reliability of operation of computing equipment are guaranteed.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Task scheduling method and device for attention calculation, medium, equipment and product

The invention discloses a task scheduling method and device for attention calculation, a medium, equipment and a product. The method comprises the steps that a loading task of a query block is decomposed into N1 loading subtasks, and a calculation task of a jth attention score block is decomposed into N2 calculation subtasks; decomposing an updating task of a middle accumulation block corresponding to the attention output block in the current iteration round into N3 parallel updating sub-tasks; asynchronously submitting all the subtasks to different task slots of N4 synchronization channels; and performing polling operation on the task slot in each synchronization channel to execute the sub-tasks stored in the task slot, and controlling an execution time sequence among the sub-tasks with the dependency relationship in each synchronization channel by configuring a number threshold value of the uncompleted sub-tasks in the waiting instruction. According to the method, the covering capability and the parallel computing capability of hardware can be improved, and the attention computing efficiency is further improved.
Owner:SHANGHAI BIREN TECH CO LTD

Infrared bidirectional heat effect simulation method based on radiation intensity and GPU acceleration

The invention discloses an infrared bidirectional thermal effect simulation method based on radiation and GPU acceleration, and belongs to the field of thermal radiation simulation and parallel computing. Dispersing the complex geometry into patch units, constructing an inter-patch radiation energy balance equation based on a radiance theory, and iteratively solving the final temperature of the patches through a radiance method; a three-level parallel strategy of a task level, a data level and an instruction level is designed, data intensive tasks such as shape factor calculation and radiance iteration are migrated to a GPU to be executed, data storage is optimized in combination with a structure array (SoA) layout and a sparse matrix compression technology, and efficient data sharing of a CPU end and a GPU end is achieved through a CUDA zero copy technology. According to the method, the dependence of a traditional method on regular grids is broken through, the calculation efficiency and precision of million-level surface patch heat radiation transmission in a complex scene are remarkably improved, and an efficient tool is provided for heat radiation indirect transmission calculation and a global illumination model in the field of three-dimensional scene infrared simulation.
Owner:ZHEJIANG UNIV

Metal additive manufacturing three-dimensional temperature field Gaussian process prediction method

The invention relates to a metal additive manufacturing three-dimensional temperature field Gaussian process prediction method, belongs to the technical field of material science, and particularly relates to a metal additive manufacturing three-dimensional temperature field prediction method. Logarithmic transformation is adopted to preprocess a temperature field, and prediction difficulty caused by extreme gradient near a molten pool is avoided; dividing the overall computational domain into a plurality of sub-domains by adopting a domain decomposition strategy to reduce the problem dimension; for each sub-domain, further combining singular value decomposition to extract a temperature field reduced-order base; establishing a local Gaussian process regression model based on a Maren kernel function and carrying out parallel training so as to establish rapid mapping from process parameters to reduced-order output; during online prediction, efficient and accurate prediction of a complete temperature field is realized through parallel calculation and full-field assembly of each local model.
Owner:BEIJING INST OF TECH

Graph convolution traffic flow prediction method based on attention mechanism and hub node enhancement

The invention discloses a graph convolution traffic flow prediction method based on an attention mechanism and hub node enhancement. The method comprises the steps that an input layer receives traffic flow data and transmits the traffic flow data to a space-time layer after dimension adjustment; the space-time layer firstly extracts traffic flow data of different time periods through a multi-scale feature fusion module, fuses a time sequence mode vector and node features, then calculates a node association strength matrix through an encoder module, identifies hub nodes in combination with improved indexes and generates an importance adjacency matrix, and finally performs data processing on the basis of the importance adjacency matrix. Generating a hidden state by using a dynamic graph convolution gating loop unit; the multi-head time attention module keeps a time sequence, and after attention weights are calculated in parallel, residual connection and layer normalization processing are carried out; the decoder module generates a multi-step prediction result through two-dimensional graph convolution; the output layer converts and outputs the prediction data. The method improves the traffic flow prediction accuracy, and provides support for traffic control.
Owner:SOUTHWEST JIAOTONG UNIV

Spectral red shift measurement method, device, equipment and medium

The embodiment of the invention provides a spectrum red shift measurement method, device and equipment and a medium, which can be applied to the fields of astronomical information technology and high-performance computing technology, and the method comprises the following steps: receiving N observation spectrum data to be measured; determining a processing mode of the observed spectrum data; in response to the processing mode being a first processing mode, processing the observation spectral data through a first calculation path, the first calculation path performing red shift measurement calculation based on a first numerical calculation library by matching the observation spectral data with a set of template spectral data to obtain a calculation result, the calculation result of the first numerical calculation library is consistent with the calculation result of a reference algorithm within a preset precision range; in response to the processing mode being a second processing mode, the observed spectral data is processed through a second computational path that reconstructs the redshift measurement calculations into a tensor parallel model based on a second library of numerical calculations and performs on parallel computing hardware, the second library of numerical calculations supporting tensor calculations.
Owner:NAT ASTRONOMICAL OBSERVATORIES CHINESE ACAD OF SCI

Hand-eye cooperative robot control system and method based on dynamic operator arrangement

The invention discloses a hand-eye cooperative robot control system and method based on dynamic operator arrangement, an upper computer planning layer operates a master control computer to periodically trigger task scheduling, a joint controller of a lower computer execution layer receives an instruction through a redundant bus to perform servo control, hand-eye camera data triggers a visual assembly line through an interrupt event, and a visual assembly line is controlled through a control interface. According to the method, pressure is calculated through visual processing, motion planning and joint control in a load sharing mode, a double-buffering mechanism is adopted to enable current frame visual processing and previous frame motion control to be executed in an overlapping mode, real-time scheduling and parallel computing of tasks are completed, hot data are cached in a memory database, cold data are archived to an HDFS and migrated through an LRU strategy, and the real-time scheduling and parallel computing of the tasks are completed. Data type conversion is automatically derived and performed based on a feature type registry, and conditional branch execution is triggered according to real-time sensor data. According to the control system, cross-hardware plug and play, algorithm flow configurability and data flow real-time sharing are achieved, and the requirement for improving the efficiency of complex operation tasks is met.
Owner:SUPER HIGH VOLTAGE BRANCH OF STATE GRID JIANGXI ELECTRIC POWER CO LTD

Dynamic batching for inference system for transformer-based generation tasks

An inference system applies a machine-learning transformer model to a batch of requests with variable input length or variable target length or variable internal sate length by selectively batching a subset of operations in the transformer model but processing requests in the batch individually for a subset of operations in the transformer model. In one embodiment, the operation to be processed individually is an attention operation of an encoder or a decoder of the transformer model. By selective batching, the inference system can allow batching operations to be performed for a batch of requests with variable input or target length or internal state length to utilize the parallel computation capabilities of hardware accelerators while preventing unnecessary computations that occur for workarounds that restrain the data of a batch of requests to a same length.
Owner:FRIENDLIAI

Coast erosion rate prediction method

The invention provides a coastal erosion rate prediction method, and belongs to the technical field of coastal erosion, and the method comprises the steps: building a three-dimensional seabed grid model, solving a hydrodynamic field through depth average simplification and GPU parallel calculation, and calculating sediment flux distribution based on shear stress discrimination and a high-order windward format. Outputting an erosion mode category and a local erosion strength coefficient by using a coastline erosion feature recognition model comprising a Josephh ring screening layer and a self-organizing mapping projection layer, starting adaptive grid encryption when the local erosion strength exceeds a threshold value, and calling a corresponding parameter set according to an erosion mode to calculate a seabed elevation change and a coastline erosion rate; and a wave energy spectrum reconstruction algorithm is adopted to generate future wave sequence cycle prediction, so that the technical problem that calculation precision and calculation efficiency are difficult to consider in coast erosion rate prediction under a complex wave power condition is solved.
Owner:SHANDONG MARINE FORECASTING & DISASTER REDUCTION CENT

Efficient matrix engine architecture based on RISC-V matrix extension and calculation method

The invention provides a high-efficiency matrix engine (RVME) architecture based on RISC-V matrix extension and a calculation method, and the architecture comprises an instruction buffering and decoding module, a matrix loading / storage module, a matrix register file, a parallel outer product array and an element-by-element operation module; the matrix register file comprises a Tile register and an Acculator register; the storage modules are respectively used for storing an input matrix and an accumulation result and supporting efficient data access and parallel computing; the matrix loading / storage module significantly improves the data loading efficiency through cache line alignment and matrix transposition optimization; the instruction buffering and decoding module cooperates with a main processor through a reordering buffer area and an instruction buffer area to ensure efficient scheduling and execution of instructions. The parallel outer product array is adopted to replace a traditional systolic array, the idle period in the calculation process is eliminated through multicast data flow scheduling and a ping-pong buffer read-write mechanism, and matrix multiplication and addition operation with high calculation utilization rate and low delay is achieved.
Owner:SHANGHAI JIAOTONG UNIV

User input information processing method and system for model reasoning stage

The invention discloses a user input information processing method and system for a model reasoning stage, and belongs to the technical field of artificial intelligence, and the method comprises the steps: obtaining a plurality of complete prompts inputted by a user; periodically acquiring hardware resource information and current task information; determining hardware resource allocation information corresponding to the complete prompt; pre-executing parallel computing operation on the complete prompt according to the hardware resource allocation information in the pre-filling stage; according to the hardware resource allocation information, calling the first type of computing equipment with the computing power reaching a preset computing power threshold value in the feed-forward network stage to read the intermediate characteristics from the preset resource pool for decoding and storage operation, wherein the computing power of the first type of computing equipment reaches the preset computing power threshold value; and calling the second type of computing equipment with the memory access performance meeting the preset memory access performance condition in the attention stage to read data from the preset resource pool to perform attention computing storage operation, obtaining a model reasoning result and outputting the model reasoning result, thereby solving the technical problem of hardware resource waste, and achieving the technical effects of improving the resource utilization rate and reasoning efficiency.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Business receipt and invoice data matching method and system

The embodiment of the invention discloses a business receipt and invoice data matching method and system, and the method comprises the following steps: accessing original business receipt data and invoice data from a plurality of heterogeneous data sources through a uniform interface, and carrying out the data cleaning according to preset mapping, merging, splitting and conversion rules, generating a standardized matching flow order; adopting streaming data loading and a Worker parallel computing mechanism to carry out multi-dimensional matching on business receipt and invoice data in the standardized flow bill, and generating a relation matrix representing a corresponding relation between the receipt and the invoice based on a matching result; and automatically checking the relation matrix by calling a rule engine and a machine learning model, marking the matching result of the business document and the invoice which are checked to be abnormal, pushing the matching result to manual auditing, and feeding back the manual auditing result to a knowledge base to optimize the subsequent matching and checking rule. According to the embodiment of the invention, the matching efficiency of the business document and the invoice data and the accuracy of the matching result can be improved.
Owner:百望股份有限公司

Long text abstract generation method and device

The invention provides a long text abstract generation method and device, and belongs to the technical field of natural text processing, the method comprises the following steps: using Euclidean norm to carry out importance sorting on Tokens and carrying out compression storage according to a sparse rate, so that global key information is completely reserved and memory occupation is obviously reduced; in the decoding stage, local attention scores and global attention scores are calculated in parallel, entropy differences are mapped into fusion weights through Sigmoid by combining temperature adjusting parameters, and dynamic balance of local details and long-distance dependence is achieved. The local key value pairs and the global key value pairs are subjected to weighted integration based on the fusion weight, a continuous semantic spectrum is formed in a single decoding layer, splicing breakage caused by traditional partitioning is eliminated, the problems of input limitation and semantic splitting are effectively relieved, the context length capable of being processed by a model is expanded under the condition that the calculation amount is not remarkably increased, and the method has the advantages of being simple in structure and convenient to operate. And local and global context information is adaptively fused, so that the accuracy and continuity of the abstract are effectively improved.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719 +1