Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

633 results about "Computation graph" patented technology

In simple terms, a computation graph is a DAG in which nodes represent variables (tensors, matrix, scalars, etc.) and edge represent some mathematical operations (for example, summation, multiplication). The computation graph has some leaf variables.

Federated Distributed Computational Graph Platform for Advanced Robotic Integration in Precision Oncological and Gene Therapies

A federated distributed computational system enables secure oncological therapy optimization through robotic integration. The system establishes a distributed graph architecture with secure communication channels connecting computational nodes, implementing encryption protocols for cross-institutional data exchange. Each node contains processing capabilities for fluorescence-guided imaging, uncertainty quantification, and expert knowledge integration while maintaining hierarchical knowledge graphs of oncological biomarkers, interventions, and outcomes. The system coordinates domain-specific knowledge through token-space communication and implements an advanced robotic integration system for surgical interventions using spatiotemporal tumor mapping, multi-modal fluorescence imaging, surgical robot coordination, and space-time stabilized mesh management. Key capabilities include wavelength-specific multi-modal fluorescence detection, combined epistemic and aleatoric uncertainty estimation, tensor-based data integration with adaptive dimensionality control, and light cone search for adaptive treatment optimization—all while maintaining strict privacy controls.
Owner:QOMPLX INC

Federated Distributed Computational Graph Platform with Advanced Multi-Expert Integration and Adaptive Uncertainty Quantification for Precision Oncological Therapy

A federated distributed computational system enables secure oncological therapy optimization through multi-expert integration and advanced uncertainty quantification. The system implements a multi-expert integration framework that coordinates domain-specific knowledge through token-space communication for precision oncological treatment, while maintaining secure cross-institutional data exchange. The architecture coordinates multi-scale spatiotemporal synchronization across computational nodes, with each node containing local processing capabilities for fluorescence-guided imaging, uncertainty quantification, and expert knowledge integration. Through a distributed graph architecture, the system enables advanced fluorescence imaging with wavelength-specific targeting, multi-level uncertainty estimation combining epistemic and aleatoric approaches, and multi-scale tensor-based integration with adaptive dimensionality control. The system implements light cone search and planning for adaptive treatment strategy optimization, enabling medical institutions and research organizations to collaborate on complex oncological therapy projects while maintaining strict data privacy controls.
Owner:QOMPLX INC

Neural network large model efficient reasoning method based on multiple GPGPUs

The invention belongs to the technical field of artificial intelligence and high-performance computing, and particularly relates to a neural network large model efficient reasoning method based on multiple GPGPUs. The method aims to solve the problems of high communication overhead, non-uniform load, low resource utilization rate, high data transmission delay and the like among multiple processors. Dividing a calculation task into a plurality of sub-graphs through static analysis and mixed granularity partitioning of a model calculation graph; distributing the sub-graphs to the optimal GPGPU based on a weighted cost function in combination with heterogeneous resource perception and a dynamic mapping strategy; a global pipeline scheduling plan is constructed by using communication topology perception, and calculation and communication overlap are maximized; data are loaded in advance through a host side hierarchical caching and asynchronous prefetching mechanism, and transmission delay is hidden; multi-stream concurrent execution and event-based lightweight synchronization are adopted on each GPGPU, so that waiting overhead is reduced. According to the method, the reasoning delay can be remarkably reduced, the throughput and the hardware utilization rate are improved, and the method has good adaptivity and expandability.
Owner:BEIJING TOPMOO TECH

Model performance estimation method and device and computer equipment

The invention relates to the technical field of artificial intelligence chips, and discloses a model performance estimation method and device and computer equipment, and the method comprises the steps: determining model configuration data and candidate distributed strategies of a target model; converting the original calculation graph corresponding to the single-card deployment state based on the model configuration data and the strategy configuration data of the candidate distributed strategies to obtain a distributed overhead calculation graph corresponding to the multi-card deployment state; and performing performance estimation on the candidate distributed strategy according to the basic overhead and the additional overhead in the distributed overhead calculation graph to obtain strategy performance data of the target model under the candidate distributed strategy, thereby realizing conversion of a single-card model into a multi-card model. And based on the multi-card model, simulation calculation of the multi-card interconnection mode is realized on the premise of limited hardware resources, so that the influence of a communication operator and a topological structure corresponding to the multi-card interconnection mode in the overall operation of the model is reflected, and the upper limit of the model performance can be accurately evaluated in a simulator verification stage before silicon is applied.
Owner:SHANGHAI BIREN TECH CO LTD

Multi-view clustering method and device

The invention relates to the technical field of multi-view clustering, in particular to a multi-view clustering method and device, and can solve the problem that the overall effect of an existing method in a large-scale clustering task is limited due to the fact that the existing method has problems in the aspects of calculation efficiency, robustness and multi-view information integration to a certain extent. The method comprises the following steps: dynamically learning an anchor matrix and a projection matrix for each view, and constructing a bipartite graph to generate a similarity matrix; calculating a graph Laplacian matrix based on the similarity matrix of each view, and extracting spectrum embedding; the spectrums of multiple views are embedded and stacked into a third-order tensor, and cross-view shared information is extracted by using a low-rank tensor constraint; multi-view atlas embedding is aligned through a spectrum rotation technology, and a discrete clustering indication matrix is directly output.
Owner:CHANGZHOU UNIV

Heterogeneous computing power cooperative scheduling system and method for mixed precision training

The invention discloses a heterogeneous computing power cooperative scheduling system and method for mixed precision training, and belongs to the technical field of artificial intelligence computing. The system comprises a computational graph analysis and operator portrait module which is used for analyzing and dividing a model computational graph and extracting operator features; the heterogeneous hardware capability sensing and matching module is used for managing performance files and real-time states of heterogeneous hardware in the cluster and matching optimal execution hardware for each calculation partition; and the data flow coordination and pipeline parallel controller is used for generating a global execution plan, managing cross-device data dependence and communication and calculating overlapping optimization execution efficiency through communication. According to the method, the problem of low scheduling efficiency of mixed precision training in a heterogeneous environment is solved, automatic and accurate mapping from a calculation task to heterogeneous hardware is realized, the training speed is remarkably improved, the training cost is reduced, and the overall resource utilization rate of a cluster is improved.
Owner:HANHOU (BEIJING) TECH CO LTD

Federated Distributed Computational Graph Platform for Oncological Therapy and Biological Systems Analysis With Neurosymbolic Deep Learning

A federated distributed computational system enables secure drug discovery and resistance tracking through hybrid simulation capabilities. The system implements a hybrid simulation orchestrator that coordinates molecular dynamics simulations with machine learning models for drug discovery analysis, while maintaining secure cross-institutional data exchange. The architecture coordinates multi-scale spatiotemporal synchronization across computational nodes, with each node containing local processing capabilities for molecular dynamics simulation and resistance pattern detection. Through a distributed graph architecture, the system enables real-world clinical data integration, resistance evolution tracking, and multi-scale tensor-based analysis with adaptive dimensionality control. The system implements real-time drug response prediction through multi-modal data analysis, enabling pharmaceutical companies and research institutions to collaborate on complex drug discovery projects while maintaining strict data privacy controls.
Owner:QOMPLX INC

Artificial intelligence model training resource adaptive distribution system

The invention belongs to the technical field of artificial intelligence, and discloses an artificial intelligence model training resource adaptive distribution system. The method comprises the following steps: acquiring and calculating graph structure data and hardware resource state data in real time, and calculating a data reuse rate and generating a candidate operator fusion scheme by constructing an operator execution time sequence constraint matrix and identifying an operator cluster of data locality characteristics; a resource competition hotspot prediction mechanism is introduced, memory bandwidth occupation fluctuation characteristics are analyzed, a resource conflict probability is calculated for a fusion scheme, and a dynamic balance optimization model of fusion income and resource conflicts is constructed. An optimal operator fusion decision sequence and a resource allocation strategy are generated through iterative solution, and accurate dynamic adjustment of computing resources in the training process is achieved. The training efficiency and the resource utilization rate are improved, the energy consumption is reduced, and the system stability is enhanced.
Owner:YANGZHOU HUAZHISHENG INFORMATION TECHNOLOGY CO LTD

End side malicious behavior security detection method

The invention relates to an end-side malicious behavior security detection method, which comprises the following steps: acquiring multi-source data, carrying out standardization processing on the multi-source data, generating a unified quintuple sequence, and constructing a quaternary dynamic behavior graph; respectively extracting time sequence features, semantic features and context features by utilizing a quintuple sequence, calculating confidence coefficients of cross-modal association, mapping the confidence coefficients to a quaternary dynamic behavior graph, and generating an individual behavior graph, a current behavior graph and a group baseline graph; and performing projection processing on the individual behavior graph, the current behavior graph and the group baseline graph to calculate a deviation degree score and a structural anomaly contribution degree of graph-level comprehensive deviation characterization, and triggering a grading response strategy by taking the deviation degree score and the structural anomaly contribution degree as double indexes of safety detection to solve an anomaly root. According to the method, internal and external threats, especially abnormal behaviors and potential risks which are difficult to find by traditional security defense means, are detected by analyzing behavior modes of users and entities.
Owner:SHAOGUAN DATA IND RESEARCH INSTITUTE

Heterogeneous calculation-based large model AI reasoning acceleration and deployment method

The invention provides a large model AI reasoning acceleration and deployment method based on heterogeneous calculation, and relates to the technical field of AI.The method comprises the steps that a neural network model and a reasoning request are obtained, and a calculation graph is analyzed; decomposing a calculation task into sub-tasks and marking calculation characteristics; establishing an adaptive mapping relationship between the subtasks and the heterogeneous computing units; constructing a cross-computing-unit data flow path and executing a scheduling strategy; and executing reasoning calculation and returning a result. According to the method, the heterogeneous computing resource characteristics can be fully utilized, the reasoning delay is reduced, the computing resource utilization efficiency is improved, and efficient deployment of the large model AI is realized.
Owner:BEIJING TRUSTFAR TECH CO LTD

Tensor segmentation and mapping method of neural network operator on wafer chip

The invention provides a tensor segmentation and mapping method of a neural network operator on a wafer chip, and the method comprises the steps: expanding a forward calculation graph into a training calculation graph containing forward propagation, back propagation and gradient updating, and employing different operators in the calculation of these stages; carrying out dimension segmentation on the input tensor of each operator in the training calculation graph to generate a plurality of global candidate segmentation strategies; aiming at each candidate segmentation strategy, enumerating feasible resource mapping strategies of all operators, and combining to generate candidate design points; and screening an optimal design point through cost evaluation, and outputting a corresponding candidate segmentation strategy and a resource mapping strategy. According to the method, forward and backward operators in a neural network training process are decoupled through expansion of a calculation graph, and a segmentation scheme for avoiding tensor copying in forward and backward calculation in the training process and a resource mapping strategy corresponding to the segmentation scheme are explored, so that tensor copying is avoided, global resource occupation is reduced, and the training efficiency is improved. Therefore, the same wafer-level chip can support larger model training.
Owner:TSINGHUA UNIVERSITY

Compatibility expansion system based on PyTorch framework

The invention relates to the technical field of deep learning, in particular to a PyTorch framework-based compatibility expansion system, which comprises a cross-framework model converter, a heterogeneous hardware abstraction layer, a hybrid computational graph execution engine, an intelligent distributed trainer and a self-adaptive optimizer, the cross-frame model converter converts input model formats of different frames into PyTorch executable formats, the heterogeneous hardware abstraction layer supports rear ends of various hardware and automatically selects an optimal calculation path, and the hybrid calculation graph execution engine fuses dynamic graph flexibility and static subgraph optimization and supports dynamic control function execution flow. An intelligent distributed trainer automatically selects a parallel scheme and is compatible with multi-protocol communication, a self-adaptive optimizer performs dynamic optimization based on hardware characteristics and a PyTorch tensor, a cross-frame model converter supports model conversion of multiple deep learning frames, the cost of migration among different frames by a user is reduced, and the universality of the PyTorch frame is improved.
Owner:SHANGHAI KUANFAN TECH CO LTD

Heterogeneous computing power adaptive compiling method and system for large model

The invention provides a large-model-oriented heterogeneous computing power adaptive compiling method, which comprises the following steps that: a user inputs a trained large model through a system interface, and a system front-end conversion module analyzes a computational graph of the model and converts the computational graph into an intermediate representation based on a unified operator description language (UDL); the system hardware sensing module automatically detects and extracts hardware feature fingerprints of at least one piece of target hardware; based on the unified operator description language UDL intermediate representation and the hardware feature fingerprint, automatically generating an optimization adaptation rule oriented to at least one piece of target hardware; wherein the basis of parameterized filling comprises specific parameters of hardware feature fingerprints and optimized attribute tags carried in an intermediate representation of a unified operator description language (UDL); and generating and deploying multiple back-end codes. The method has the beneficial effects that intelligent compiling based on hardware features can be realized, and the deployment efficiency and the operation performance of a large model in a complex heterogeneous computing power cluster are remarkably improved.
Owner:SHENZHEN XINGSHENG DIGITAL TECH CO LTD

Three-dimensional point cloud segmentation method and system based on bidirectional fusion of point cloud and aerial view

The invention provides a three-dimensional point cloud segmentation method and system based on bidirectional fusion of a point cloud and an aerial view, and belongs to the technical field of three-dimensional scene environment perception, and the method comprises the steps: obtaining to-be-segmented original point cloud data; processing the obtained original point cloud data by using a pre-trained segmentation model to obtain a segmentation result; wherein the segmentation model comprises an encoder module, a backbone network and a fusion segmentation head. According to the method, a bidirectional fusion mechanism of the point cloud and the aerial view is provided, through interactive fusion of the point cloud and the BEV features in multiple stages, the expression ability of the model for local and global features is improved, the problems of fuzzy semantic boundary, difficulty in small target recognition and the like are remarkably improved, and the segmentation precision is obviously superior to that of an existing BEV method. Compared with a voxel method and a multi-representation fusion method, the method has the advantages that the weight of a calculation graph is kept light in structural design, compared with the voxel method and the multi-representation fusion method, the method has remarkable advantages in the aspects of parameter quantity and reasoning speed, precision and real-time performance are better considered, and the method is suitable for scenes with extremely high efficiency requirements such as automatic driving.
Owner:BEIJING JIAOTONG UNIV +1

Model training method, device, system and equipment

The invention relates to a model training method, device, system and equipment. In the process of executing at least part of the forward process for the Nth batch of training samples, segmenting the dynamically generated calculation graph into a plurality of sub-forward processes, and obtaining a plurality of sub-reverse processes corresponding to the plurality of sub-forward processes; after the forward process of the Nth batch of training samples is executed and at least part of the forward process of the Pth batch of training samples is executed; in at least one time period, a calculation unit is used for executing one sub-forward process in at least part of forward processes of the Pth batch of training samples and one sub-reverse process in multiple sub-reverse processes of the Nth batch of training samples in parallel, and one of the sub-forward process and the sub-reverse process executed in parallel belongs to calculation operation, and the other belongs to calculation operation; the other belongs to communication operation. Therefore, in the model training process, at least part of communication cost overhead can be masked.
Owner:BEIJING WUWEN CORE TECH CO LTD

Artificial intelligence chip pre-silicon verification system and method

The invention provides an artificial intelligence chip pre-silicon verification system and method, and relates to the technical field of artificial intelligence chips, and the system comprises a model importing layer which is used for importing a current model; the model analysis layer is used for analyzing the current model to obtain a calculation graph; the distributed strategy configuration layer is used for configuring a distributed topology strategy of a hardware simulator of the artificial intelligence chip based on the distributed strategy of the current model, and adjusting a computational graph of the current model to enable the computational graph to be matched with the distributed topology strategy; the performance prediction layer is used for determining the operation performance data of the current model in the artificial intelligence chip based on the operation performance data of each operator; or determining a performance prediction result of the current model in the artificial intelligence chip based on the performance prediction result of each operator; and the operator implementation layer is used for controlling the hardware simulator to run each operator in the computational graph. According to the system and the method provided by the invention, the end-to-end whole-process verification from model input to result output is carried out on the artificial intelligence chip.
Owner:SHANGHAI BIREN TECH CO LTD

Smart home distributed heterogeneous computing power collaborative reasoning method

The invention relates to the technical field of smart home distributed heterogeneous computing power cooperative reasoning methods, and particularly discloses a smart home distributed heterogeneous computing power cooperative reasoning method. The objective of the invention is to solve the problems that heterogeneous equipment in an existing smart home system is uneven in computing power utilization, depends on a central scheduling node, is difficult to ensure privacy security and is insufficient in dynamic adaptability. The method comprises the following steps: constructing a dynamic computing power portrait of household equipment and updating the dynamic computing power portrait in real time; performing semantic analysis and computational graph decomposition on the intelligent reasoning task to generate a fine-grained reasoning unit with computing power and delay constraints; based on a decentralized broadcast protocol and a weighted Hungary algorithm, the reasoning unit is matched to the optimal local device; executing cross-device collaborative reasoning through an encrypted publishing-subscribing mechanism; and finally aggregating a result and returning a log for optimization. According to the technical scheme, efficient cooperation of family heterogeneous computing power can be achieved, end-side reasoning delay is lower than 200 milliseconds, the comprehensive utilization rate of resources exceeds 85%, and meanwhile data privacy and system robustness are guaranteed.
Owner:NINGXIA HUIWAN NETWORK TECH CO LTD

Oracle bone structure identification method of glyph graph isomorphic network

The invention belongs to the technical field of character pattern recognition, and particularly provides an oracle structure recognition method of a font pattern isomorphic network, which comprises three stages of font pattern structure feature extraction, font pattern isomorphic network model and oracle font structure recognition. Through oracle font skeleton extraction, oracle font skeleton singular point detection, skeleton burr and bifurcation optimization, and key point extraction as a node communication path as an edge, an oracle font graph structure is established, and graph structure nodes, edges and graph level features are calculated. The font graph isomorphic network integrates the advantages of a graph isomorphic network and dynamic edge and graph embedding, and learns vector representation of multi-level features such as nodes, edges and graphs of the oracle font graph structure so as to train and identify the font structure of the oracle. The method can be used for structure matching recognition of oracle character pattern images, has the advantage of being high in recognition accuracy, and is suitable for recognition of character pattern structures such as gold texts and seal scripts and recognition of social networks, biological information and molecular structures.
Owner:ANYANG NORMAL UNIV

Heterogeneous processor-oriented deep neural network reasoning task dynamic scheduling method and system

The invention discloses a heterogeneous processor-oriented deep neural network reasoning task dynamic scheduling method and system, and belongs to the technical field of computer systems. The method comprises the steps that a DNN model is converted into a computational graph in a user space, and the computational graph is divided into a plurality of fine-grained micro-tasks through key path analysis; monitoring resource states of the CPU and the GPU in a kernel space in real time; dynamically selecting an optimal processor for each microtask in a kernel layer based on a cost model comprising execution time, queue length and utilization rate; and executing the micro-task on the corresponding processor according to a scheduling result, and triggering cross-processor task dynamic migration when detecting that the load is unbalanced, including context storage, data transmission and execution recovery. The system adopts a user space and kernel space collaborative architecture to realize the method. Through fine-grained division, kernel-level real-time scheduling and a dynamic migration mechanism, the reasoning efficiency, the resource utilization rate and the system stability of the heterogeneous system under a complex load are effectively improved.
Owner:EAST CHINA NORMAL UNIV

Neural network model processing method and device, equipment and storage medium

The invention provides a neural network model processing method and device, equipment and a storage medium, and relates to the technical field of computers, in particular to the technical field of Shenzhou networks, data processing and the like. According to the specific implementation scheme, a target neural network model is converted into a calculation graph; based on homeomorphic transformation rules, converting a target sub-graph in the calculation graph into a target sketch; the number of operators in the target sketch is smaller than the number of operators in the target sub-graph; and under the condition that the target sketch is contained in the sketch template set, operator fusion is performed on the target sub-graph, so that the access stock aiming at m operators in the target sub-graph is reduced to the access stock of n fused operators in the reasoning stage, n and m are positive integers, and n is smaller than m.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Intelligent sensing method and system for corrosion state of offshore wind power steel structure

The invention relates to the technical field of corrosion monitoring and evaluation, in particular to an intelligent sensing method and system for the corrosion state of an offshore wind power steel structure. The method comprises the following steps: monitoring data by arranging a time sequence; constructing an initial neural network model taking the time sequence monitoring data as input and the corrosion state parameters as output; acquiring a physical law mathematical expression for analyzing the corrosion process, converting the physical law mathematical expression into a differentiable computational graph, comparing the output of the computational graph with the output of the initial neural network model, and taking the difference between the output of the computational graph and the output of the initial neural network model as a physical constraint term; and adding the physical constraint item into a loss function of the initial neural network model to form a mixed loss function, training the initial neural network model containing the mixed loss function by using the time sequence monitoring data until convergence, and obtaining a physical information neural network model. The corrosion prediction accuracy of the offshore steel structure can be improved.
Owner:CHINA MACHINERY (SHANXI) INSPECTION & TESTING CO LTD +1

Dynamic operator-oriented incremental compiling optimization method and system

The invention discloses a dynamic operator-oriented incremental compilation optimization method and system, belongs to the technical field of deep learning, and aims to solve the technical problem of how to provide a fine-grained, reusable and extensible compilation strategy for dynamic operator-oriented incremental compilation optimization. Comprising the following steps: extracting attributes of a dynamic operator and constructing a dynamic operator hash signature according to the extracted attributes; dividing the calculation graph into a static sub-graph and a dynamic sub-graph, establishing a mapping table of a dynamic operator hash signature and a compiling kernel, and caching the mapping table in an incremental compiling cache; solving and instantiating the dynamic sub-graph through a dynamic specifying engine, generating a specialized sub-graph, querying whether a matched compiling result exists in an incremental cache or not, and updating the incremental cache; and for the static sub-graph which is not changed, directly multiplexing the last compiling result of the static sub-graph, executing cross-sub-graph fusion optimization work between the dynamic sub-graph and the static sub-graph which are subjected to specialized processing, and submitting a fusion execution plan to a runtime engine.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Mitigation of computational graph lineage growth

Example implementations include a system and a computer-implemented method for breaking computational graph lineages configured for identifying at least one mutable variable of a given database table in a block of structured query language (SQL) instructions, the block containing at least one loop, the at least one mutable variable including one or more of a first variable declared and initialized outside a given loop and used and updated in the given loop or a second variable defined inside the given loop and assigned value as a result of an operation in the given loop. The implementations further include inserting a breakage of a computational graph lineage of at least one of the first variable or the second variable.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Inference method and device, equipment and storage medium

The invention provides an inference method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, large language models and the like. The method comprises the steps of determining target metadata of a target operator through a predefined function interface in response to an identified extended dependency package corresponding to the reference target operator in an inference framework; a calculation graph is generated according to the target metadata and native metadata, and the native metadata is metadata of other operators except the target operator in the reasoning framework; and reasoning on the target hardware according to the input data and the calculation graph to obtain a reasoning result. According to the method, the operator and the frame are decoupled, meanwhile, the fracture condition of the computational graph is avoided, and the compatibility and reasoning efficiency of the reasoning frame are improved.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Space-time constraint end-to-end automatic driving system and method based on unified VLA model

The invention relates to the technical field of artificial intelligence and automatic driving, and discloses a space-time constraint end-to-end automatic driving system and method based on a unified VLA model, and the system comprises a unified encoder, a computing power management unit, a core neural network and an output decoder. A gating operator embedded into a core neural network is used for calculating characteristic information entropy and computing power consumption in real time, and through topology reconstruction logic based on entropy and computing power joint constraint, computational graph collapse operation is dynamically triggered at an instruction level, so that network reasoning depth is adaptively expanded along with physical time constraint, and the network reasoning accuracy is improved. According to the method, isomorphic mapping of the topological structure of the calculation model and the physical time constraint is established, the deterministic delay upper limit under the extreme working condition is determined, the problem of dimension mismatching of an existing static calculation graph under the hard real-time constraint is solved, and optimal bit allocation of calculation resources in the time dimension is achieved.
Owner:GUANGZHOU SMART BODY TECH CO LTD

Cloud edge-end collaborative deployment method and system of image reasoning multi-mode neural network

The invention provides a cloud edge-end collaborative deployment method and system for an image reasoning multi-mode neural network, and is applied to the field of artificial intelligence, and the method comprises the steps: abstracting a preset multi-mode neural network, and obtaining an original calculation graph; based on the calculation node set and the data dependence edge set, splitting the original calculation graph to obtain a plurality of connected sub-graphs, the categories of the plurality of connected sub-graphs including a single-modal processing sub-graph, a cross-modal interaction sub-graph and a task specific sub-graph; based on the topological sequence of the plurality of connected sub-graphs, determining segmentation points of each connected sub-graph; dividing the plurality of connected sub-graphs according to the segmentation points to obtain a plurality of segmentation parts; and integrally dividing the multi-modal neural network based on the plurality of segmentation parts according to a genetic algorithm, and respectively deploying the integrally divided multi-modal neural network to a terminal device, an edge server and a cloud platform. According to the invention, a flexible and efficient model segmentation and deployment mechanism can be realized.
Owner:BEIJING UNIV OF POSTS & TELECOMM +2

Sugar net image grading method and system based on multi-modal decoupling knowledge distillation

The invention discloses a sugar net image grading method and system based on multi-modal decoupling knowledge distillation, and belongs to the technical field of image processing and artificial intelligence. The method constructs a teacher-student architecture: a teacher model generates an ordered DR language prototype and calculates an image-text similarity matrix through a ViT image encoder and a text encoder with a Prompt module; according to the student model, image features are extracted by using light-weight EfficientNet-B0. In the training stage, a decoupling distillation mechanism is designed, target class KL loss, non-target class KL loss and grade sorting loss generated by Prompt are jointly optimized, and migration of multi-modal ordered knowledge is achieved; in the reasoning stage, only student models are deployed for efficient grading. According to the method, the problems of boundary fuzziness, long-tail distribution and sequential modeling in DR classification are solved, lightweight deployment is realized while the classification precision is improved, and the method is suitable for clinical edge equipment.
Owner:TIANJIN UNIVERSITY OF TECHNOLOGY

End-side large language model reasoning acceleration optimization method based on NPU

The invention discloses an end-side large language model reasoning acceleration optimization method based on NPU, and belongs to the technical field of computers. According to the method, the speculative decoding technology is applied to a decoding stage to reduce the influence of NPU bandwidth bottleneck on the decoding stage, and an end-side large language model is partitioned to execute pre-filling calculation and NPU calculation graph loading for the decoding stage in parallel, so that the loading delay and the NPU calculation delay are overlapped, and the overall time and memory overhead are reduced.
Owner:PEKING UNIV

Diffusion model GPU reasoning accelerator and diffusion model GPU reasoning acceleration method

The invention provides a diffusion model GPU reasoning accelerator and a diffusion model GPU reasoning acceleration method. The accelerator comprises a static feature extraction module which is configured to extract static statistical features, not changing along with cue words, of normalized operators of all layers in a diffusion model in an offline stage; the dynamic decision-making module is configured as follows: dynamic statistical characteristics are generated in the online stage, and a decision-making instruction of whether to reuse the statistical characteristics of the preorder time step or not is generated in combination with the static statistical characteristics and the dynamic statistical characteristics; the reuse layer normalization generation module is configured to multiplex statistical characteristics of preorder time steps and optimize a layer normalization operator into a reuse layer normalization operator; the attention-normalization fusion module is configured to combine multi-head attention with the reuse layer normalization operator to generate a fusion kernel; and the computational graph optimization module is configured to replace a layer normalization operator with a reuse layer normalization operator, replace the sub-graph fusing multi-head attention and layer normalization with a fusion kernel, and optimize the computational graph.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Dynamic cybersecurity scoring using traffic fingerprinting and risk score improvement

A system for dynamic cybersecurity scoring using traffic fingerprinting and score improvement, that uses a web crawler that sends message prompts to external hosts and receives responses from external hosts, a time-series data store that produces time-series data from the message responses, and a directed computational graph module that analyzes the time-series data to produce a weighted score representing the overall cybersecurity state of an organization.
Owner:QOMPLX INC