Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1182results about "Computer simulations" patented technology

Parameter-efficient large-language fine-tuning federated learning framework

Provided in the present invention is a parameter-efficient large-language fine-tuning federated learning framework, comprising the following steps: performing modeling on LoRA adapters of different edge clouds; since different weights exhibit different average performances on the LoRA adapters, using singular values to quantify the importance of the weights, and therefore, before each round of independent training of the LoRA adapters using N edge clouds, using a matrix singular value to decompose a BA matrix in the LoRA adapter for each trainable weight; configuring heterogeneous LoRA adapters on the basis of the importance of the weights; and using different numbers of quantization bits to quantize a pre-trained model, and performing high-precision inverse quantization on the pre-trained model only when matrix multiplication is executed, wherein the pre-trained model is quantized to the maximum number of quantization bits on the basis of the memory budget of the edge clouds. The present invention has the following beneficial effects: the present invention determines the optimal fine-tuning model structure, thereby improving the performance of LLM fine-tuning, and adapts to heterogeneous and resource-constrained edge clouds.
Owner:FUDAN UNIVERSITY

Method and apparatus for implementing ai-ML in a wireless network

A method, an apparatus, and a computer readable medium for storing instructions are described for a user terminal and a base station for updating an AI / ML configuration in case of a handover. The method performed by a user equipment comprising operating a first AI / ML configuration in a coverage area of a first base station; receiving an AI / ML configuration information indicating a second AI / ML configuration; operating the second AI / ML configuration indicated by the AI / ML configuration information in the coverage area of a second base station. Operating a first AI / ML configuration comprises operating a first AI / ML Model or a first AI / ML Model with a first AI / ML Model configuration in the coverage area of the first base station. The first AI / ML Model is associated with a first AI / ML Model identifier and the first AI / ML Model configuration is associated with a first AI / ML Model configuration identifier. Operating the second AI / ML configuration comprises operating a second AI / ML Model or the first AI / ML Model with a second configuration in the coverage area of the second base station. The second AI / ML Model is associated with a second AI / ML Model identifier and the second AI / ML Model configuration is associated with a second AI / ML Model configuration identifier. Indicating the second AI / ML configuration comprises indicating a second AI / ML Model identifier and / or a second AI / ML Model configuration identifier.
Owner:HARFANG IP INVESTMENT CORP

Temporal dynamics simulation in matmul-free neural architectures

A method is provided for processing data in a neural network system. The method includes receiving input data; processing the input data through a first set of neural network layers configured to perform data processing using MatMul-free techniques to produce intermediate data; further processing the intermediate data through a second set of neural network layers configured to simulate spiking neural network (SNN) functionalities using MatMul-free techniques; and outputting a result based on the processed data from the second set of neural network layers.
Owner:LEPTUDE INC

Graph neural network execution on neural processing unit

Workloads for executing a graph neural network (GNN) may be divided among various processing units, such as a central processing unit (CPU) and a neural processing unit (NPU). The NPU may include a data processing unit (DPU) and a digital signal processor (DSP). The CPU may perform precomputation, model optimization, hardware optimization, and compilation. For example, the CPU may precompute a parameter matrix and use the parameter matrix as internal parameters of a GNN. The CPU may also perform node padding, approximation computation, or transfer of DSP operations to DPU to optimize the GNN. The CPU may also perform sparsity data compute and storage, vertical fusion of DSP operations and DPU operations, or data quantization to optimize performance of the NPU. The compiled GNN may be provided to the NPU, and the DPU and DSP may perform the operations in the compiled GNN to produce a prediction of the GNN.
Owner:INTEL CORP

User Behavior Modeling for Detecting and Containing Malicious Activity in a Storage System

An illustrative method includes monitoring operations performed with respect to a storage system by an entity using an identity associated with a particular role, the particular role providing the entity with a set of permissions associated with the storage system; determining, based on the monitoring, that one or more operations of the operations deviate from an expected activity profile associated with the role by more than a threshold; and performing, based on the determining that the one or more operations deviate from the expected activity by more than the threshold, a remedial action with respect to the entity.
Owner:PURE STORAGE INC

Large model dynamic compression optimization method and system based on sparse pruning

The invention relates to the technical field of large model algorithms, in particular to a large model dynamic compression optimization method and system based on sparse pruning, and the method comprises the steps: capturing original weight fluctuation data generated by resource fluctuation in reasoning, and obtaining sparse weight reference data through sparse processing; analyzing calculation complexity through model reasoning delay data, and separating reasoning delay amount caused by a model scale; dynamically controlling the model compression ratio within a preset performance range based on the delay amount and the sparse reference data, and collecting reasoning precision distribution data under different compression parameters; evaluating a model performance state under each parameter by means of a neural network simulation method, and generating a performance state simulation result; determining a model quality optimization compensation parameter based on a simulation result by combining resource fluctuation data acquired in real time in a compression process; and finally, the compression strategy is adaptively regulated and controlled through the compensation parameters, and collaborative optimization of model calculation complexity, reasoning precision and delay during dynamic change of hardware resources is realized.
Owner:NOVNET COMPUTING SYST TECH CO LTD

Automatic kernel network parameter optimization method

The invention relates to the technical field of parameter optimization, in particular to an automatic kernel network parameter optimization method, which comprises the steps of constructing an enhanced deep Q network model, and integrating the enhanced deep Q network model with a priority playback buffer area, a meta learning module, a Bayesian optimizer and a neural architecture search module; using performance index data to train an enhanced deep Q network model, the training process including using a priority playback buffer to store and sample empirical data, using a meta-learning module to perform task adaptation, and monitoring training indexes of multiple dimensions to evaluate the convergence state of the model; selecting a kernel parameter adjustment action according to the current state through the trained enhanced deep Q network model; executing the selected kernel parameter adjustment action, and evaluating a parameter adjustment effect based on the multi-target reward function; and updating the enhanced deep Q network model according to an evaluation result, wherein the priority playback buffer area and the Bayesian optimizer are utilized in the updating process.
Owner:GUANGZHOU CITY UNIV OF TECH

Workload balance with prompt and token routing for expert models

PendingUS20250356216A1Computer simulationsMixture of expertsWorkload
Systems and methods are provided for determining expert placement layouts and / or prompt and toke routing for a mixture of experts (MoE) model. In some instances, an expert workload distribution with respect to a plurality of experts is determined based on a gating neural network. In some instances, an expert placement layout with respect to a plurality of computing units is determined based on the determined expert workload distribution. In some instances, a two-level routing strategy is provided to first adaptively route an incoming prompt to a suitable computing device, then perform a token routing within that computing device to ensure workload balance between different computing units of that computing device.
Owner:BYTEDANCE TECHNOLOGY LTD

Artificially intelligent routing agent for routing portions of a task through multiple customized agents, and systems, devices, and methods of use thereof

This application describes, amongst other things, methods and systems for building and deploying agents. An example method includes obtaining orchestration data about a set of task-specific components selected to provide a response to the prompt, where each respective task-specific components in the set of task-specific components is configured to assist with a respective clinical task of the one or more clinical tasks. The method further includes, determining an order in which each respective task-specific components of the set of task-specific components should be utilized to prepare a complete response to the prompt that address the one or more clinical tasks based on the obtained orchestration data about the set of task-specific components. The method also includes, in accordance with the determined order, providing first data related to the prompt to a first task-specific component and receiving a first response from the first task-specific component.
Owner:TEMPUS AI INC

Intelligent routing and unified adaptation method for large language model

The invention discloses an intelligent routing and unified adaptation method for a large language model, and the method comprises the steps: defining all access details of the model through a declarative configuration file, and achieving the zero-code access of the large language model without writing any code for a newly-added model; a completely consistent calling interface is provided for an upstream application, the isomerism of all downstream large language models is shielded, and a unified request and response abstraction layer is constructed; through a strategy engine and a JSON path technology, complex streaming response including content thinking is precisely processed, and intelligent analysis and content extraction are carried out; dynamic configuration and intelligent strategy hot update are supported; intelligent routing of the model is realized, and an optimal large language model instance is dynamically selected; and meanwhile, enterprise-level governance capability is provided, governance functions such as fusing, degradation, current limiting and monitoring are integrated, and stability guarantee is provided for model calling. Therefore, the maintainability, the expandability and the user experience consistency of the system are comprehensively improved.
Owner:NANJING INFORMATION HIGH-SPEED RAILWAY RES INST OF SCI AND TECH

Heterogeneous computing power adaptive compiling method and system for large model

The invention provides a large-model-oriented heterogeneous computing power adaptive compiling method, which comprises the following steps that: a user inputs a trained large model through a system interface, and a system front-end conversion module analyzes a computational graph of the model and converts the computational graph into an intermediate representation based on a unified operator description language (UDL); the system hardware sensing module automatically detects and extracts hardware feature fingerprints of at least one piece of target hardware; based on the unified operator description language UDL intermediate representation and the hardware feature fingerprint, automatically generating an optimization adaptation rule oriented to at least one piece of target hardware; wherein the basis of parameterized filling comprises specific parameters of hardware feature fingerprints and optimized attribute tags carried in an intermediate representation of a unified operator description language (UDL); and generating and deploying multiple back-end codes. The method has the beneficial effects that intelligent compiling based on hardware features can be realized, and the deployment efficiency and the operation performance of a large model in a complex heterogeneous computing power cluster are remarkably improved.
Owner:SHENZHEN XINGSHENG DIGITAL TECH CO LTD

Large model energy consumption optimization method and device, computer equipment, readable storage medium and program product

The invention relates to a large model energy consumption optimization method and device, computer equipment, a computer readable storage medium and a computer program product. Comprising the steps of collecting hardware state data and task data of a large model; performing stage identification according to the task data, and determining a current stage; performing state prediction according to the hardware state data and the task data through a prediction model corresponding to the current stage to obtain target state information; generating optimization parameters of the current stage through a joint optimizer according to the target state information; and adjusting the resources of the large model according to the optimization parameters. In combination with stage perception and dynamic resource adjustment, a differentiated resource adjustment strategy based on different stages is realized, specific energy consumption pain points in an AI scene are solved, high power consumption of large model training, instantaneous fluctuation of reasoning requests and the like are reduced, sustainable and efficient operation of an AI system is ensured, the computing power demand and resource consumption of a large model are balanced, and the system performance is improved. And wide landing and development of a large model in various scenes are facilitated.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Methods and apparatus for hardware-aware machine learning model training

Methods, apparatus, systems, and articles of manufacture are disclosed for hardware-aware machine learning model training. An example apparatus includes a configuration determiner to determine a hardware configuration of a target hardware platform on which the machine learning model is to be executed, a layer generator to assign sparsity configurations to layers of the machine learning model based on the hardware configuration, and a deployment controller to deploy the machine learning model to the target hardware platform in response to outputs of the machine learning model satisfying respective thresholds, the outputs including a quantity of clock cycles to execute the machine learning model with the layers having the assigned sparsity configurations.
Owner:INTEL CORP

Method for establishing and deploying spiking neural network on hardware device

The invention relates to a method for deploying a spiking neural network to a hardware device. The method includes providing the spiking neural network, training the spiking neural network to obtain a trained spiking neural network, mapping neurons and synapses in the trained spiking neural network to corresponding components of the hardware device, simulating deployment of the trained spiking neural network on a hardware device using the obtained mapping, and deploying the trained spiking neural network to the hardware device using the mapping and the simulation. The training, mapping, simulation and deployment steps are executed by using hardware information of the hardware device; wherein the hardware information comprises at least one of hardware resource constraints, hardware connection constraints, dynamic range of hardware design parameters, characterization or statistics of neurons and synapses, reconfigurability, programmability, yield, computing resources, temporal characteristics, constraints on pre-processing, interfaces and peripheral devices, and available encoders and / or decoders.
Owner:INNATERA NANOSYSTEMS BV

Intermediate representation-based spiking neural network deployment method and deployment tool chain

The invention discloses a spiking neural network deployment method and a deployment tool chain based on intermediate representation, and the method comprises the steps: carrying out the model analysis and structure mapping operation of a trained ANN-ONNX model, generating an SNN-ONNX model with a pulse characteristic, carrying out the optimization of a weight parameter through the combination of transfer learning and a precision fine tuning strategy, and carrying out the optimization of the weight parameter. Therefore, the expression ability and adaptability of the converted model are improved, and the reasoning execution performance of the SNN model on a target platform is improved by adopting an optimization means; meanwhile, an SNN-oriented modular operator component library is constructed, and the portability, maintainability and expandability of operators among different hardware platforms are enhanced by adopting an abstract interface and a design mode of specifically realizing decoupling. According to the method, efficient deployment of the SNN model on a target hardware platform is achieved, performance fidelity of the model under the function equivalent condition is ensured, and the method has wide engineering application prospects in the fields of edge calculation, low-power-consumption intelligent terminals, brain inspiration type artificial intelligence and the like.
Owner:HANGZHOU DIANZI UNIV

Model training method and device, electronic equipment, storage medium and program product

The invention provides a model training method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of artificial intelligence. The method comprises the steps of obtaining a model training task and a model structure feature corresponding to the model training task; acquiring a current resource state; dividing the model training task into a plurality of sub-tasks; selecting a target resource scheduling strategy based on the model training task, the model structure feature and the current resource state; and distributing the subtasks to corresponding computing nodes based on the target resource scheduling strategy, wherein the computing nodes are used for executing the corresponding subtasks. According to the embodiment, the target resource scheduling strategy is selected based on the model training task, the model structure feature and the current resource state, the sub-tasks are efficiently allocated to the computing nodes through the target resource scheduling strategy, and the distributed training efficiency and the resource utilization rate are improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

LLC perception type physical information nested neural network parameter estimation method suitable for LLC resonant converter

The invention relates to the technical field of converters, in particular to an LLC perception type physical information nested neural network parameter estimation method suitable for an LLC resonant converter, and the method comprises the following steps: S1, constructing a continuous time state space model of the LLC resonant converter, and carrying out the discretization processing; s2, an LLC perception type physical information nested neural network is constructed, and the LLC perception type physical information nested neural network comprises a data reconstruction network and a physical information nested neural network which are connected and is used for online parameter identification of the LLC resonant converter; the data reconstruction network comprises a resonance state judgment layer, a K2 operation layer, a pseudo label generation layer, a constraint layer and a data reconstruction layer; the physical information nested neural network comprises a middle state mapping layer and a physical layer; defining LLC perception type physical information nested neural network input; s3, constructing a loss function; and S4, deploying and executing the LLC perception type physical information nested neural network model.
Owner:CHONGQING UNIV

Apparatus and method for monitoring optimization performance of deep learning compiler

Disclosed are an apparatus and method for monitoring the optimization performance of a deep learning compiler. The method includes calculating metric information for the evaluation of the performance of a compiler, calculating a score function value corresponding to resource optimization policy information, set in the artificial intelligence (AI)-based optimizer of the compiler, based on the metric information, and providing performance analysis results of the AI-based optimizer based on the score function value.
Owner:MOBILINT INC

Learning model evaluation support device and learning model evaluation support program

PendingJP2025145519A2D-image generationMachine learningModel selectionEvaluative learning
To provide a learning model evaluation support device and a learning model evaluation support program which allow for easily evaluating a characteristic of a learning model.SOLUTION: A learning model evaluation support device 100 includes a model information display unit 110, a model selection unit 120, and an output display unit 150. The model information display unit 110 selectively displays a plurality of learning models which are stored in a prescribed storage area and share an interface, on a display device 260. The model selection unit 120 selects at least one of the plurality of learning models displayed on the display device 260. The output display unit 150 displays an output result from at least one learning model on the display device 260.SELECTED DRAWING: Figure 17
Owner:SCREEN HOLDINGS CO LTD

Neural network model processing method and device, equipment and storage medium

The invention provides a neural network model processing method and device, equipment and a storage medium, and relates to the technical field of computers, in particular to the technical field of Shenzhou networks, data processing and the like. According to the specific implementation scheme, a target neural network model is converted into a calculation graph; based on homeomorphic transformation rules, converting a target sub-graph in the calculation graph into a target sketch; the number of operators in the target sketch is smaller than the number of operators in the target sub-graph; and under the condition that the target sketch is contained in the sketch template set, operator fusion is performed on the target sub-graph, so that the access stock aiming at m operators in the target sub-graph is reduced to the access stock of n fused operators in the reasoning stage, n and m are positive integers, and n is smaller than m.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Medical examination data intelligent processing and large model offline deployment method and system

The invention provides a medical physical examination data intelligent processing and large model offline deployment method and system. The method comprises the following steps: performing online configuration in an offline deployment basic environment, deploying a physical examination data processing algorithm module, downloading a pre-training large model file and a reasoning framework from a model warehouse, and generating a distributable mirror image through a packaging command; performing offline deployment on the distributable mirror image, transmitting the mirror image to a server without an external network through a physical medium, importing and loading the mirror image, and starting a container service; cleaning and standardizing the received physical examination data, matching the standardized data with a medical data matching library, and calculating the severity level; and preprocessing the unstructured medical text, calling an off-line deployed large model to perform semantic analysis and generate an explanatory suggestion, and integrating processing results to generate a standardized physical examination report. According to the invention, off-line deployment and application of a large model in an external network-free environment are realized, and the intelligent level and efficiency of physical examination data processing are improved.
Owner:XUNKANG INFORMATION TECH (SHENZHEN) CO LTD

Method, electronic device, and computer program product for deploying machine learning model

Embodiments of the present disclosure relate to a method, an electronic device, and a computer program product for deploying a machine learning model. The method includes: acquiring a machine learning model in accordance with an open neural network exchange format; converting the machine learning model to an intermediate representation using a multi-level intermediate representation method; and deploying a computation associated with the machine learning model to at least one computing device using the intermediate representation.
Owner:EMC IP HLDG CO LLC

Convolutional neural network-based building plane element recognition model construction method

The invention relates to the technical field of image recognition, in particular to a building plane element recognition model construction method based on a convolutional neural network, and the method comprises the following steps: obtaining an input APN sample; elements in the APN sample are labeled, and a labeled file is generated; converting the annotation file into a format required by YOLO, and normalizing the annotation file; performing data enhancement on the APN sample, expanding the training data volume, and outputting to obtain an element detection result and a feature vector; installing dependency and carrying out data configuration; setting command line model parameters and customizing training scripts; and carrying out model training and model reasoning. According to the method, the plane layout is abstracted based on the machine learning method, and the vector data type has high flexibility and deformation capability, so that the vector output which keeps the original plane image form unchanged can be converted into various objects according to the purpose of a user. And classifying and identifying the plane elements by using the GNN.
Owner:BEIJING UNIV OF CIVIL ENG & ARCHITECTURE

Method and storage medium for quantion aware retraining for graph-based neural network model

A method may comprise: adding a plurality of markers to a plurality of graph modules in a first neural network (NN) model in a form of a directed acyclic graph (DAG); generating calibration data by collecting input values and output values of each of the plurality of graph modules using the plurality of markers; determining, based on the calibration data, a scale value and an offset value applicable to the first NN model; generating, based on the scale value and the offset value, a second NN model including a weight parameter in integer format through quantization; and updating at least one parameter included in the second NN model by performing a quantization-aware retraining technique on the second NN model.
Owner:DEEPX CO LTD

Dynamic multi-model monitoring and validation for artificial intelligence models

The systems and methods disclosed herein receives artifacts generated using a first set of models within a multi-model superstructure. The multi-model superstructure includes a second set of models to test the first set of models. The multi-model superstructure dynamically routes the artifacts of the first set of models to one or more models of the second set of models by (i) determining a set of dimensions of the artifacts against which to evaluate the artifacts and (ii) identifying the models in the second set used to test the particular dimension. The second set of models then assesses each artifact against a set of assessment metrics. If an artifact fails to meet one or more assessment metrics, the second set of models generates actions to align the artifact with the set of assessment metrics.
Owner:CITIBANK N A

Digital in-memory computing architecture compiling and simulating method and tool chain

The invention discloses a digital in-memory computing architecture compiling and simulating method and a tool chain, and belongs to the technical field of in-memory computing. The method comprises the following steps: acquiring a target deep neural network model and a target in-memory computing architecture configuration parameter adapted to the target deep neural network model, the target in-memory computing architecture configuration parameters comprise at least one of the following parameters: the number of cores, the on-chip communication bandwidth between the cores, the size, the number and the bit width of a storage unit of each core, and the scale of a computing unit of each core; configuring a digital in-memory computing accelerator according to the target in-memory computing architecture configuration parameters; compiling the target deep neural network model into a target optimization instruction sequence based on a digital in-memory calculation accelerator; simulating and executing the target optimization instruction sequence, and performing performance evaluation according to performance indexes in the process of executing the target optimization instruction sequence.
Owner:BEIHANG UNIV

Neural network parallel scheduling-oriented single-instruction multi-thread processor micro-architecture device

A single-instruction multi-thread processor micro-architecture device oriented to neural network parallel scheduling comprises a front-end instruction fetching module, an instruction cache module, a decoding module, an arithmetic logic operation unit, a multiplication and division module, a memory access unit and a data cache module, and the micro-architecture device allocates a unique thread number for each thread. Threads are organized into thread groups, and in each period, the fair alternate arbiter selects an instruction from an instruction buffer of the thread group and sends the instruction to a subsequent decoding stage. The micro-architecture not only solves challenges faced by end-side equipment when processing high-performance calculation requirements such as neural network reasoning, but also provides an effective solution capable of reducing energy consumption and improving calculation efficiency through an innovative architecture design. The method is of great significance in promoting development of end-side AI application.
Owner:XI AN JIAOTONG UNIV

Edge device with built-in compiler for neural network models

A system includes a substrate on which a first memory, a neural processing unit (NPU) including a plurality of processing elements (PEs) with multiplier-accumulator circuits, a controller, and a second memory, and a central processing unit (CPU) are disposed. The CPU may be configured to execute a universal compiler to perform a conversion for a particular neural network model into a machine code executable by the NPU and store the machine code in the first memory or the second memory. When the particular neural network model, generated by one among a plurality of machine learning frameworks that are incompatible with each other, is received and stored in the first memory, the universal compiler may perform the conversion based on mapping information indicating mapping between elements of machine learning frameworks and functions or operations executable by the CPU or NPU.
Owner:DEEPX CO LTD

Ai-based generation of a computer program using compiler-gathered semantic information about target code

Techniques are described herein that are capable of performing AI-based generation of a computer program using compiler-gathered semantic information about target code. A user-generated request that requests information about target code is converted into an AI prompt, which requests that the AI model generate a computer program to determine the information. An AI model is caused to generate the computer program, which comprises configuring the computer program to determine, at runtime of the computer program, the information using semantic information about the target code gathered by a compiler and provided to the computer program by an API, by providing the AI prompt as an input to the AI model. A response to the AI prompt that includes the computer program is received from the AI model. Presentation of a representation of the computer program and / or automatic execution of the computer program against the target code is triggered.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC