Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1635results about "Computer simulations" patented technology

Large Language Model System

Construct multiple retrieval augmented generation (RAG) databases for different related fields as needed to increase the recording capacity of additional information. By referring to the entire text or image of the original document page where relevant content is described based on the chunks obtained from RAG retrieval, eliminate reference omissions. At the same time, reflect the background information and related information that may be described around the retrieved chunks in the context of the question text. Databaseize the records of the question text and the answer text to make them searchable, and provide a large language model using retrieval augmented generation that eliminates the need for costly and time-consuming regenerating answers. 【Solution means】The RAG database includes a page image acquisition means, a page text database recording means, a related feature vector extraction means, and a feature vector database recording means.
Owner:INST OF MEDICAL INFORMATION TECH CO LTD

Parameter-efficient large-language fine-tuning federated learning framework

Provided in the present invention is a parameter-efficient large-language fine-tuning federated learning framework, comprising the following steps: performing modeling on LoRA adapters of different edge clouds; since different weights exhibit different average performances on the LoRA adapters, using singular values to quantify the importance of the weights, and therefore, before each round of independent training of the LoRA adapters using N edge clouds, using a matrix singular value to decompose a BA matrix in the LoRA adapter for each trainable weight; configuring heterogeneous LoRA adapters on the basis of the importance of the weights; and using different numbers of quantization bits to quantize a pre-trained model, and performing high-precision inverse quantization on the pre-trained model only when matrix multiplication is executed, wherein the pre-trained model is quantized to the maximum number of quantization bits on the basis of the memory budget of the edge clouds. The present invention has the following beneficial effects: the present invention determines the optimal fine-tuning model structure, thereby improving the performance of LLM fine-tuning, and adapts to heterogeneous and resource-constrained edge clouds.
Owner:FUDAN UNIVERSITY

Method and device for determining model training configuration parameters and storage medium

The invention discloses a model training configuration parameter determination method and device and a storage medium, and relates to the technical field of computers, and the method comprises the steps: dividing configuration items into fixed configuration items and undetermined configuration items, and under the constraint that the fixed configuration items adopt specified configuration, determining the undetermined configuration items; predicting the predicted training performance of the to-be-trained model when the to-be-trained configuration item is trained by adopting different configuration parameters, and selecting a target configuration parameter from multiple configuration parameters of a group of to-be-trained configuration items based on the predicted training performance corresponding to the multiple configuration parameters, when the to-be-trained model is subjected to model training, the target configuration parameters are selected by one group of to-be-determined configuration items, so that the cost required by model training can be remarkably reduced, the resource utilization efficiency is maximized, the resource waste is reduced, the training efficiency is improved, and the problems of high hardware resource demand and high cost in a large model training stage in the related technology are solved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Method and apparatus for implementing ai-ML in a wireless network

A method, an apparatus, and a computer readable medium for storing instructions are described for a user terminal and a base station for updating an AI / ML configuration in case of a handover. The method performed by a user equipment comprising operating a first AI / ML configuration in a coverage area of a first base station; receiving an AI / ML configuration information indicating a second AI / ML configuration; operating the second AI / ML configuration indicated by the AI / ML configuration information in the coverage area of a second base station. Operating a first AI / ML configuration comprises operating a first AI / ML Model or a first AI / ML Model with a first AI / ML Model configuration in the coverage area of the first base station. The first AI / ML Model is associated with a first AI / ML Model identifier and the first AI / ML Model configuration is associated with a first AI / ML Model configuration identifier. Operating the second AI / ML configuration comprises operating a second AI / ML Model or the first AI / ML Model with a second configuration in the coverage area of the second base station. The second AI / ML Model is associated with a second AI / ML Model identifier and the second AI / ML Model configuration is associated with a second AI / ML Model configuration identifier. Indicating the second AI / ML configuration comprises indicating a second AI / ML Model identifier and / or a second AI / ML Model configuration identifier.
Owner:HARFANG IP INVESTMENT CORP

Systems and methods for generating and executing function calls using machine learning

The methods, systems, and computer networking apparatuses described herein enable language models to receive input (e.g., a query or a request) from a user or application, and without any additional training data or instructions, determine to generate a function call based on the received input from the user and generate the function call based on the determination. In some embodiments, a language model may further access an external tool or application to request an output as a response to the generated function call. The disclosed methods, systems, and networking apparatuses improve the technical field by incorporating language model capabilities within the function calling process, and allowing for function-related information to be provided to a language model via input received in any number of formats or types, including structured or unstructured input, language or non-language input, or any combination thereof.
Owner:OPENAI OPCO LLC

Temporal dynamics simulation in matmul-free neural architectures

A method is provided for processing data in a neural network system. The method includes receiving input data; processing the input data through a first set of neural network layers configured to perform data processing using MatMul-free techniques to produce intermediate data; further processing the intermediate data through a second set of neural network layers configured to simulate spiking neural network (SNN) functionalities using MatMul-free techniques; and outputting a result based on the processed data from the second set of neural network layers.
Owner:LEPTUDE INC

Artificial intelligence-based topographic surveying and mapping method and system

The invention discloses a topographic surveying and mapping method based on artificial intelligence. The topographic surveying and mapping method comprises the steps of (1) multi-source data fusion acquisition; (2) intelligently extracting topographic features; (3) three-dimensional reconstruction optimization; (4) dynamic terrain evolution modeling; the invention also discloses a system for implementing the topographic mapping method based on artificial intelligence, and the system comprises a data acquisition and preprocessing module, a topographic feature intelligent extraction module, a three-dimensional reconstruction optimization module, a dynamic topographic evolution modeling module and an interactive visualization and decision platform. The data acquisition and preprocessing module, the terrain feature intelligent extraction module, the three-dimensional reconstruction optimization module, the dynamic terrain evolution modeling module and the interactive visualization and decision-making platform are connected in sequence, and each module communicates through a gRPC interface and supports independent upgrading.
Owner:HUATING COAL GRP CO LTD

Training a digital twin in artificial intelligence-defined networking

A system including one or more processors and one or more non-transitory computer-readable media storing computing instructions that, when executed on the one or more processors, perform certain acts. The acts can include generating a digital twin network simulation of a physical computer network controlled through a software-defined-network (SDN) control system. The acts also can include training a routing agent model on the digital twin network simulation using a reinforcement-learning model on traffic that flows through nodes of the digital twin network simulation. The routing agent model includes a machine-learning model. The acts additionally can include deploying the routing agent model, as trained, from the digital twin network simulation to the SDN control system of the physical computer network. Other embodiments are described.
Owner:WORLD WIDE TECHNOLOGY HOLDING CO LLC

Graph neural network execution on neural processing unit

Workloads for executing a graph neural network (GNN) may be divided among various processing units, such as a central processing unit (CPU) and a neural processing unit (NPU). The NPU may include a data processing unit (DPU) and a digital signal processor (DSP). The CPU may perform precomputation, model optimization, hardware optimization, and compilation. For example, the CPU may precompute a parameter matrix and use the parameter matrix as internal parameters of a GNN. The CPU may also perform node padding, approximation computation, or transfer of DSP operations to DPU to optimize the GNN. The CPU may also perform sparsity data compute and storage, vertical fusion of DSP operations and DPU operations, or data quantization to optimize performance of the NPU. The compiled GNN may be provided to the NPU, and the DPU and DSP may perform the operations in the compiled GNN to produce a prediction of the GNN.
Owner:INTEL CORP

Large model active tool calling method and system based on multi-step reasoning

The invention provides a large model active tool calling method and system based on multi-step reasoning, and relates to the technical field of artificial intelligence, and the method comprises the steps: collecting tool information, generating a structured tool description set, automatically generating an executable code sample and corresponding natural language explanation based on the tool description set, and constructing a reverse training data set; the method comprises the following steps of: marking tool use necessity tags in a real problem, prompting a large model to generate a multi-step reasoning path combined with a natural language and codes, verifying and filtering the correctness of the reasoning path, and constructing a forward training data set; and performing three-stage training on the large model based on the reverse training data set and the forward training data set. According to the method, the problem that the active tool calling capability of a large model is insufficient is solved, and the code execution rate and problem solving efficiency of a complex reasoning task are improved.
Owner:李俊涛

Multi-agent arrangement method and device based on service grid and storage medium

The embodiment of the invention provides a multi-agent arrangement method and device based on a service grid and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: generating corresponding domain-specific language description information based on workflow information of a plurality of agents; and based on the domain-specific language description information, generating routing information corresponding to the workflow information of the plurality of agents, the routing information being used for defining a calling relationship of the plurality of agents in the service grid. According to the scheme, the labor cost is reduced, and the arrangement efficiency is improved.
Owner:ALIBABA CLOUD COMPUTING CO LTD

User Behavior Modeling for Detecting and Containing Malicious Activity in a Storage System

An illustrative method includes monitoring operations performed with respect to a storage system by an entity using an identity associated with a particular role, the particular role providing the entity with a set of permissions associated with the storage system; determining, based on the monitoring, that one or more operations of the operations deviate from an expected activity profile associated with the role by more than a threshold; and performing, based on the determining that the one or more operations deviate from the expected activity by more than the threshold, a remedial action with respect to the entity.
Owner:PURE STORAGE INC

Parallel strategy optimal selection method, and neural network solver training method and apparatus

The present application discloses a parallel strategy optimal selection method, an electronic device, a device, and a neural network solver training method and apparatus. The parallel strategy optimal selection method comprises: on the basis of a neural network solver and a heuristic candidate strategy, quickly acquiring a plurality of candidate hybrid parallel strategies; predicting overhead values of the candidate hybrid parallel strategies; and using the predicted overhead values as guidance to perform optimal selection among the candidate hybrid parallel strategies so as to quickly determine a hybrid parallel strategy. The present application improves the parallel training efficiency of large models, and reduces energy consumption.
Owner:HUAWEI TECH CO LTD

Large language model multi-agent cooperative work method and system

The invention relates to the technical field of large language models, in particular to a large language model multi-agent cooperative working method, which comprises the following steps: S1, designing a plurality of agents to simulate social division and cooperation, so that each agent executes different professional abilities, the plurality of agents including a planning agent, a tool agent and a reflection agent; s2, when task input is obtained, the planning agent filters retrieved tools, and a complex task is decomposed into subtasks easy to understand in a code and annotation mode; s3, performing targeted tool enhancement on the filtered tool set by the tool agent, and executing specific operation planned by the planning agent to obtain a subtask result; s4, evaluating a subtask result by the reflection agent, and planning the task from different angles by combining static planning and dynamic planning; and S5, the planning agent feeds back a summary result according to execution of the subtasks.
Owner:NO 63921 UNIT OF PLA +1

Cost-aware efficient tool planning method based on large model

The invention provides a cost-aware efficient tool planning method based on a large model. According to the method, efficiency and cost optimization of tool scheduling is realized by constructing a CATP-LLM framework. The method comprises the steps that a tool planning language TPL is designed, a non-linear multi-branch parallel scheme is generated through structured Token support, and the task execution efficiency is remarkably improved; in combination with a cost-aware offline reinforcement learning CAORL algorithm, a tool selection strategy is dynamically optimized based on historical data, and task performance and resource consumption are balanced; and the balance relationship between the performance and the cost is quantified through the QoP of the scheme quality index, and dynamic adjustment is realized to output an efficient and low-cost planning scheme. Compared with the prior art, parallel tool calling, dynamic cost optimization and complex task adaptability are supported, the execution cost is remarkably reduced while the task quality is ensured, and an efficient solution is provided for multi-scene task planning.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL +1

Method for improving calculation speed of model based on mercuric chloride AI processor

The invention relates to a method for improving the calculation speed of a model based on a mercuric chloride AI processor. The method comprises the following steps: deploying the mercuric chloride AI processor and a deep learning model in a server, and carrying out adaptation and optimization on the mercuric chloride AI processor; before the data in the first buffer area is read, predicting and preloading the data to be processed, and loading the data from the global memory to the second buffer area in advance; a double-buffer mechanism is arranged in the mercuration AI processor, interrupt and event trigger points are set, the operation of loading data to a next buffer area is immediately started when a specified calculation stage is finished, and the time sequence of data circulation is accurately controlled; and a control parameter is automatically adjusted based on monitoring data of the real-time monitoring module, the calculation process is decomposed into a plurality of stages, different buffer areas are allocated for each stage, and access to the shared memory is optimized by setting a specified cache replacement strategy and the size of a cache line. According to the process, the calculation speed of the model in the deep learning field is improved, the resource utilization rate is improved, and the defect of manual adjustment and optimization is overcome.
Owner:四川华鲲振宇智能科技有限责任公司

Large model dynamic compression optimization method and system based on sparse pruning

The invention relates to the technical field of large model algorithms, in particular to a large model dynamic compression optimization method and system based on sparse pruning, and the method comprises the steps: capturing original weight fluctuation data generated by resource fluctuation in reasoning, and obtaining sparse weight reference data through sparse processing; analyzing calculation complexity through model reasoning delay data, and separating reasoning delay amount caused by a model scale; dynamically controlling the model compression ratio within a preset performance range based on the delay amount and the sparse reference data, and collecting reasoning precision distribution data under different compression parameters; evaluating a model performance state under each parameter by means of a neural network simulation method, and generating a performance state simulation result; determining a model quality optimization compensation parameter based on a simulation result by combining resource fluctuation data acquired in real time in a compression process; and finally, the compression strategy is adaptively regulated and controlled through the compensation parameters, and collaborative optimization of model calculation complexity, reasoning precision and delay during dynamic change of hardware resources is realized.
Owner:NOVNET COMPUTING SYST TECH CO LTD

Automatic kernel network parameter optimization method

The invention relates to the technical field of parameter optimization, in particular to an automatic kernel network parameter optimization method, which comprises the steps of constructing an enhanced deep Q network model, and integrating the enhanced deep Q network model with a priority playback buffer area, a meta learning module, a Bayesian optimizer and a neural architecture search module; using performance index data to train an enhanced deep Q network model, the training process including using a priority playback buffer to store and sample empirical data, using a meta-learning module to perform task adaptation, and monitoring training indexes of multiple dimensions to evaluate the convergence state of the model; selecting a kernel parameter adjustment action according to the current state through the trained enhanced deep Q network model; executing the selected kernel parameter adjustment action, and evaluating a parameter adjustment effect based on the multi-target reward function; and updating the enhanced deep Q network model according to an evaluation result, wherein the priority playback buffer area and the Bayesian optimizer are utilized in the updating process.
Owner:GUANGZHOU CITY UNIV OF TECH

Workload balance with prompt and token routing for expert models

Systems and methods are provided for determining expert placement layouts and / or prompt and toke routing for a mixture of experts (MoE) model. In some instances, an expert workload distribution with respect to a plurality of experts is determined based on a gating neural network. In some instances, an expert placement layout with respect to a plurality of computing units is determined based on the determined expert workload distribution. In some instances, a two-level routing strategy is provided to first adaptively route an incoming prompt to a suitable computing device, then perform a token routing within that computing device to ensure workload balance between different computing units of that computing device.
Owner:BYTEDANCE TECHNOLOGY LTD

Neural network image-text analysis and cross-framework code generation method and system

The invention discloses a neural network image-text analysis and cross-framework code generation method and system, and the method comprises the steps: analyzing a neural network architecture picture through a visual large model, and extracting a layer type, a connection relation and a topological structure feature; performing semantic understanding on text description by combining a large language model, and extracting layer parameters and configuration information in a standardized manner; utilizing a multi-modal alignment mechanism to fuse vision and text features, and generating unified model representation; and finally, directly generating an executable code supporting a mainstream framework based on a large language model and grammar check. According to the method, the limitation of traditional manual coding is broken through, end-to-end generation from a complex framework to a multi-framework code is achieved, the problems of cross-framework adaptation and semantic understanding are solved, the development efficiency of deep learning is improved, and the method is suitable for scientific research verification and industrial deployment scenes.
Owner:HARBIN INST OF TECH

Artificially intelligent routing agent for routing portions of a task through multiple customized agents, and systems, devices, and methods of use thereof

This application describes, amongst other things, methods and systems for building and deploying agents. An example method includes obtaining orchestration data about a set of task-specific components selected to provide a response to the prompt, where each respective task-specific components in the set of task-specific components is configured to assist with a respective clinical task of the one or more clinical tasks. The method further includes, determining an order in which each respective task-specific components of the set of task-specific components should be utilized to prepare a complete response to the prompt that address the one or more clinical tasks based on the obtained orchestration data about the set of task-specific components. The method also includes, in accordance with the determined order, providing first data related to the prompt to a first task-specific component and receiving a first response from the first task-specific component.
Owner:TEMPUS AI INC

Systems and methods of large language model driven orchestration of task-specific machine learning software agents

Systems and methods of the present disclosure may receive, from a user computing device, a user-provided data record query including a natural language request for information associated with one or more data sources. User persona attributes of the user may be determined, such as a user role or security parameters or both. Based on the user persona attributes a context query may be generated to obtain context attributes associated with the user-provided query. The natural language request and the context attributes are input into the model orchestration large language model (LLM) to output instructions to machine learning (ML) agents based on the context attributes. The ML agents output responses associated with the user-provided data record query based on the instructions, and the responses are input into the model orchestration LLM to output to the user computing device a natural language response based on the context attributes.
Owner:BROADRIDGE FINANCIAL SOLUTIONS

Intelligent routing and unified adaptation method for large language model

The invention discloses an intelligent routing and unified adaptation method for a large language model, and the method comprises the steps: defining all access details of the model through a declarative configuration file, and achieving the zero-code access of the large language model without writing any code for a newly-added model; a completely consistent calling interface is provided for an upstream application, the isomerism of all downstream large language models is shielded, and a unified request and response abstraction layer is constructed; through a strategy engine and a JSON path technology, complex streaming response including content thinking is precisely processed, and intelligent analysis and content extraction are carried out; dynamic configuration and intelligent strategy hot update are supported; intelligent routing of the model is realized, and an optimal large language model instance is dynamically selected; and meanwhile, enterprise-level governance capability is provided, governance functions such as fusing, degradation, current limiting and monitoring are integrated, and stability guarantee is provided for model calling. Therefore, the maintainability, the expandability and the user experience consistency of the system are comprehensively improved.
Owner:NANJING INFORMATION HIGH-SPEED RAILWAY RES INST OF SCI AND TECH

Optimization method for distributed execution of deep learning task, and distributed system

Disclosed are an optimization method for distributed execution of a deep learning task and a distributed system. The method includes that: a computation graph is generated based on a deep learning task and hardware resources are allocated for the distributed execution of the deep learning task; the allocated hardware resources are grouped to obtain at least one grouping scheme; for each grouping scheme, tensor information related to multiple operators contained in the computation graph is split based on the value of at least one factor under this grouping scheme to obtain multiple candidate splitting solutions; and an optimal efficiency solution for executing the deep learning task of the hardware resources is selected by using a cost model. Through operator splitting based on device grouping combined with optimization solving based on the cost model, automatic optimization of distributed execution for various deep learning tasks is realized. Furthermore, computation graph partitioning based on grouping can be introduced, and the solving space can be restricted according to different levels of optimization, thereby generating a distributed execution solution of required optimization level within controllable time.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Heterogeneous computing power adaptive compiling method and system for large model

The invention provides a large-model-oriented heterogeneous computing power adaptive compiling method, which comprises the following steps that: a user inputs a trained large model through a system interface, and a system front-end conversion module analyzes a computational graph of the model and converts the computational graph into an intermediate representation based on a unified operator description language (UDL); the system hardware sensing module automatically detects and extracts hardware feature fingerprints of at least one piece of target hardware; based on the unified operator description language UDL intermediate representation and the hardware feature fingerprint, automatically generating an optimization adaptation rule oriented to at least one piece of target hardware; wherein the basis of parameterized filling comprises specific parameters of hardware feature fingerprints and optimized attribute tags carried in an intermediate representation of a unified operator description language (UDL); and generating and deploying multiple back-end codes. The method has the beneficial effects that intelligent compiling based on hardware features can be realized, and the deployment efficiency and the operation performance of a large model in a complex heterogeneous computing power cluster are remarkably improved.
Owner:SHENZHEN XINGSHENG DIGITAL TECH CO LTD

Orchestrate events in Distributed DevOps Apparatus Leveraging Generative AI

Systems and methods for orchestrating events in distributed DevOps apparatus leveraging generative AI are disclosed to automate and streamline software development and deployment in distributed DevOps environments using Generative Adversarial Networks (GANs) and other AI techniques. The method involves interpreting UML diagrams, design documents, or the like with generative AI and computer vision to create DevOps tasks, integrating with various DevOps tools for task management, and deploying generated rules for automated event execution. Metadata is generated and processed in order to facilitate AI analysis. The systems and methods reduce manual intervention, increase efficiency, and improve accuracy, scalability, security, and compliance in DevOps workflows.
Owner:BANK OF AMERICA CORP

Methods and systems for the automated generation of neural network architectures

Presented herein are systems and methods for generating a neural network architecture (NNA) tailored for a given task. In certain embodiments, the technology automatically identifies a neural network architecture appropriate to perform the task, tweaks the architecture to meet one or more particular use-case requirements, and trains the most optimal neural network model for the task.
Owner:UNIVERSITY OF MAINE

Large model energy consumption optimization method and device, computer equipment, readable storage medium and program product

The invention relates to a large model energy consumption optimization method and device, computer equipment, a computer readable storage medium and a computer program product. Comprising the steps of collecting hardware state data and task data of a large model; performing stage identification according to the task data, and determining a current stage; performing state prediction according to the hardware state data and the task data through a prediction model corresponding to the current stage to obtain target state information; generating optimization parameters of the current stage through a joint optimizer according to the target state information; and adjusting the resources of the large model according to the optimization parameters. In combination with stage perception and dynamic resource adjustment, a differentiated resource adjustment strategy based on different stages is realized, specific energy consumption pain points in an AI scene are solved, high power consumption of large model training, instantaneous fluctuation of reasoning requests and the like are reduced, sustainable and efficient operation of an AI system is ensured, the computing power demand and resource consumption of a large model are balanced, and the system performance is improved. And wide landing and development of a large model in various scenes are facilitated.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Methods and apparatus for hardware-aware machine learning model training

Methods, apparatus, systems, and articles of manufacture are disclosed for hardware-aware machine learning model training. An example apparatus includes a configuration determiner to determine a hardware configuration of a target hardware platform on which the machine learning model is to be executed, a layer generator to assign sparsity configurations to layers of the machine learning model based on the hardware configuration, and a deployment controller to deploy the machine learning model to the target hardware platform in response to outputs of the machine learning model satisfying respective thresholds, the outputs including a quantity of clock cycles to execute the machine learning model with the layers having the assigned sparsity configurations.
Owner:INTEL CORP