Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

242 results about "Network processor" patented technology

A network processor is an integrated circuit which has a feature set specifically targeted at the networking application domain. Network processors are typically software programmable devices and would have generic characteristics similar to general purpose central processing units that are commonly used in many different types of equipment and products.

Complex scene-oriented AI large model lightweight deployment method

The invention provides a complex scene-oriented AI large model lightweight deployment method, and relates to the technical field of edge computing, and the method comprises the steps: carrying out the structured pruning of a pre-trained Transform network based on the attention head importance score, carrying out the dynamic sparsification of the activation state of a feedforward network according to the input tensor entropy value, employing the dynamic mixing precision quantization, and carrying out the reconstruction of an AI large model. Obtaining network parameters after pruning quantization; deploying the pruned and quantized network parameters to an edge computing device, distributing a feature extraction operator to a neural network processor through a heterogeneous computing scheduler, and unloading a classification operator to a multi-core central processing unit; and managing an on-chip memory in combination with a virtual memory paging mechanism, realizing zero-copy data transmission by utilizing a direct memory access controller, and outputting a reasoning result tensor. According to the method, efficient and reliable operation of the large model at the resource-constrained edge node is realized.
Owner:XIAN XINGXUN INTELLIGENT COMM TECH CO LTD

Coagulant adding control system and method based on multi-scale alumen ustum characteristic analysis

The invention provides a coagulant addition control system and method based on multi-scale alumen ustum characteristic analysis, and relates to the technical field of coagulant addition control. Comprising an underwater camera device, an embedded prediction module, an integrated neural network processor NPU, a dosing device, a feature extraction module and a time sequence feature coding module. According to the method, underwater camera devices are arranged in a flocculation area and a settlement area respectively, spatial multi-dimensional feature analysis is carried out on images acquired at double view angles, spatial multi-scale parameters are obtained, and real-time prediction of coagulant dosage is realized through an improved deep learning model; the improved deep learning model is embedded into a controller of an integrated neural network processor NPU, after the controller receives a model prediction result, regulation and control signals of PAC and PAM are generated in combination with the current operation state of the dosing pump, and dosing equipment is driven to execute actions; meanwhile, a regulation and control result is fed back to the feature extraction module in real time, and accurate matching of the coagulant dosage and dynamic evolution of alumen ustum is ensured.
Owner:NORTHEASTERN UNIV CHINA

Intelligent control method, device, equipment and system for automatic stamping

The invention relates to the technical field of office automation, in particular to an intelligent control method, device, equipment and system for automatic stamping, and aims to recognize and semantically analyze the recognized file content of a paper material through a multi-modal stamping feature recognition model, realize intelligent determination and calibration of a stamping position, improve the stamping accuracy and improve the stamping efficiency. The stamping method can adapt to different file formats and is flexible without presetting a fixed template, so that the applicability of the stamping method is improved; edge AI computing power is deployed in the stamping equipment, and a neural network processor is locally integrated, so that data processing localization is realized, data security is improved, data response delay in the stamping processing process is reduced, and network dependence is reduced; and based on local powerful computing power support of the stamping equipment, large-batch stamping processing can be achieved, the stamping efficiency is improved, and therefore the working efficiency is improved, and the human resource cost is reduced.
Owner:GUANGDONG PLANNING & DESIGNING INST OF TELECOMM

Code compiling method and device based on P4 assembler, assembler and medium

The invention provides a code compiling method and device based on a P4 assembler, the assembler and a medium. The code compiling method comprises the steps that a to-be-processed assembly code file is obtained; splitting and compiling the assembly code file to obtain a plurality of code segment compiling results; according to the auxiliary information, by taking Parser components, Mat components and Deparser components contained in the Ingress unit and the Egress unit as classification types, classifying the code segment compiling results to obtain code segment groups; linking the P4 table item data and the machine instruction code contained in the code segment group, and endowing a link result with a global instruction identifier and code segment header information to obtain a code segment binary code; and generating a firmware file based on the code segment binary code. The compiling process of the network processor can be simplified, so that the development efficiency is improved.
Owner:YIHUA TECHNOLOGY (BEIJING) CO LTD

Quantization and inverse quantization method in large language model and neural network processor

The invention relates to a quantization method, an inverse quantization method and a neural network processor in a large language model. A core (namely a quantization unit) in a neural network processor is arranged to execute quantization processing in a large-scale language model reasoning process, and the quantization unit has a data partitioning function, so that the quantization processing efficiency is greatly improved compared with an existing quantization processing core which is limited by the data size (such as block wise) when receiving data to be quantized. And frequent interaction with a cache is not needed, so that online high-efficiency large-model quantification processing is realized. And performing an inverse quantization process in a large language model inference process by setting a core (i.e., an inverse quantization unit) in a neural network processor, and the inverse quantization unit having a data partitioning function, a matrix multiplication function, and a high precision accumulation function (e.g., multiplying by a corresponding quantization parameter), the inverse quantization processing does not need to frequently carry data in a plurality of cores, and the inverse quantization processing efficiency of a large model is improved.
Owner:北京凌川科技有限公司

Flexible and scalable thermal test vehicle design for electronics cooling solutions

PCT designated stageWO2026084737A1Analog circuit testingDigital circuit testingTransistor arrayNetwork processing unit
The density and power consumption of modern integrated circuits, such as Graphic Processing Units (GPUs), Central Processing Units (CPUs), and Network Processing Units (NPUs) is growing rapidly, which necessitates designing advanced cooling systems. Existing solutions for characterizing and validating these cooling system are inadequate. A flexible, scalable Thermal Test Vehicle (TTV) is disclosed which is based on an array of power transistors, measurement / control circuitry, and onboard computer. The TTV is configured for characterizing the performance of electronic cooling solutions under a variety of operating conditions.
Owner:RGT UNIV OF CALIFORNIA +1

Software and hardware collaborative optimization method of hybrid in-memory architecture

The invention discloses a software and hardware collaborative optimization method for a hybrid in-memory architecture, which comprises the following steps of: performing joint feature representation on a target AI algorithm and an in-memory computing architecture, extracting context features, and parameterizing in-memory computing unit configuration, a neural network processor assembly line, a multi-core interconnection topology and a storage level interface of the in-memory computing architecture; an off-line reference data set is constructed, a predictive agent model is trained, discrete architecture parameters are processed by the model by adopting an embedding method, feature association is learned by applying an encoder with a self-attention mechanism, joint prediction of multi-dimensional PPA indexes is realized through a parallel prediction network, and a feasible region constraint learning mechanism is introduced in training; the trained agent model is embedded into a multi-objective evolutionary algorithm, the energy efficiency ratio, the average computing power utilization rate, the model execution delay and the like serve as optimization objectives, chip area efficiency and power consumption constraints are met at the same time, and a Pareto optimal in-memory computing architecture configuration set is searched.
Owner:SOUTH CHINA UNIV OF TECH

High-speed multi-port cache arbiter with variable word length

The invention discloses a high-speed multi-port cache arbiter with a variable word length. Along with the high-speed development of the modern network technology, storage management and scheduling take up most of time in the processing of data packets by network equipment, so that the low-speed caching capability of most memories becomes a bottleneck for limiting the further improvement of a network processor. Therefore, by managing the shared cache of not less than 32 blocks of 256Kbit SRAM units and supporting simultaneous cache writing of a plurality of ports through an arbitration mechanism, write scheduling and read scheduling can be realized without mutual influence, and concurrent read-write conflicts are avoided; meanwhile, each port supports eight priority queues, so that the data transmission bandwidth of each port can reach 1G bps; by processing a data packet, caching and data scheduling according to packets are supported, the length of the data packet is 64-1024 bytes variable, and the influence of the length of the data packet on memory resources can be avoided; and the purpose of storing resources can be achieved by dynamically adjusting space saving through memory recovery, and the data storage efficiency is improved.
Owner:NANJING UNIV

Configurable decompression circuit supporting COO and Bitmap compression algorithms

The invention relates to a configurable decompression circuit supporting COO and Bi tmap compression algorithms, and belongs to the field of integrated circuits. According to the hybrid compression method suitable for the circuit, after sparseness analysis is carried out on weight matrixes of all layers of a neural network, COO compression based on a coordinate type sparse matrix or Bitmap compression based on bitmap masks is selected according to the sparseness characteristics of different layers, so that the storage space is optimized; comprising a control module, a first selector, a second selector, a bitmap description memory, a coordinate index memory, a numerical memory, a converter, a COO decoder and a decompression data storage module, decoding of two compression formats of COO and Bi tmap can be supported, and the operation efficiency and flexibility of a neural network processor are improved. The method effectively reduces the weight storage demand of the neural network, enhances the adaptability of the compression algorithm, and is suitable for hardware implementation of an efficient neural network model.
Owner:BEIJING MXTRONICS CORP +1

Deep learning model reasoning acceleration method and system based on heterogeneous computing architecture

The invention relates to the technical field of deep learning reasoning acceleration, and discloses a deep learning model reasoning acceleration method and system based on a heterogeneous computing architecture. The method comprises the steps of obtaining a calculation intensity index by analyzing a calculation operation type of a model layer, and dividing the model into a front calculation area, a dynamic calculation area and a rear calculation area according to the calculation intensity index. And allocating the calculation-intensive and regular front and rear areas to a neural network processor, and allocating the dynamic sparse area to a central processing unit. The neural network processor executes preposition calculation, generates a preposition activation value vector and transmits the preposition activation value vector to the central processing unit; after completing dynamic calculation, the central processing unit transmits a generated dynamic activation value vector back to the neural network processor; and the neural network processor executes post calculation and outputs a result. According to the method, the utilization efficiency of heterogeneous hardware resources is improved and the overall reasoning time delay is reduced through calculation feature guided fine task division and cooperative scheduling between processors.
Owner:STORAGEX TECH INC

Dynamic slimmable neural network for sequential data processing

There is disclosed an apparatus (100, 200, 300, 400, 500) comprising: an input interface (101) to receive an input segment (102) of an input sequential data; a neural network, NN, processor (110) to derive an output result (104) by processing the input segment (102) through a NN having a number of layers (111) from a first layer to a last layer, the NN using a predetermined number of deactivatable units which are selectively deactivatable; a gating module (120) configured to deactivate at least one deactivatable unit based on the input segment (102) and / or on at least one intermediate output segment (112, 123), an output interface (103) configured to provide an output (104) derived from the last layer of the NN.
Owner:FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV

Apparatus, method and computer program for processing an audio signal using feature segmentation and feature combination

An apparatus for processing an information signal has: a feature extractor for extracting a set of features having a first dimension; a feature segmenter for segmenting into a first subset having a second dimension and a second subset having a third dimension, which overlap, both being lower than the first dimension; a neural network processor for processing the first and second subsets using a first and a second neural network to obtain a first and a second result, respectively; a feature combiner for combining the first and second results using a third neural network, having a third complexity lower than a first or a second complexity of the first and second neural network to obtain a result set of features having a result dimension; and an output post-processor for post-processing the result set of features to obtain a processed information signal.
Owner:FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV

Photonic NPU embedded neural network processor

The invention discloses a photon NPU (Network Processing Unit) embedded neural network processor, which comprises (1) an electromagnetic wave bus, (2) a multiply-add module, (3) an activation function module, (4) a two-dimensional data operation module and (5) a decompression module, the NPU embedded neural network processor is divided into independent modules with different sizes according to functions and purposes, each functional module is provided with an independent input and output end, and each functional module is provided with an independent input end and an independent output end. All input and output ends are connected with a transmitting end and a receiving end of an electromagnetic wave bus, the bandwidth advantage of electromagnetic waves is utilized, a processing terminal is further formed through interconnection, and the electromagnetic wave bus is used for replacing a control bus, a data bus, an address bus and all replaceable circuits to transmit needed control signals, data and data addresses. An electromagnetic wave bus is used as a carrier of data needing to be processed, a data address and control signal transmission, and a photon (NPU) embedded neural network processor processes and outputs a result after passing through a radio frequency chip and a baseband chip.
Owner:刘国栋

Image processing method and chip

Embodiments of the present application provide an image processing method. The method comprises: an image signal processor in a chip processing data received from a camera to generate a target image; the image signal processor writing a first patch in the target image into a system cache in the chip; a neural network processor of the chip reading the first patch from the system cache; the neural network processor processing the first patch based on a neural network model to obtain a first output patch; and the neural network processor writing the first output patch into the system cache. The processing includes at least one of color interpolation or high dynamic range (HDR) processing. The technical solution provided by the present application saves the time-consuming of image processing and saves the system power consumption.
Owner:HUAWEI TECH CO LTD

Server information collection, updating method and device, system, medium, product and equipment

This application discloses a method, apparatus, system, medium, product, and device for server information collection and updating. The method includes: acquiring a target image; loading the acquired target image into a target server, wherein the target server includes a neural network processor (NPU); and collecting hardware information of the target server through the loaded target image, wherein the hardware information includes the network interface card (NIC) information of the NPU. This improves the efficiency and accuracy of server hardware information collection, thereby increasing the construction efficiency and scalability of large-scale clusters and reducing the operational complexity of the cluster.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

24-bit intermediate accumulation processing method and device for NPU convolution operation and server

The invention provides a 24-bit intermediate accumulation processing method and device for NPU convolution operation and a server, and relates to the technical field of neural networks, and the method comprises the steps: determining data path information of original data according to a multiplication and addition calculation result between a preset weight and an activation value corresponding to the input original data when a neural network processor NPU carries out convolution operation, the preset accumulation frequency corresponds to the data path information; the data path information is compared with a preset data path threshold value, a target processing strategy of the original data is determined according to a comparison result, and the target processing strategy comprises a low input bit width processing strategy and a high input bit width processing strategy; and performing segmented accumulation processing on the original data according to a preset accumulation frequency, and executing 24-bit intermediate value read-back, addition and write-back processing corresponding to the target processing strategy to obtain a target convolution operation output result. According to the method, the operation cost can be reduced while the error is reduced.
Owner:ZHEJIANG XINMAI SILICON CO LTD

Neural network processor

The invention discloses a neural network processor, which is applied to the field of artificial intelligence and comprises a systolic array, a multiplication accumulator array, a processing scheduler, a data dispatcher and a memory. The invention provides a hybrid SA / MAC array collaborative architecture, which solves the problem of high vacancy rate of a single systolic array or a multiplication accumulator array by dynamically allocating tasks through a processing scheduler, can adapt to various network layer structures of large, medium and small sizes, and improves the processing efficiency; compared with a pure multiplication accumulator array, the hardware area and the implementation complexity are reduced, and the hardware cost is reduced; compared with a pure systolic array, the data loading and carrying-out time is shortened, a small amount of tasks which cannot be efficiently completed by the systolic array are processed through the multiplication accumulator array, the systolic array is prevented from being repeatedly started, and the total calculation time consumption is reduced; the data dispatcher manages the data flow direction in a unified mode, it is ensured that data interaction between the systolic array and the multiplication accumulator array is accurate and efficient, and respective operation is not affected.
Owner:CCORE TECH CO LTD

A lightweight neural network processor storage architecture co-optimization method

The present application relates to neural network inference hardware technical field, particularly to a kind of lightweight neural network processor storage architecture collaborative optimization method, comprising: step 1: the difference of on-chip data on bandwidth demand and access mode is analyzed, and the differential single-port storage setting is carried out to the buffer storage of each data in neural network processor;Step 2: the GEMM module of neural network processor is based on semi-pulsating array setting, and ALU operation fusion processing module is used, to reduce the storage bandwidth and bit width of data;Step 3: the storage structure is optimized by introducing the dynamic adjustment strategy based on timing analysis, to output the optimized neural network processor.The present application guarantees the computing performance, significantly improves the area efficiency and access efficiency of storage system, realizes the collaborative optimization of bandwidth, area and power consumption.
Owner:ZHEJIANG UNIV

Intelligent control method, device, equipment and system for automatic stamping

The present application relates to the technical field of office automation, and more particularly to an intelligent control method, device, equipment and system for automatic stamping, which identifies and performs semantic analysis on the file content of the recognized paper materials through a multi-modal stamping feature recognition model, realizes intelligent determination and calibration of the stamping position, improves the stamping accuracy, and is flexible in adapting to different file formats and not requiring a preset fixed template, thereby improving the applicability of the stamping method; and by deploying edge AI computing power in the stamping device, integrating a local neural network processor, realizing data processing localization, improving data security, reducing data response delay in the stamping processing process, and reducing network dependency; and based on the strong local computing power support of the stamping device, a large number of stamping processes can be realized, the stamping efficiency is improved, and thus the work efficiency is improved and the human resource cost is reduced.
Owner:GUANGDONG PLANNING & DESIGNING INST OF TELECOMM

A task compilation method, device and compiler

Embodiments of the present application provide a task compiling method, device and compiler. The method comprises: the compiler receiving at least one compiling task input by a user, the compiler judging whether the compiling task comprises multiple branch tasks, if it is judged that the compiling task comprises multiple branch tasks, then dividing the hardware resources of an embedded neural network processor (NPU) obtained according to the multiple branch tasks, generating a hardware allocation result, and generating a first compiling instruction according to the hardware allocation result, the compiling parameters input by the user and the core parameters required by each NPU core calculation obtained; the compiler sending the first compiling instruction to a scheduler, so that the scheduler schedules the hardware resources according to the first compiling instruction, thereby realizing reasonable allocation of the hardware resources of the NPU through the compiler and improving the operation speed of the NPU.
Owner:SPREADTRUM COMM (TIANJIN) INC

QUERY CHANGES BY NEURAL NETWORKS

Processors, systems, and techniques for predicting database queries are described. In at least one embodiment, one or more previous query results are obtained from a database, and one or more neural networks are used to predict a database query, at least partially, based on one or more previous query results.
Owner:NVIDIA CORP

Intelligent control method for aero-engine and related equipment

The embodiment of the invention provides an intelligent control method for an aero-engine and related equipment, and belongs to the technical field of data processing. The method comprises the steps that if it is determined that the type of a calculation task is a general calculation task, a central processing unit is scheduled to execute processing of to-be-processed data; if it is determined that the calculation task type is the intelligent calculation task, a neural network processor unit is scheduled to execute processing of the to-be-processed data; if it is determined that the calculation task type is a digital signal processing task, a digital signal processor platform is scheduled to execute processing of the to-be-processed data; when the neural network processor units are scheduled to execute the intelligent computing task, if it is detected that the computing power of the neural network processor unit of the main platform is insufficient, part of the to-be-processed data or the computing sub-graph is distributed to the neural network processor unit of at least one auxiliary platform for parallel processing. According to the embodiment of the invention, the real-time and reliability requirements of a complex intelligent processing task can be met in a computing environment of an aero-engine.
Owner:CHINA AERONAUTICAL CONTROL SYST RES INST +1

Voltage frequency adjustment method and apparatus, neural network accelerator, and storage medium

The embodiment of the application discloses a voltage frequency adjustment method and device, a neural network accelerator and a storage medium. The method comprises the following steps: for each calculation layer in a preset neural network, the size of a corresponding feature image and the size of a contained convolution kernel are used to determine a corresponding total calculation amount, and the time length of a calculation engine for completing the corresponding total calculation amount is estimated to determine the corresponding calculation time length; for each calculation layer in the preset neural network, the storage amount of the corresponding feature image in each memory of a plurality of memories and the transmission bandwidth of each memory of the plurality of memories are used to estimate the time length of the transmission of the corresponding feature image to the calculation engine, and the corresponding memory access time length is determined; during the execution of each calculation layer in the preset neural network by the calculation engine in the neural network processor, the voltage frequency of the neural network processor is dynamically adjusted based on the corresponding calculation time length and the memory access time length of each calculation layer.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

An intelligent auxiliary control electrical system based on a special excavator

The application discloses an intelligent auxiliary control electrical system based on a special excavator, an auxiliary data processing module, the auxiliary data processing module is a neural network processor, receives data and processes, outputs a control instruction, a whole machine controller, collects whole vehicle data, an execution module, accepts the control instruction and executes the instruction, a data display module, through the cooperation of hardware and software, the state of the equipment is monitored and displayed in real time, the mode of the equipment is switched and the function is selected in the man-machine interaction mode, the device reduces hardware modification through a standardization module, adjusts configuration through a data display module, such as the logic and rules of a safety warning lamp, the research and development cycle is greatly shortened, the electrical design efficiency is improved, the electro-hydraulic cooperative control of the whole machine and various special functions are realized, the whole machine control system program is effectively avoided from changing with machines, the device further realizes the switching of different modes through the setting of a double electromagnetic valve group, and the application scene and the capacity of the excavator are greatly improved.
Owner:SHANZHONG JIANJI CO LTD

Dynamic forward information base prefix optimization

A network device includes a routing information base including a first plurality of entries, a forwarding information base including a second plurality of entries; a forwarding information base entry optimizer that programs the second plurality of entries of the forwarding information base using, at least in part, the first plurality of entries; and a network processor that forwards packets based on the second plurality of entries of the forwarding information base. The second plurality of entries is less than the first plurality of entries.
Owner:ARISTA NETWORKS INC

Counter-attack circuitry

PendingCN122640098ALoop controlControl switch
The present application provides a counter-attack circuit system, including a power supply, an operating switch, a plurality of current sources, a plurality of control switches, a control loop and a neural network processor. The power supply outputs an output voltage at a first node, the operating switch is coupled between the first node and a second node, and the current sources are coupled to the second node through corresponding control switches. The control loop controls the operating switch to be turned on or off according to an operating voltage, a target voltage and a control voltage, and controls the current sources and the control switches to change the operating voltage. The neural network processor is coupled to the second node, wherein the neural network processor is powered by at least one of the current sources in response to the operating switch being turned off.
Owner:NUVOTON

Network device for forwarding address and port mapping data packets with aid of network processor and / or software address and port mapping information table

The invention provides a network device which comprises a storage device and a network processor. The storage device stores an address and port mapping information table, the address and port mapping information table comprises a plurality of table entries for storing a plurality of second internet protocol headers, and each second internet protocol header conforms to a second internet protocol version. The network processor is configured to receive a first data packet having a first Internet Protocol header conforming to a first Internet Protocol version different from the second Internet Protocol version, and to generate and forward a second data packet. And the second data packet includes a second Internet Protocol header captured from the storage device and at least a portion of the first data packet.
Owner:AIROHA TECH (SUZHOU) LTD

Method for operating a ring communication network

ActiveUS12603797B2Data switching networksAuxiliary memoryNetwork communication
A method of operating a ring communication network, including providing a management device having a first processor, a first memory, a first communicator, and a computer program stored in the first memory and capable of running on the first processor; providing a network device having a network processor, a network memory, and a network communicator; disposing an auxiliary memory in the network device; organizing the management device and the network device to form the ring communication network; determining a location of the auxiliary memory in the ring communication network; and assigning a transmission direction such that a transmission distance between the management device and the auxiliary memory is minimized.
Owner:CATERPILLAR INC

Model deployment method, device, equipment and program product

The invention relates to the field of artificial intelligence, in particular to a model deployment method and device, equipment and a program product. The method comprises: using an operator and / or a structure adapted to a neural network processor, training a to-be-deployed model through a predetermined first training framework to obtain a trained model in a first format, the first training framework being different from a second training framework adapted to the neural network processor, and the second training framework being different from the second training framework; the model in the first format is a model obtained after training of a first training framework; performing parameter conversion and / or operator combination reconstruction on the model in the first format, and converting the model in the first format into a model in a second format which is a format supported by a neural network processor; and compiling the model in the second format, and deploying the compiled model to the edge device. As excessive computational nodes or intermediate conversion layers are not introduced in the method, the operator redundancy can be reduced, and the reasoning speed and the quantization precision can be improved.
Owner:UNILUMIN GRP

Lightweight deployment method of ai large model for complex scene

The application provides an AI large model lightweight deployment method for complex scenes, relates to the technical field of edge computing, and comprises the following steps: performing structural pruning on a pre-trained Transformer network based on attention head importance scores, dynamically sparsifying feedforward network activation states according to input tensor entropy values, adopting dynamic mixed precision quantization to obtain pruned and quantized network parameters; deploying the pruned and quantized network parameters to an edge computing device, distributing feature extraction operators to a neural network processor through a heterogeneous computing scheduler, and offloading classification operators to a multi-core central processing unit; combining a virtual memory paging mechanism to manage on-chip storage, using a direct memory access controller to realize zero-copy data transmission, and outputting inference result tensors. The application realizes efficient and reliable operation of a large model on a resource-limited edge node.
Owner:XIAN XINGXUN INTELLIGENT COMM TECH CO LTD