A method and system for lightweight deployment of a large model on an edge computing device

By constructing an acceleration capability profile and using an optical guide switch matrix to switch data transmission paths, the problem of low computational efficiency for large model deployments on edge computing devices is solved, achieving efficient and lightweight deployment and real-time inference, and improving inference efficiency to adapt to different hardware platforms.

CN120821479BActive Publication Date: 2025-12-05LUSTER LIGHTWAVE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511294327.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-12-05
Estimated Expiration
2045-09-11

AI Technical Summary

Technical Problem

In existing technologies, edge computing devices suffer from low computational efficiency when deploying large models. In particular, the transmission path relies on traditional buses, which causes communication to occupy too many inference cycles. The double buffering mechanism cannot break through the physical latency limit of hardware transmission. At the same time, the statically compiled quantized model is not adapted to the differences in device instruction sets, resulting in significant differences in inference efficiency.

Method used

By extracting the hardware instruction set architecture type and the number of parallel computing units of edge devices, an acceleration capability profile is constructed. A unified lookup table vectorization engine is used to divide the weights of the large model into multiple weight subsets by bit grouping, generating a pre-computation vector that matches the target instruction set architecture. Data transmission paths are switched during the hardware instruction execution cycle through a light guide switch matrix to eliminate data dependency conflicts.

Benefits of technology

It enables efficient lightweight deployment and real-time inference on edge computing devices, significantly reduces communication latency and computational burden, improves instruction parallel execution efficiency, overcomes traditional bus transmission bottlenecks, and adapts to the inference efficiency of different hardware platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821479B_ABST
    Figure CN120821479B_ABST
Patent Text Reader

Abstract

The application provides a lightweight deployment method and system of a large model on an edge computing device. In the application, the hardware instruction set architecture type and the number of parallel computing units of the target edge device are extracted to construct an acceleration capability portrait. Based on the instruction type, the large model weight is grouped and divided by a unified lookup table vectorization engine for precalculation, and a precalculation vector matching the target instruction set is generated. According to the number of parallel units and the instruction level parallelism capability, the precalculation vector is compiled to generate an adaptive parallel lookup table instruction block, the execution threads equal to the number of parallel units are allocated, and the data dependency conflict is eliminated. Finally, the instruction block is loaded into the shared memory area, the topology logic of the light guide switch matrix is configured based on the instruction type, and the data transmission path is dynamically switched within the hardware instruction cycle. The application realizes efficient deployment of the large model on the edge device and low-delay inference in the resource-constrained environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of lightweight AI deployment technology driven by edge computing, and in particular to a lightweight deployment method and system for large models on edge computing devices. Background Technology

[0002] In edge computing scenarios such as industrial IoT, autonomous driving, and medical terminals, to meet the requirements of real-time performance, high accuracy, and device universality, large models with tens of billions of parameters must be deployed to resource-constrained edge devices. This requires the deployment scheme to maintain model accuracy under low latency conditions, while being flexible and compatible with different edge hardware platforms, resolving the contradiction between the fragmentation of computing power at the edge and the computational intensity of large models.

[0003] Currently, the mainstream solution adopts a CPU-FPGA heterogeneous collaborative framework based on a double buffer mechanism. The model weights are compressed into INT8 format and stored in the FPGA memory through a fixed-precision quantization strategy. The processor schedules the computation task flow, and the programmable gate array executes the core matrix operation. The double buffer technology is used to preload data to reduce communication latency.

[0004] However, this framework has significant drawbacks. The reliance on a traditional bus for transmission paths results in communication accounting for over 40% of inference cycles, and the double-buffering mechanism cannot overcome the physical latency limitations of hardware transmission. Static compilation of the quantization model ignores instruction set differences, leading to inference efficiency differences of several times between devices with the same model on different architectures. Summary of the Invention

[0005] This application provides a lightweight deployment method and system for large models on edge computing devices to solve the problem of low computing efficiency of edge devices in the prior art.

[0006] Firstly, this application provides a lightweight deployment method for large models on edge computing devices, including:

[0007] Extract the hardware instruction set architecture type and the number of parallel computing units of the target edge device, identify instruction-level parallel capabilities based on the hardware instruction set architecture type, and construct an acceleration capability profile that includes instruction type and the number of parallel computing units.

[0008] Based on the instruction type in the acceleration capability profile, the large model weights are divided into multiple weight subsets by bit grouping through a unified lookup table vectorization engine. Pre-computation operations are performed on each weight subset to generate a pre-computation vector that matches the target instruction set architecture.

[0009] based on the number of parallel computing units and the instruction-level parallelism capability in the acceleration capability profile, compiling the pre-computed vector to generate a parallel lookup table instruction block adapted to the target instruction set architecture, assigning an execution thread equal to the number of parallel computing units to the parallel lookup table instruction block, reordering the pipeline operation sequence according to the instruction-level parallelism capability, and eliminating data dependency conflicts;

[0010] loading the parallel lookup table instruction block into a shared memory region of a processing unit and a field programmable gate array of the target edge device, configuring the topology connection logic of the optical switch matrix based on the instruction type in the acceleration capability profile, switching the data transmission path between the processing unit and the field programmable gate array through the optical switch matrix within a hardware instruction execution cycle, and completing the lightweight deployment and inference execution of the large model.

[0011] Optionally, according to the instruction type in the acceleration capability profile, the large model weight is divided into multiple weight subsets by bit grouping through the uniform lookup table vectorization engine, and a pre-computed vector matching the target instruction set architecture is generated by performing a pre-computation operation on each weight subset, including:

[0012] extracting the maximum parallel operation bit width from the instruction type in the acceleration capability profile, obtaining a pre-set bit grouping unit fixed length, dividing the maximum parallel operation bit width by the bit grouping unit fixed length to calculate the number of bit grouping units contained in a single weight subset;

[0013] cutting the large model weight sequence into a continuous unit sequence according to the bit grouping unit fixed length, and dividing the continuous unit sequence into multiple weight subsets by extracting an equal number of continuous units from the continuous unit sequence according to the number of bit grouping units;

[0014] for each weight subset, enumerating the complete value range of the input value in the uniform lookup table vectorization engine, and calculating the set of simulated output values of the input value on all bit grouping units in the weight subset for each input value of the weight subset within the complete value range;

[0015] concatenating the set of simulated output values in order to form a vectorization result with a bit width equal to the maximum parallel operation bit width, arranging the vectorization result in order of input value from small to large, and forming a pre-computed vector in a continuous storage region.

[0016] Optionally, for each weight subset, the complete value range of the input value is enumerated in the uniform lookup table vectorization engine, and the set of simulated output values of the input value on all bit grouping units in the weight subset is calculated for each input value of the weight subset within the complete value range, including:

[0017] determining an input value complete value range of the weight subset, the complete value range being defined by a binary value space boundary corresponding to a fixed length of the bit grouping unit;

[0018] traversing each input value in the complete value range in ascending order, taking the current input value as an input item, and inputting the uniform lookup table vector engine;

[0019] synchronously calculating independent output values of the input item on all bit grouping units in the weight subset, generating an output value sequence equal to the number of bit grouping units;

[0020] combining the output value sequence in the physical storage order of the bit grouping unit to form a set of analog output values corresponding to the current input value.

[0021] Optionally, the parallel lookup table instruction block is loaded into the shared memory region of the processing unit and the field programmable gate array of the target edge device, the topology connection logic of the light guide switch matrix is configured based on the instruction type in the acceleration capability profile, the data transmission path between the processing unit and the field programmable gate array is switched through the light guide switch matrix within a hardware instruction execution period, and the lightweight deployment and inference execution of the large model are completed, including:

[0022] loading the reorganized parallel lookup table instruction block into the shared memory region of the processing unit and the field programmable gate array of the target edge device, extracting all vector operation code types from the instruction type of the acceleration capability profile to generate a set of vector operation types;

[0023] Based on the set of vector operation types, the topology connection logic of the first data channel is configured in the light guide switch matrix, the channel connects the processing unit vector register and the field programmable gate array high-speed interface, and the topology connection logic of the second data channel is set, the channel connects the processing unit scalar register and the shared memory region;

[0024] Monitoring the opcode type in the hardware instruction execution period, when a vector opcode is detected, triggering the light guide switch matrix to switch to the first data channel, when a non-vector opcode is detected, triggering the light guide switch matrix to switch to the second data channel, and completing the lightweight deployment and inference execution of the large model.

[0025] Optionally, monitoring the opcode type in the hardware instruction execution period, when a vector opcode is detected, triggering the light guide switch matrix to switch to the first data channel, when a non-vector opcode is detected, triggering the light guide switch matrix to switch to the second data channel, and completing the lightweight deployment and inference execution of the large model, including:

[0026] A monitoring point of an operation code type is established after an instruction decoding stage, an operation code field of a current instruction is intercepted through a hardware logic circuit, a preset identification bit sequence of the operation code field is extracted, and the identification bit sequence is matched with a predefined vector operation code identifier set;

[0027] When the identification bit sequence matches the vector operation code identifier set, it is determined that a current instruction type is a vector operation code, and a first path switching control signal is generated;

[0028] When the identification bit sequence does not match the vector operation code identifier set, it is determined that the current instruction type is a non-vector operation code, and a second path switching control signal is generated;

[0029] The path switching control signal is input into a physical channel selector of a light guide switch matrix, wherein the first path switching control signal triggers a physical connection of a first transmission path, and the second path switching control signal triggers a physical connection of a second transmission path, to complete lightweight deployment and inference execution of a large model.

[0030] Optionally, a hardware instruction set architecture type and a number of parallel computing units of a target edge device are extracted, an instruction level parallelism capability is identified based on the hardware instruction set architecture type, and an acceleration capability image containing an instruction type and a number of parallel computing units is constructed, including:

[0031] An identification code of the hardware instruction set architecture type is read from a system configuration storage area of the target edge device, all operation codes and bit width parameters supported by the hardware instruction set architecture type are obtained by querying a predefined architecture capability mapping table according to the identification code;

[0032] The number of independent operation cores in a processing unit is collected through an operation unit management interface of the target edge device, and the number of configurable computing blocks is collected through a configuration state interface of a field programmable gate array;

[0033] Based on the operation codes and the bit width parameters, an instruction level parallelism capability feature is identified, the number of independent operation cores is added to the number of configurable computing blocks to obtain a total number of parallel computing units;

[0034] The all operation codes and bit width parameters, the instruction level parallelism capability feature, and the total number of parallel computing units are fused to construct the acceleration capability image containing the instruction type and the number of parallel computing units.

[0035] Optionally, based on the number of parallel computing units and the instruction-level parallelism capability in the acceleration capability profile, the pre-computed vector is compiled to generate a parallel lookup table instruction block adapted to the target instruction set architecture, an execution thread equal to the number of parallel computing units is assigned to the parallel lookup table instruction block, the order of pipeline operations is reorganized according to the instruction-level parallelism capability, and data dependency conflicts are eliminated, including:

[0036] According to the instruction-level parallelism capability of the acceleration capability profile, the continuous memory distribution characteristics of the pre-computed vector are analyzed, and the segments of the pre-computed vector stored continuously are compiled into the same instruction block based on the number of parallel computing units, and an equal number of independent instruction blocks are generated;

[0037] In each of the independent instruction blocks, a lookup table operation code of the target instruction set architecture and a corresponding pre-computed vector start address offset are filled in to obtain a parallel lookup table instruction block adapted to the target instruction set architecture, and an execution thread identifier equal to the number of parallel computing units is created, and each of the execution thread identifiers is assigned to the parallel lookup table instruction block;

[0038] The output-input relationship between all the parallel lookup table instruction blocks is scanned, and instruction block pairs with data dependency are identified, the execution time sequence order of the identified instruction block pairs is adjusted, a synchronization control operation code is inserted to eliminate data dependency conflicts, and the instruction block sequence is rearranged according to the adjusted time sequence order.

[0039] In a second aspect, the present application provides a lightweight deployment system of a large model on an edge computing device, including:

[0040] The construction module is configured to extract the hardware instruction set architecture type and the number of parallel computing units of the target edge device, identify the instruction-level parallelism capability based on the hardware instruction set architecture type, and construct an acceleration capability profile containing the instruction type and the number of parallel computing units;

[0041] The matching module is configured to divide the large model weight into a plurality of weight subsets according to the bit grouping based on the instruction type in the acceleration capability profile through a uniform lookup table vectorization engine, perform a pre-computation operation on each of the weight subsets, and generate a pre-computed vector matched to the target instruction set architecture;

[0042] The generation module is configured to compile the pre-computed vector based on the number of parallel computing units and the instruction-level parallelism capability in the acceleration capability profile, generate a parallel lookup table instruction block adapted to the target instruction set architecture, assign an execution thread equal to the number of parallel computing units to the parallel lookup table instruction block, reorganize the order of pipeline operations according to the instruction-level parallelism capability, and eliminate data dependency conflicts;

[0043] The switching module is configured to load the parallel lookup table instruction block to a shared memory region of a processing unit and a field programmable gate array of the target edge device, configure topology connection logic of a light guide switch matrix based on instruction types in the acceleration capability profile, switch a data transmission path between the processing unit and the field programmable gate array through the light guide switch matrix in a hardware instruction execution cycle, and complete lightweight deployment and inference execution of the large model.

[0044] In a third aspect, the present application provides a computing device including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the lightweight deployment method of a large model on an edge computing device as described in the first aspect above.

[0045] In a fourth aspect, the present application provides a computer storage medium storing a computer program, which, when executed by a computer, implements the lightweight deployment method of a large model on an edge computing device as described in the first aspect.

[0046] The present application extracts the hardware instruction set architecture type and the number of parallel computing units of the target edge device, identifies instruction-level parallel capability based on the architecture type, and constructs an acceleration capability profile. The effect is to quantify the hardware computing power characteristics of the device in depth, accurately identify the parallel potential of the instruction set, and provide hardware customization optimization basis for model deployment. According to the instruction types in the acceleration capability profile, the large model weight is divided into bit groups by the unified lookup table vectorization engine, and a pre-computed vector matching the target instruction set is generated for each weight subset. This step significantly reduces the real-time computing burden, and converts complex operations into efficient lookup table operations through the pre-computation strategy. Based on the number of parallel computing units and the instruction-level parallel capability in the acceleration capability profile, the pre-computed vector is compiled to generate a parallel lookup table instruction block that adapts to the instruction set architecture, allocates an equal number of execution threads, and reorganizes the pipeline operation sequence. This process significantly improves the instruction parallel execution efficiency, and completely eliminates data dependency conflicts through dynamic thread allocation and pipeline reorganization. The parallel lookup table instruction block is loaded to the shared memory region, the topology connection logic of the light guide switch matrix is configured based on the instruction types, and the data transmission path is dynamically switched in the hardware instruction execution cycle. Finally, ultra-low delay communication at the hardware level is realized, the traditional bus transmission bottleneck is overcome, and lightweight deployment and real-time inference are completed.

[0047] Further, by extracting the maximum parallel operation bit width from the instruction type of the acceleration capability profile, the number of bit grouping units contained in a single weight subset is calculated in combination with the preset bit grouping unit fixed length. This ensures that the weight grouping strictly matches the maximum parallel bit width of the hardware instruction, laying the foundation for single-cycle multi-data parallel processing. The large model weight sequence is cut into a continuous unit sequence according to the fixed length, and an equal number of continuous units are intercepted according to the number of bit grouping units and divided into weight subsets. This operation forms a continuous and regular memory storage structure, significantly improving cache hit efficiency and reducing memory fragmentation access overhead. Enumerate the complete value range of the input value for each weight subset, and calculate the simulated output set of each input value on all bit grouping units. This pre-computation mechanism completely eliminates the need for multiplication and addition operations in real-time inference, converting dynamic computation tasks into static lookup operations. The simulated output set is spliced into a vectorized result with the same width as the maximum parallel operation bit width in order, and arranged in the order of the input value to form a pre-computed vector stored continuously. Finally, the memory access mode is deeply adapted to the continuous addressing characteristics of the hardware, significantly reducing memory access delay and improving computing throughput.

[0048] These and other aspects of the present application will become more apparent from the following description of the embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0050] Figure 1 A flow chart of a method for lightweight deployment of a large model on an edge computing device is shown;

[0051] Figure 2 A scenario diagram of a method for lightweight deployment of a large model on an edge computing device is shown;

[0052] Figure 3 A structural schematic diagram of a lightweight deployment system of a large model on an edge computing device is shown;

[0053] Figure 4 A structural schematic diagram of a computing device is shown. DETAILED DESCRIPTION

[0054] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application.

[0055] In some of the flowcharts depicted in the description and claims of the present application and in the above mentioned figures, a plurality of operations is included which occur in a specific order, but it should be clearly understood that these operations can be executed not in the order in which they appear herein or in parallel, and the serial numbers of the operations such as 101, 102, etc. are merely used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these flowcharts can include more or fewer operations, and the operations can be executed in sequence or in parallel. It should be noted that the descriptions herein such as "first", "second", etc. are used to distinguish different messages, devices, modules, etc. and do not represent the order of precedence, nor do "first" and "second" represent different types.

[0056] Researchers found that existing edge computing devices have significant bottlenecks when deploying large models. The transmission path relies on traditional bus communication to occupy a large number of inference cycles, and the double buffering mechanism is difficult to break through the physical delay constraints of hardware transmission; at the same time, the static compiled quantization model is not adapted to the differences in device instruction sets, resulting in significant gaps in inference efficiency for the same model on heterogeneous devices. Therefore, there is an urgent need for a lightweight deployment method that can dynamically adapt to edge hardware features and cooperatively optimize communication and computing load.

[0057] To solve the above problems, the present application proposes a hardware adaptive instruction level cooperative deployment method, the core of which is to dynamically couple hardware resources and computing tasks through acceleration capability profiling. Specifically, first, the hardware instruction set architecture features and parallel unit scale of the edge device are extracted, and an acceleration capability profile containing instruction types and parallel computing capabilities is constructed; using the profile, the model weight is vectorized as a bit grouping vector, generating a pre-computed vector matching the target architecture; then, according to the actual parallel capability, a parallel lookup table instruction block is compiled, and execution threads are dynamically allocated and reorganized into a pipeline operation according to the instruction level parallel feature; finally, through the light guide switch matrix, the data transmission path between CPU and FPGA is intelligently switched within the hardware execution cycle. This method eliminates the communication delay bottleneck through hardware-aware instruction level optimization, solves the efficiency gap caused by instruction set differences through architecture adaptive pre-compilation mechanism, and reconstructs the pipeline based on real-time parallel capability to eliminate data dependency conflicts, achieving end-to-end performance leap in lightweight deployment.

[0058] The technical solutions in the embodiments of the present application will be described clearly and completely in the embodiments of the present application combined with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0059] Figure 1A flowchart of a method for lightweight deployment of a large model on an edge computing device is provided for embodiments of the present application, as shown in Figure 1 The method comprises:

[0060] 101. Extracting the hardware instruction set architecture type and the number of parallel computing units of the target edge device, identifying the instruction-level parallelism capability based on the hardware instruction set architecture type, and constructing an acceleration capability profile containing the instruction type and the number of parallel computing units.

[0061] Optionally, step 101 can specifically include the following steps:

[0062] 1011. Read the identification code of the hardware instruction set architecture type from the system configuration storage area of the target edge device, and according to the identification code, query the pre-defined architecture capability mapping table to obtain all operation codes and bit width parameters supported by the hardware instruction set architecture type;

[0063] 1012. Collect the number of independent operation cores in the processing unit through the operation unit management interface of the target edge device, and collect the number of configurable computing blocks through the configuration state interface of the field programmable gate array;

[0064] 1013. Based on the operation code and the bit width parameter, identify the instruction-level parallelism capability feature, add the number of independent operation cores and the number of configurable computing blocks to obtain the total number of parallel computing units;

[0065] 1014. Fuse the all operation codes and bit width parameters, instruction-level parallelism capability features, and total number of parallel computing units to construct an acceleration capability profile containing instruction type and number of parallel computing units.

[0066] In the above steps, the target edge device refers to the computing hardware deployed in the edge computing location, such as a smart camera or a router; the hardware instruction set architecture type represents the instruction set format supported by the device processor, such as ARM or x86, which is used to define the basic operation capability of the device; the number of parallel computing units refers to the total number of units that can be executed independently by the device at the same time; the instruction-level parallelism capability describes the way the hardware supports simultaneous execution of multiple instructions; constructing an acceleration capability profile is to create a data structure to record the computing performance features of the device; the operation code is the type of processor instruction, such as the code corresponding to the addition or subtraction operation; the bit width parameter is the number of data bits operated by the processing unit at a time, such as 32 bits representing the processing of 4 bytes of data; the independent operation core is an independent physical processing unit in the processor that can run tasks independently; the configurable computing block is a user programmable functional block in the field programmable gate array, which is used to customize computing tasks.

[0067] In the embodiments of the present application, first, the identification code of the hardware instruction set architecture type is read and the operation code and bit width parameter are obtained through step 1011. For example, the target edge device is an intelligent industrial sensor, and the system configuration storage area stores the identification code as "ARMv8". The pre-defined architecture capability mapping table is an electronic file storing data of different architecture types, and the query process uses a string matching algorithm. In a specific implementation, the device startup script automatically accesses the system configuration storage area to extract the identification code string "ARMv8". Then, the preloaded mapping table file is searched in the local file system, and the corresponding row is accurately matched according to the identification code. All operation codes supported by the architecture are read, such as ADD / SUB / MUL and bit width parameters (64 bits). The obtained data will be output as an operation code list and a bit width value for subsequent steps.

[0068] Secondly, the number of independent operation cores of the processing unit and the number of configurable computing blocks of the field programmable gate array are collected through step 1012. For example: the processing unit management interface calls the "getCoreCount" function of the operating system, and returns the number value, such as 4, indicating that the device has 4 processing cores; the field programmable gate array configuration status interface calls the "fpgaConfigStatus" function of the hardware driver to obtain the number of configurable computing blocks, such as 8, indicating that there are 8 available computing blocks. The collection process is automatically executed by the background service program at the device startup. Finally, the number of independent cores 4 and the number of computing blocks 8 are output, and these two values will be added to calculate the total number of parallel units in the next step.

[0069] Then, two operations are performed through step 1013: identifying the instruction-level parallelism capability feature based on the operation code and bit width parameter; and adding the number of independent operation cores and the number of configurable computing blocks. For example: when the operation code contains the vector instruction "VADD" and the bit width parameter is 64 bits, the identification process uses the bit width comparison and operation code scanning algorithm. First, the operation code set is scanned to detect whether there is a vector instruction code, such as VADD; then it is checked whether the bit width is ≥ 64 bits. If both conditions are met, the instruction-level parallelism capability is marked as supporting single instruction multiple data stream operation (SIMD). The value addition operation needs to be verified before execution: if they are both positive integers, such as core number 4 and computing block number 8, then 4+8=12 is calculated; if they are not integers, they are reset to 0. Finally, the instruction parallelism capability feature SIMD and the total number of parallel units 12 are output.

[0070] Finally, the three-part data is fused to construct the acceleration capability profile through step 1014. For example, the supported operation codes (ADD / SUB / VADD), the bit width parameter (64 bits), the instruction-level parallelism capability feature (SIMD), and the number of parallel units (12) are combined into structured data. The fusion process calls the encapsulation function to create a JSON object, and the key-value pair structure is: {"operationcodes":["ADD","SUB","VADD"],"bit_width":64,"paralleltype":"SIMD","unitcount":12}. The constructed profile is stored as a device local configuration file, such as devicecapability.json, for direct calling by the subsequent task scheduling module.

[0071] In actual application, in a certain edge computing scenario, the target edge device A is applied to the real-time analysis task of the intelligent security system, the identification code of the hardware instruction set architecture type such as architecture type B is read from the device system configuration storage area, the pre-defined architecture capability mapping table is queried, all operation codes supported by the type such as integer multiplication, floating point addition, vector loading and bit operation, and the corresponding bit width parameters such as 32 bits and 64 bits are obtained; then, the number of independent operation cores in the processing unit is collected through the operation unit management interface as 4, and the number of configurable computing blocks is collected through the configuration state interface of the field programmable gate array as 6; subsequently, the instruction-level parallelism capability feature such as supporting data-level parallelism is identified based on the obtained operation codes and bit width parameters, and the number of independent operation cores and the number of configurable computing blocks are added to obtain the number of overall parallel computing units as 10; finally, all operation codes and bit width parameters, instruction-level parallelism capability features, and the number of overall parallel computing units are fused to construct the acceleration capability profile containing instruction types and the number of parallel computing units, thereby enhancing the adaptability of the device to the computing task.

[0072] In the overall scheme of the above step 101, the supported operation codes and bit width parameters are obtained by reading the hardware instruction set architecture identification code stored by the target edge device and querying the pre-defined mapping table, the number of independent operation cores of the processor is collected through the operation unit management interface, and the number of configurable computing blocks of the field programmable gate array is collected through the configuration state interface; the instruction-level parallelism capability feature is identified based on the operation codes and bit width parameters, and the total number of overall parallel computing units is obtained by comprehensively counting the number of two kinds of operation units; finally, the instruction type feature and the number of parallel computing units are fused to form a complete hardware acceleration capability profile, realizing multi-dimensional quantification representation of the computing architecture characteristics of the edge device.

[0073] 102、according to the instruction type in the acceleration capability profile, the unified lookup table vectorization engine divides the large model weight into multiple weight subsets by bit grouping, performs pre-computation operation on each weight subset, and generates a pre-computed vector matching the target instruction set architecture;

[0074] Optionally, step 102 can specifically include the following steps:

[0075] 1021、extract the maximum parallel operation bit width from the instruction type of the acceleration capability profile, obtain the preset bit grouping unit fixed length, divide the maximum parallel operation bit width by the bit grouping unit fixed length to calculate the number of bit grouping units contained in a single weight subset;

[0076] 1022、cut the large model weight sequence into a continuous unit sequence according to the bit grouping unit fixed length, and according to the number of bit grouping units, cut equal amounts of continuous units from the continuous unit sequence to divide into multiple weight subsets;

[0077] 1023、for each weight subset, enumerate the complete value range of the input value in the unified lookup table vectorization engine, and for each input value of the weight subset in the complete value range, calculate the set of simulated output values of the input value on all bit grouping units in the weight subset;

[0078] Specifically, step 1023 can include the following process: determine the complete value range of the input value of the weight subset, which is defined by the binary value space boundary corresponding to the bit grouping unit fixed length; traverse each input value in the complete value range in ascending order, input the current input value into the unified lookup table vectorization engine as an input item; simultaneously calculate the independent output values of the input item on all bit grouping units in the weight subset within the complete value range, and generate an output value sequence equal to the number of bit grouping units; combine the output value sequence in the physical storage order of the bit grouping unit to form the set of simulated output values corresponding to the current input value.

[0079] 1024、concatenate the set of simulated output values in order to form a vectorization result with a bit width equal to the maximum parallel operation bit width, arrange the vectorization result in ascending order of input values, and form a pre-computed vector in a continuous storage area.

[0080] In the above steps, the acceleration capability profile is a performance data structure extracted from the target edge device, containing instruction type information and computing capability data supported by the device; the large model weight refers to a sequence of parameter values in a large machine learning model, used for model inference calculation; the unified lookup table vectorization engine is a computing system module that implements efficient computation and parallel data processing operations based on pre-stored lookup tables; bit grouping is a data division method that splits the large model weight sequence into fixed-size binary bit group units; the maximum parallel operation bit width is the maximum number of bits that the device instruction set can process at a time, such as 64 bits indicating that the device can process 64 bits of data simultaneously; the bit grouping unit fixed length is a pre-set size of the bit grouping unit, for example, 8 bits as a basic unit; the bit grouping unit quantity is an integer value calculated to represent the number of fixed-length units contained in each weight subset; the weight subset is a collection of sub-units obtained by dividing the large model weight sequence, each subset containing the same number of bit grouping units; the input value complete value range is the entire numerical interval that defines the possible input values of the bit grouping unit, with the fixed length determining the boundary, for example, an 8-bit unit corresponds to the integer range of 0 to 255; the simulated output value set is a pre-computed sequence of simulation results on each weight subset, with each result corresponding to an input value; the pre-computed vector is a wide-bit vector data formed by concatenating all simulated output values and arranging them in input value order, matching the device instruction set bit width for easy acceleration calculation.

[0081] In the embodiments of the present application, first, the maximum parallel operation bit width value is extracted from the acceleration capability profile through step 1021, and the pre-set bit grouping unit fixed length value is obtained; then, the bit grouping unit quantity is calculated, which is an integer operation implemented by a division algorithm, and the specific formula is defined as: the bit grouping unit quantity is equal to the maximum parallel operation bit width divided by the bit grouping unit fixed length, i.e. wherein represents the bit grouping unit quantity, represents the maximum parallel operation bit width, represents the bit grouping unit fixed length; if the division result is a decimal number, the system automatically performs a floor function to convert it to an integer to ensure that the unit quantity is an integer value; the entire calculation process is implemented by calling mathematical operation functions by the system background service, for example, the acceleration capability profile of the target edge device shows that the maximum parallel operation bit width = 64 bits, and the pre-set bit grouping unit fixed length = 8 bits; the system first reads the and values, and then applies the formula to calculate the unit quantity = 8, this value is directly passed to the next step for dividing the weight sequence, the system will check whether the value is a positive integer to prevent incorrect input.

[0082] Secondly, by step 1022, based on the bit grouping unit fixed length and the unit number value, the large model weight sequence is cut and divided into multiple weight subsets; the implementation process involves data cutting algorithm and sequence truncation operation, first, load the large model weight sequence, such as binary data stream or numerical array, according to the fixed length, it is divided into multiple equal length continuous unit sequence; for example, a large model weight sequence length 1024 bits, fixed length is set to 8 bits, then the cutting process is executed 1024 divided by 8 equals 128, get 128 continuous 8 bit unit sequence; then, from the beginning of the sequence, the unit number calculated by step 1021 is Truncate 8 continuous units as a group of weight subsets, such as truncating units 1 to 8 as the first group, and then continue to truncate units 9 to 16 as the second group, and execute in a loop until the end of the sequence, finally, generate 128 divided by 8 equals 16 complete weight subsets; the division process uses memory slice function to ensure data boundary alignment, each weight subset is stored as an independent array block for subsequent processing.

[0083] Then, by step 1023, pre-compute operation is performed on each weight subset to generate a set of simulation output values; the execution process includes the following steps: first, determine the complete value range of the input value, the maximum value of the range is calculated as , wherein, represents the maximum value of the value range, represents the fixed length of the bit grouping unit, the minimum value is fixed as 0; for example, the fixed length = 8, the calculation process is = 28 1= 256 1= 255, the value interval is all integers from 0 to 255; then, complete traversal of the value range in ascending order of integers, such as from 0 to 255, each input value is passed into the uniform lookup table vectorization engine as an input item; the engine calculates the independent output value of the input item on all units of the weight subset based on the pre-stored calculation rule such as multiplication or addition operation, the output value sequence length is equal to the unit number obtained by step 1021 ; finally, combine the output sequence to generate the simulation output value set of the current input value according to the physical storage order of the bit grouping unit; for example, a weight subset contains =8 single length 8-bit, unit weight value is 01 respectively, 01 represents decimal 1 value 02 represents decimal 2, and so on, when the input value is 5, the engine applies 5 to each unit weight 01 at the same time, 01 outputs 5 times 1 equals 5, 02 outputs 5 times 2 equals 10, and so on, 8 times are calculated synchronously, forming a sequence, such as 5, 10, 15, 20, 25, 30, 35, 40, and combined as an output value set 5, 10, 15, 20, 25, 30, 35, 40 in the order of units; The whole process adopts parallel processing technology to reduce delay and ensure efficiency.

[0084] Finally, the simulation output value set of each input value is spliced into a vectorized result with bit width matching by step 1024, and arranged in the order of input values to form a pre-computed vector; The implementation process includes vector splicing and data sorting operations, first, for each input value corresponding to the simulation output value set generated by step 1023, such as set 5, 10, 15, 20, 25, 30, 35, 40, perform a splicing algorithm, that is, all output value sequences are converted to binary and merged into a single wide-bit vector, the bit width is equal to the maximum parallel operation bit width extracted by step 1021, such as bit; For example, the output value sequence of input value 5 has 8 outputs, each with 8 bits, which are spliced into a 64-bit binary sequence such as 0101001010100010, and a complete 64-bit vector; Then, arrange all vectorized results in the order of input values from small to large, such as 0 to 255, use a linear array write function to store the results to a continuous storage area such as a memory buffer or a file to form the final pre-computed vector; For example, store all input value vectors in the order of 0, 1, 2, etc. in an array for subsequent model inference to directly load and utilize device parallelism to improve efficiency.

[0085] In practical applications, in the image recognition edge deployment scenario of an agricultural Internet of Things system, when optimizing the inference of a large model based on the acceleration capability profile of device A, which contains vector load instructions and 10 parallel units, first, the maximum parallel operation bit width of 128 bits supported is parsed from the profile, and the fixed length of the 32-bit grouping unit is combined to calculate that a single weight subset needs to contain 4 grouping units; then the crop disease recognition model weight sequence to be deployed is continuously cut into 70000 basic unit sequences in units of 32 bits, and the sequence is divided into 17500 weight subsets according to the principle of one group per 4 units and distributed to different parallel computing units for processing; for each weight subset, a complete input value enumeration space (0 to 4294967295) is established in the unified lookup table vectorization engine, and a parallel pipeline mechanism is used to synchronously calculate the operation results of each input value on the 4-bit grouping units, for example, when the input value is 137, the engine synchronously outputs the convolution kernel mapping value of unit 1, the activation function prefetch value of unit 2, the feature scaling factor of unit 3, and the weight compensation parameter of unit 4, and combines them into a 4-dimensional output sequence in physical storage order; finally, the output sequence corresponding to each input value is spliced into a 128-bit vectorization result, and a structured precomputed vector matrix is generated in the continuous storage area in the order of 0 to 4294967295, so that the special instruction set of edge device C can directly call the vector to perform single instruction multiple data stream operations.

[0086] In the overall scheme of the above step 102, based on the instruction type characteristics in the acceleration capability profile, the maximum parallel operation bit width is extracted, and the number of bit grouping units contained in a single weight subset is calculated in combination with the fixed length of the preset bit grouping unit; after cutting the large model weight sequence into a continuous unit sequence in a fixed length, a plurality of weight subsets are divided according to the calculated number of units; for each weight subset, the complete input value range defined by the bit grouping unit length is determined in the unified lookup table vectorization engine, each input value in the range is traversed in ascending order, the independent output values generated by the input value acting on all bit grouping units in the weight subset are synchronously calculated, an output value sequence matching the number of units is formed, and an analog output value set is combined in physical storage order; finally, the set is converted into a vectorization result with a bit width equal to the maximum parallel operation bit width, and the precomputed vector is arranged in a continuous storage area in the input value ascending order, to realize efficient instruction-level parallel computing structure matching the large model weight. This process reorganizes the weight structure and optimizes the precomputation driven by hardware instruction characteristics, significantly improves the model inference efficiency on edge devices, and reduces the runtime repeated computation overhead.

[0087] 103. based on the number of parallel computing units and the instruction-level parallelism capability in the acceleration capability profile, compiling the pre-computed vector to generate parallel lookup table instruction blocks adapted to the target instruction set architecture, assigning an equal number of execution threads to the parallel computing units to the parallel lookup table instruction blocks, reordering the pipeline operation sequence according to the instruction-level parallelism capability, and eliminating data dependency conflicts;

[0088] Optionally, step 103 can specifically include the following steps:

[0089] 1031. according to the instruction-level parallelism capability in the acceleration capability profile, analyzing the continuous memory distribution characteristics of the pre-computed vector, taking memory address continuity as the division basis, compiling the continuously stored segments of the pre-computed vector to the same instruction block, and generating an equal number of independent instruction blocks based on the number of parallel computing units;

[0090] 1032. filling the lookup table operation codes of the target instruction set architecture and the corresponding pre-computed vector start address offsets in each of the independent instruction blocks to obtain parallel lookup table instruction blocks adapted to the target instruction set architecture, creating an equal number of execution thread identifiers to the number of parallel computing units, and assigning each of the execution thread identifiers to the parallel lookup table instruction blocks;

[0091] 1033. scanning the output-input relationships among all the parallel lookup table instruction blocks, identifying instruction block pairs with data dependencies, adjusting the execution time sequence order of the identified instruction block pairs, inserting synchronization control operation codes to eliminate data dependency conflicts, and rearranging the instruction block sequence according to the adjusted time sequence order.

[0092] In the above steps, the acceleration capability profile is a performance data structure extracted from the target edge device, containing the number of parallel computing units and instruction-level parallelism capability features; the pre-computed vector is the vectorized data generated in step 102; the compiled pre-computed vector is the conversion of the data into instruction code executable by the device; the parallel lookup table instruction block adapted to the target instruction set architecture is the compiled instruction unit that enables the device to efficiently query the pre-computed vector; the execution thread allocation is the allocation of independent processing units for each instruction block; the instruction-level parallelism capability is the feature of hardware supporting simultaneous execution of multiple instructions such as single instruction multiple data stream operations; the number of parallel computing units is the total number of units that the device can handle simultaneously; the reordering pipeline operation sequence is the adjustment of the instruction execution order to optimize the flow; the data dependency conflict is the conflict that occurs when an instruction requires the output of another instruction; the contiguous memory distribution feature describes the way the pre-computed vector is stored continuously in memory; the independent instruction block is the instruction unit that can run independently after compilation; the execution thread identifier is a symbol used to uniquely mark each processing thread; the data-dependent instruction block pair is a combination of instruction blocks that have an input-output relationship; the synchronization control operation code is a code unit inserted into the instruction to ensure sequential execution; the starting address offset is the position offset value of the pre-computed vector in memory; and the lookup table operation code is a data query operation code supported by the device instruction set.

[0093] In the embodiments of the present application, first, the continuous memory distribution feature of the pre-computed vector is analyzed by step 1031, and an equal number of independent instruction blocks is generated based on the number of parallel computing units. For example, the pre-computed vector is stored in a continuous memory region such as an array data block with addresses 1000 to 2000. The system calls the analysis function to scan the memory address continuity, checks whether the starting address and ending address of each segment are continuous, uses the address difference calculation algorithm to determine the segment size, and uses the padding function to align to ensure continuity if the addresses are not continuous. Then, based on the instruction-level parallelism capability of the acceleration capability profile such as single instruction multiple data stream operation support, the system uses a segmentation algorithm with memory address continuity as the basis for division. The system divides the pre-computed vector segments from the beginning of the address. The size of each segment is determined by dividing the total length by the number of parallel computing units. For example, a pre-computed vector length of 1024 bytes and a number of parallel computing units of 8, the system calculates the size of each segment as 1024 divided by 8, which equals 128 bytes. Then, the segments are divided into addresses 1000 to 1127 as the first instruction block corresponding to the memory, and the subsequent is similarly divided to generate 8 complete instruction blocks. All instruction blocks are stored in the instruction cache area for subsequent use. The entire process ensures that the segments are continuous and non-overlapping.

[0094] Secondly, fill in the table lookup operation code and the start address offset to get the parallel table lookup instruction block, and create the allocation execution thread identifier. For example, in each independent instruction block, the system loads the target instruction set architecture such as ARM supported table lookup operation code such as data loading instruction LDR; then, fill in the corresponding pre-computed vector start address offset, such as the first instruction block offset 0, and the subsequent increases by 128 bytes per segment size, so the second instruction block offset is 128. The calculation process uses the memory address calculation function, and the start address plus the segment size obtains the real address value written to the instruction block; then, the system creates the execution thread identifier number equal to the parallel computing unit number such as 8, uses the thread management function, and generates unique identifiers such as integers 1 to 8, which are marked as thread 1, thread 2, etc.; finally, allocate each thread identifier to a parallel table lookup instruction block, such as thread 1 to the first instruction block, thread 2 to the second instruction block, and so on. The entire allocation process is based on the sequential mapping algorithm, which ensures that each thread has a corresponding instruction block. All instruction block and thread information are stored in the thread queue for execution.

[0095] Finally, through step 1033, scan the output-input relationship of all parallel table lookup instruction blocks, identify the data-dependent instruction block pair, adjust the timing to eliminate data dependency conflicts, and rearrange the instruction block sequence. For example, the system calls the scanning function to check the output and input of each instruction block, such as output variables and input variables, and uses the relationship detection algorithm to compare variable names or address ranges. If it is found that the output result of instruction block 1 is used as input by instruction block 2, it is identified as a pair of instruction blocks with data dependency; then, adjust the execution timing sequence. For the identified instruction block pair, the system uses the scheduling function to force instruction block 1 to execute before instruction block 2, ensuring that the output precedes the input. Other non-dependent blocks remain parallel execution; then, insert synchronization control operation codes between dependent pairs, such as inserting the waiting control instruction SYNC, to ensure that data is ready before executing the subsequent block; finally, rearrange the sequence of all instruction blocks according to the adjusted timing, and use the sequence reorganization function to write the new order to the execution queue, such as the first-order parallel structure to eliminate conflicts; for example, among the three instruction blocks, block A and block B are dependent, and block C is independent. The system identifies A to B as a dependent pair, adjusts the order to execute A first and then B, inserts the SYNC operation code before B, and finally, the sequence is arranged as A followed by the synchronization instruction, then B, and finally C is stored. This can be directly used in the device execution pipeline to improve efficiency.

[0096] In practical applications, during the edge inference deployment of an industrial AI model, device A performs compilation optimization on the pre-computed vector based on the 10 parallel computing units and data-level parallel instruction capabilities in the acceleration capability profile. First, the continuous memory distribution characteristics of the pre-computed vector are analyzed according to the instruction-level parallelism feature. The vector is distributed in a linear storage form in a 256 KB memory region with a starting address of 0x1000. According to the address continuity principle, the vector is divided into 10 independent segments with equal capacity, each occupying a 25.6 KB storage space, and 10 independent instruction blocks are generated. Then, the specific lookup table operation code of the target instruction set architecture, such as VLOAD, is filled in each instruction block, and the corresponding memory segment starting address offset is associated, for example, instruction block 1 is associated with 0x1000 address offset, and instruction block 2 is associated with 0x19600 address offset, thereby forming a parallel lookup table instruction block set adapted to the architecture. At the same time, 10 execution thread identifiers numbered 1 to 10 are created and assigned to each instruction block in a 1:1 mapping relationship. The subsequent process scans the input-output relationship of all instruction blocks, identifies the instruction block combination with data dependency, and a typical case is that the output data of instruction block 3 is used as the input dependent source of instruction block 7. Through the timing adjustment mechanism, the execution order of instruction block 7 is moved forward to before instruction block 3, and a SYNC synchronization control operation code is inserted at the key node. Finally, the reconstructed pipeline execution sequence is arranged in the order of instruction block 1, instruction block 2, instruction block 7, instruction block 3, instruction block 4, etc. The parallel lookup table instruction block sequence that eliminates data conflicts significantly improves the utilization of computing resources.

[0097] In the overall scheme of the above step 103, relying on the number of parallel computing units and instruction-level parallelism capability characteristics in the acceleration capability profile, equal independent instruction blocks are generated according to the continuous memory distribution characteristics of the pre-computed vector based on address continuity; the lookup table operation code of the target instruction set architecture and the starting address offset of the pre-computed vector are filled in each instruction block to form parallel lookup table instruction blocks, and execution thread identifiers matching the number of parallel computing units are created and assigned to each instruction block; the output-input relationship between instruction blocks is scanned to identify instruction block pairs with data dependency, the execution timing order is adjusted, and a synchronization control operation code is inserted to eliminate conflicts, and finally the pipeline structure is optimized according to the reorganized instruction block sequence. This process fully utilizes the characteristics of hardware parallel resources to achieve deep optimization of lookup table operations, effectively avoids data waiting delay between computing units, and maximizes the inference throughput of large models on edge devices.

[0098] 104. loading the parallel lookup table instruction block to a shared memory region of a processing unit and a field programmable gate array of a target edge device, configuring topology connection logic of a light guide switch matrix based on instruction types in the acceleration capability profile, switching data transmission paths between the processing unit and the field programmable gate array through the light guide switch matrix in a hardware instruction execution cycle, completing lightweight deployment and inference execution of a large model.

[0099] Optionally, step 104 can specifically include the following steps:

[0100] 1041. loading the reorganized parallel lookup table instruction block to a shared memory region of a processing unit and a field programmable gate array of a target edge device, extracting all vector operation code types from instruction types in the acceleration capability profile to generate a vector operation type set;

[0101] 1042. configuring topology connection logic of a first data channel in the light guide switch matrix based on the vector operation type set, channel connecting a processing unit vector register and a field programmable gate array high-speed interface, setting topology connection logic of a second data channel, channel connecting a processing unit scalar register and a shared memory region;

[0102] 1043. monitoring operation code types in a hardware instruction execution cycle, triggering the light guide switch matrix to switch to the first data channel when a vector operation code is detected, triggering the light guide switch matrix to switch to the second data channel when a non-vector operation code is detected, completing lightweight deployment and inference execution of a large model.

[0103] Wherein, step 1043 can specifically include the following process: establishing an operation code type monitoring point after the instruction decoding stage, intercepting the operation code field of the current instruction through a hardware logic circuit, extracting a preset identifier bit sequence of the operation code field, and matching the identifier bit sequence with a predefined vector operation code identifier set; when the identifier bit sequence matches the vector operation code identifier set, it is determined that the current instruction type is a vector operation code, and a first path switching control signal is generated; when the identifier bit sequence does not match the vector operation code identifier set, it is determined that the current instruction type is a non-vector operation code, and a second path switching control signal is generated; inputting the path switching control signal to a physical channel selector of the light guide switch matrix, wherein the first path switching control signal triggers physical connection of a first transmission path, and the second path switching control signal triggers physical connection of a second transmission path, completing lightweight deployment and inference execution of a large model.

[0104] In the above steps, the field programmable gate array is a programmable hardware component used to accelerate the computation of specific tasks, whose structure can be dynamically adjusted. The shared memory area is a common storage space accessible by both the processing unit and the field programmable gate array, facilitating data exchange. The instruction type distinguishes different operation commands such as addition or multiplication. The vector opcode type refers to a specific command type that processes batch data such as vector operations. The vector operation type set is a list that stores all supported vector opcodes. The optical switch matrix is a hardware device that can dynamically change the connection path, used to quickly switch data transmission channels. The topology connection logic defines the physical connection rules of the optical switch matrix, including specific connection points. The first data channel refers to the direct path connecting the processing unit vector register and the field programmable gate array high-speed interface, where the vector register is used to store batch data and the high-speed interface is used for high-speed data transmission. The second data channel refers to the backup path connecting the processing unit scalar register and the shared memory area, where the scalar register is used to store single data points. The hardware instruction execution cycle is the total process time from obtaining the instruction to completing the execution. The opcode type indicates whether the instruction is a vector or scalar operation. The instruction decoding stage is the stage where the processing unit interprets the meaning of the instruction. The opcode type monitoring point is set at the hardware location after the decoding stage, used to check the opcode type opcode field. The opcode field is the part of the binary sequence in the instruction that defines the operation. The identification bit sequence is a fixed bit sequence extracted from the opcode field, indicating the operation type. The predefined vector opcode identifier set is a list of known vector opcode values used for comparison matching. The path switching control signal is an electrical signal that triggers the optical switch matrix to change the path, the first path switching control signal enables vector operations, and the second path switching control signal enables scalar operations. The physical channel selector is the component in the optical switch matrix that actually switches the physical connection. The lightweight deployment of large models and inference execution refers to the process of efficiently running model prediction tasks on devices with limited resources.

[0105] In the embodiment of the present application, first, the reorganized parallel lookup table instruction block is loaded to the shared memory area of the processing unit and the field programmable gate array of the target edge device through step 1041, all vector operation code types are extracted from the instruction type of the acceleration capability profile, and a vector operation type set is generated. In a specific implementation, the reorganized instruction block is copied to the shared memory area as basic data using a memory read-write command; then, the acceleration capability profile is read, which is like a configuration file or table storing supported instruction type details, and all vector operation code types are filtered from it using a parsing algorithm, such as extracting these types by traversing the profile data to match the keyword vector; finally, the extracted vector operation code types are combined into a list as the vector operation type set. This process forms a complete flow: shared memory loading ensures data availability, and parsing the profile to extract instruction types generates a set based on actual hardware capabilities. For example, the acceleration capability profile may list supported operation codes including VADD vector addition, VMUL vector multiplication, and ADD normal addition; VADD and VMUL are extracted as vector types when parsing, and the set [VADD; VMUL] is generated for subsequent step configuration.

[0106] Secondly, based on the vector operation type set, the topology connection logic of the first data channel is configured in the optical switch matrix through step 1042, the channel connection processing unit vector register and the field programmable gate array high-speed interface are set, the topology connection logic of the second data channel is set, and the channel connection processing unit scalar register and the shared memory area are set. In a specific implementation, the physical connection mode is defined according to the logical rule of the optical switch matrix set by the vector operation type set, and the configuration command is used to write the logic; wherein the first data channel is configured as: the vector register data output of the processing unit is directly connected to the high-speed interface of the field programmable gate array to accelerate transmission; the second data channel is configured as: the scalar register of the processing unit is connected through the data channel of the shared memory area when normal operation is used. This process forms a complete flow: the set provides input guidance to configure the logical rule to ensure that different types of instructions have optimized path configurations. For example, the optical switch matrix is configured for the set [VADD; VMUL] to select the first channel to directly transmit a large amount of data when vector operation is performed, and to select the second channel to save resources through the shared memory area buffer when non-vector operation such as ADD is performed.

[0107] Finally, the opcode type in the hardware instruction execution cycle is monitored by step 1043, and when a vector opcode is detected, the light guide switch matrix is triggered to switch to the first data channel, and when a non-vector opcode is detected, the light guide switch matrix is triggered to switch to the second data channel to complete the lightweight deployment and inference execution of the large model. In the specific implementation, after the instruction decoding stage is completed, the hardware checkpoint is set, and the opcode field part of the current instruction is captured through the hardware logic circuit; then, the preset identification bit sequence in the opcode field is extracted, which specifies the operation type, such as obtaining the key bit position through the shift operation, for example, extracting the high 4 bits from the 8-bit opcode as the identification sequence; then, the identification bit sequence is compared with the predefined vector opcode identifier set one by one, such as using a logic gate circuit to match the same bit value; when the identification sequence matches the set, the first path switching control signal is generated, which indicates the vector operation, otherwise the second path switching control signal is generated; finally, these signals are input to the physical channel selector, for example, the electronic controller of the switch matrix triggers the actual physical connection to switch the first channel or the second channel, the path optimization is completed before instruction execution, the data transmission efficiency is guaranteed, and thus the lightweight model inference is executed. This process forms a complete flow: the monitoring point triggers the type checking sequence extraction and matching control signal generation and dynamically switches the path in the execution cycle. For example, assuming that the opcode field is 1100, the identification bit sequence is 11, and the predefined set is [11 vector] and [00 scalar]; comparison shows that 11 matches the set to generate the first signal, which triggers the light guide switch matrix to switch to the first channel for fast vector data processing, and realizes efficient model inference deployment to end the large model task.

[0108] In practical applications, when performing lightweight deployment of a plant disease recognition model in an agricultural Internet of Things edge computing device A, the reorganized parallel lookup table instruction block is first loaded into the 256 KB memory area shared by the processing unit and the field programmable gate array, with the starting address being 0x2000. The set of vector operation code types extracted from the acceleration capability profile includes VLOAD vector loading, VSHR logical shifting, VAND bit operation, and other 8 types of operation, and the topology connection logic of the light guide switch matrix is configured based on the set. The first data channel establishes a direct link between the processing unit vector register group and the field programmable gate array high-speed interface, and the transmission bandwidth is set to 128 bits per cycle. The second data channel establishes an access path between the processing unit scalar register and the shared memory area. In the design of the hardware instruction execution phase, the hardware logic circuit intercepts the operation code field after instruction decoding. When the operation code bit[5:3] is detected as 101 binary sequence, the feature matches the predefined VLOAD operation code identifier, and the first path switching control signal is generated immediately. In this case, the light guide switch matrix is switched to the first data channel within 4 nanoseconds, for example, when executing the VLOAD instruction, the calculation block of the field programmable gate array directly obtains the weight parameters in the vector register through the channel, realizing zero-copy data transmission. Conversely, when the operation code bit[5:3] is 010, the scalar operation type is matched, and the second path switching control signal is triggered. At this time, the processing unit accesses the intermediate calculation result in the shared memory through the second channel. This dynamic path switching mechanism realizes 27 channel switching operations in a single reasoning period, ensuring that large models maintain millisecond-level reasoning latency in resource-constrained agricultural edge devices.

[0109] In the overall scheme of the above step 104, the reorganized parallel lookup table instruction block is loaded into the shared memory area of the processing unit and the field programmable gate array, and based on the set of vector operation code types extracted from the acceleration capability profile, the first data channel topology logic connecting the vector register and the high-speed interface and the second data channel topology logic connecting the scalar register and the shared memory are respectively configured in the light guide switch matrix. Through the hardware logic circuit, the operation code field identification bit sequence is intercepted after instruction decoding, and it is matched with the predefined set of vector operation code identifiers in real time: when the vector operation code is detected, the first path switching control signal is generated to trigger the light guide switch matrix to establish the physical connection of the vector register to the high-speed interface, and when the non-vector operation code is detected, the second path switching control signal is generated to trigger the physical connection of the scalar register to the shared memory. Thereby, the data transmission path between the two architectures is switched immediately within the hardware instruction execution period, realizing adaptive data routing optimization based on operation code type, and finally completing the lightweight deployment and efficient reasoning execution of large models.

[0110] The following is a complete embodiment for steps 101 to 104:

[0111] AsFigure 2 As shown, in a certain smart agriculture edge computing device deployment, when implementing a large model lightweight solution in device A, first read the instruction set architecture identification code 0xAA01 from the device system configuration area, query the architecture table to confirm support for 8 instruction types such as vector multiplication addition and bit expansion, and 128-bit operation bit width. Through the operation management interface, 4 physical cores in the CPU and 6 reconfigurable computing blocks in the FPGA are collected to form an acceleration capability image containing 10 parallel units (4+6) and specific instruction characteristics. Based on the 128-bit maximum bit width in the image, the 2.5MB weight of the plant disease recognition model is cut into 65536 units with a fixed length of 32 bits. Every 4 units form a weight subset (128 / 32=4), a total of 16384 subsets are generated. In the unified lookup table engine, each subset is completely enumerated from 0 to 4294967295 input values: when the input value is 261, the engine synchronously calculates the output of this value in the 4 units - unit 1 outputs the convolution coefficient 0.74, unit 2 outputs the feature scaling factor 1.2, unit 3 outputs the bias compensation -0.15, and unit 4 outputs the normalization parameter 0.98, combined into a four-dimensional output sequence; finally, a 128-bit precomputed vector matrix is generated in ascending order of input value and stored in the continuous memory area 0x2000~0x3FFFF.

[0112] According to the number of 10 parallel units, the memory area is equally divided into 10 25.6KB segments (0x2000-0x7FFF, 0x8000-0xDFFF...). 10 instruction blocks are compiled: block 1 fills the "VLDR128 R1,0x2000" opcode, block 2 fills "VLDR128 R2,0x8000", etc. Assign 10 execution threads to the corresponding instruction blocks. Scanning found that block 3 (starting at 0x14000) outputs are dependent on block 7 (starting at 0x2A000), by inserting the SYNC instruction and moving block 7 to the front of block 3, the pipeline sequence is reconstructed to block 1→2→7→SYNC→3→4...

[0113] In the hardware execution phase, the instruction blocks are loaded into the CPU-FPGA shared memory starting at 0x20000. Based on the vector operation types VLDR / VSHR / VFMA in the image, configure the light guide matrix double-channel topology: channel 1 directly connects the CPU vector register and the FPGA high-speed interface (128-bit wide); channel 2 connects the CPU scalar register and the shared memory area. In the instruction decoding phase, when the opcode low 3 bits are 101 for vector operations, the matrix is switched to channel 1 within 0.8ns; when the low 3 bits are 010 for scalar operations, the matrix is switched to channel 2, achieving single-cycle dynamic path switching and completing efficient inference of the model on the agricultural edge device.

[0114] Figure 3A structural schematic diagram of a system for lightweight deployment of a large model on an edge computing device is provided for an embodiment of the present application, as shown in Figure 3 The system includes:

[0115] The construction module 31 is configured to extract a hardware instruction set architecture type and a number of parallel computing units of a target edge device, identify instruction-level parallelism capability based on the hardware instruction set architecture type, and construct an acceleration capability image containing an instruction type and the number of parallel computing units.

[0116] The matching module 32 is configured to group and divide large model weights into a plurality of weight subsets according to the instruction type in the acceleration capability image by using a unified lookup table vectorization engine, perform pre-computation operations on each weight subset, and generate pre-computed vectors matching a target instruction set architecture.

[0117] The generation module 33 is configured to compile the pre-computed vectors based on the number of parallel computing units and the instruction-level parallelism capability in the acceleration capability image, generate parallel lookup table instruction blocks adapted to the target instruction set architecture, assign a number of execution threads equal to the number of parallel computing units to the parallel lookup table instruction blocks, reorganize pipeline operation sequences according to the instruction-level parallelism capability, and eliminate data dependency conflicts.

[0118] The switching module 34 is configured to load the parallel lookup table instruction blocks into a shared memory area of a processing unit and a field programmable gate array of a target edge device, configure topology connection logic of a light guide switch matrix based on the instruction type in the acceleration capability image, switch data transmission paths between the processing unit and the field programmable gate array within a hardware instruction execution period through the light guide switch matrix, and complete lightweight deployment and inference execution of the large model.

[0119] Figure 3 The lightweight deployment system of the large model on the edge computing device can perform Figure 1 The implementation principle and technical effects of the lightweight deployment method of the large model on the edge computing device in the embodiment shown in are not described again. The specific manner in which each module, unit in the lightweight deployment system of the large model on the edge computing device in the above embodiment performs operations has been described in detail in the embodiment related to the method, and will not be described in detail here.

[0120] In one possible design, Figure 3 The lightweight deployment system of the large model on the edge computing device in the embodiment shown in can be implemented as a computing device, as shown in Figure 4 The computing device can include a storage component 41 and a processing component 42.

[0121] The storage component 41 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 42.

[0122] The processing component 42 is configured to perform the above Figure 1 The method for lightweight deployment of a large model on an edge computing device in the embodiment.

[0123] The processing component 42 can include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component can also be one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components, for executing the above method.

[0124] The storage component 41 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or their combination, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0125] Of course, the computing device can also include other components, such as input / output interfaces, display components, communication components, etc.

[0126] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.

[0127] The communication component is configured to facilitate wired or wireless communication between the computing device and other devices, etc.

[0128] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform, and the computing device can be a cloud server, and the processing component, the storage component, etc. can be basic server resources rented or purchased from the cloud computing platform.

[0129] The embodiment of the present application also provides a computer storage medium, which stores a computer program, and the computer program can implement the above Figure 1 The method for lightweight deployment of a large model on an edge computing device in the embodiment.

[0130] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0131] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0132] Through the description of the foregoing embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and a necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the foregoing technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0133] Finally, it should be noted that: the foregoing embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A lightweight deployment method for large models on edge computing devices, characterized in that, The method comprises the following steps: extracting the hardware instruction set architecture type and the number of parallel computing units of the target edge device, identifying the instruction level parallelism capability based on the hardware instruction set architecture type, and constructing an acceleration capability image containing the instruction type and the number of parallel computing units; According to the instruction type in the acceleration capability image, the large model weight is divided into multiple weight subsets by the uniform lookup table vectorization engine according to the bit grouping unit, the pre-computation operation is performed on each weight subset, and the pre-computation vector matching the target instruction set architecture is generated; Based on the number of parallel computing units and the instruction level parallelism capability in the acceleration capability image, the pre-computation vector is compiled to generate a parallel lookup table instruction block adapted to the target instruction set architecture, an execution thread equal to the number of parallel computing units is assigned to the parallel lookup table instruction block, the pipeline operation order is reorganized according to the instruction level parallelism capability, and the data dependency conflict is eliminated; The parallel lookup table instruction block is loaded into the shared memory area of the processing unit and the field programmable gate array of the target edge device, the topology connection logic of the light guide switch matrix is configured based on the instruction type in the acceleration capability image, the data transmission path between the processing unit and the field programmable gate array is switched within the hardware instruction execution period through the light guide switch matrix, and the lightweight deployment and inference execution of the large model are completed; According to the instruction type in the acceleration capability image, the large model weight is divided into multiple weight subsets by the uniform lookup table vectorization engine according to the bit grouping unit, the pre-computation operation is performed on each weight subset, and the pre-computation vector matching the target instruction set architecture is generated, which comprises: Extracting the maximum parallel operation bit width from the instruction type in the acceleration capability image, obtaining a pre-set fixed length of bit grouping unit, dividing the maximum parallel operation bit width by the fixed length of bit grouping unit to calculate the number of bit grouping units contained in a single weight subset; The large model weight sequence is cut into a continuous unit sequence according to the fixed length of the bit grouping unit, and the continuous units are intercepted from the continuous unit sequence according to the number of bit grouping units to divide them into multiple weight subsets; For each weight subset, enumerate the complete value range of the input value in the uniform lookup table vectorization engine, and calculate the set of simulated output values of the input value on all bit grouping units in the weight subset for each input value in the complete value range of the weight subset; The set of simulated output values is spliced into a vectorization result with a bit width equal to the maximum parallel operation bit width in order, and the vectorization result is arranged in order of input value from small to large to form a pre-computation vector in a continuous storage area.

2. The method of claim 1, wherein, For each weight subset, enumerate the complete value range of the input value in the uniform lookup table vectorization engine, and calculate the set of simulated output values of the input value on all bit grouping units in the weight subset for each input value in the complete value range of the weight subset, which comprises: Determine the complete value range of the input value of the weight subset, which is defined by the binary value space boundary corresponding to the fixed length of the bit grouping unit; Traverse each input value in the complete value range in ascending integer order, take the current input value as an input item, and input the uniform lookup table vectorization engine; Synchronously calculate the independent output values of the input item on all bit grouping unit cells in the weight subset of the complete value range, generating an output value sequence equal to the number of bit grouping units; Combine the output value sequence in the physical storage order of the bit grouping units to form a set of analog output values corresponding to the current input value.

3. The method of claim 1, wherein, Load the parallel lookup table instruction block into the shared memory area of the processing unit and the field programmable gate array of the target edge device, configure the topology connection logic of the light guide switch matrix based on the instruction type in the acceleration capability profile, switch the data transmission path between the processing unit and the field programmable gate array through the light guide switch matrix within the hardware instruction execution cycle, and complete the lightweight deployment and inference execution of the large model, including: Load the reorganized parallel lookup table instruction block into the shared memory area of the processing unit and the field programmable gate array of the target edge device, extract all vector operation code types from the instruction type of the acceleration capability profile to generate a set of vector operation types; Based on the set of vector operation types, configure the topology connection logic of the first data channel in the light guide switch matrix, which connects the processing unit vector register and the field programmable gate array high-speed interface, and set the topology connection logic of the second data channel, which connects the processing unit scalar register and the shared memory area; Monitor the opcode type in the hardware instruction execution cycle, trigger the light guide switch matrix to switch to the first data channel when a vector opcode is detected, and trigger the light guide switch matrix to switch to the second data channel when a non-vector opcode is detected, to complete the lightweight deployment and inference execution of the large model.

4. The method of claim 3, wherein, Monitor the opcode type in the hardware instruction execution cycle, trigger the light guide switch matrix to switch to the first data channel when a vector opcode is detected, and trigger the light guide switch matrix to switch to the second data channel when a non-vector opcode is detected, to complete the lightweight deployment and inference execution of the large model, including: Establish an opcode type monitoring point after the instruction decoding stage, intercept the opcode field of the current instruction through a hardware logic circuit, extract a preset identifier bit sequence of the opcode field, and match the identifier bit sequence with a predefined set of vector opcode identifiers; When the identifier bit sequence matches the set of vector opcode identifiers, determine that the current instruction type is a vector opcode, and generate a first path switching control signal; When the identifier bit sequence does not match the set of vector opcode identifiers, determine that the current instruction type is a non-vector opcode, and generate a second path switching control signal; Input the path switching control signal into the physical channel selector of the light guide switch matrix, wherein the first path switching control signal triggers the physical connection of the first transmission path, and the second path switching control signal triggers the physical connection of the second transmission path, to complete the lightweight deployment and inference execution of the large model.

5. The method of claim 1, wherein, Extracting a hardware instruction set architecture type and a parallel computing unit quantity of a target edge device, identifying instruction level parallelism capability based on the hardware instruction set architecture type, and constructing an acceleration capability image containing instruction types and parallel computing unit quantities, including: reading an identification code of the hardware instruction set architecture type from a system configuration storage area of the target edge device, querying a predefined architecture capability mapping table according to the identification code, and obtaining all operation codes and bit width parameters supported by the hardware instruction set architecture type; collecting a quantity value of independent operation cores in a processing unit through an operation unit management interface of the target edge device, and collecting a quantity value of configurable computing blocks through a configuration state interface of a field programmable gate array; based on the operation codes and the bit width parameters, identifying instruction level parallelism capability features, adding the quantity value of the independent operation cores and the quantity value of the configurable computing blocks to obtain a total parallel computing unit quantity; fusing the all operation codes and bit width parameters, the instruction level parallelism capability features, and the total parallel computing unit quantity to construct an acceleration capability image containing instruction types and parallel computing unit quantities.

6. The method of claim 1, wherein, Based on the parallel computing unit quantity and the instruction level parallelism capability in the acceleration capability image, the pre-computed vector is compiled to generate a parallel lookup table instruction block adapted to the target instruction set architecture, an execution thread equal to the parallel computing unit quantity is assigned to the parallel lookup table instruction block, the pipeline operation order is reorganized according to the instruction level parallelism capability, and data dependency conflicts are eliminated, including: According to the instruction level parallelism capability of the acceleration capability image, the continuous memory distribution characteristics of the pre-computed vector are analyzed, the memory address continuity is taken as the division basis, the segments of the pre-computed vector stored continuously are compiled into the same instruction block, and based on the parallel computing unit quantity, an equal number of independent instruction blocks are generated; filling the lookup table operation code of the target instruction set architecture and the corresponding pre-computed vector starting address offset in each of the independent instruction blocks to obtain a parallel lookup table instruction block adapted to the target instruction set architecture, creating an execution thread identifier equal to the parallel computing unit quantity value, and assigning each of the execution thread identifiers to the parallel lookup table instruction block; scanning the output-input relationship among all the parallel lookup table instruction blocks, identifying instruction block pairs with data dependency, adjusting the execution time sequence order of the identified instruction block pairs, inserting a synchronization control operation code to eliminate data dependency conflicts, and rearranging the instruction block sequence according to the adjusted time sequence order.

7. A lightweight deployment system for large models on edge computing devices, characterized in that, including: a construction module for extracting a hardware instruction set architecture type and a parallel computing unit quantity of a target edge device, identifying instruction level parallelism capability based on the hardware instruction set architecture type, and constructing an acceleration capability image containing instruction types and parallel computing unit quantities; a matching module for grouping the model weights into multiple weight subsets according to the instruction types in the acceleration capability image by a uniform lookup table vectorization engine, performing pre-computation on each of the weight subsets, and generating a pre-computed vector matched to the target instruction set architecture. The generating module is configured to compile the pre-computed vector based on the number of parallel computing units and the instruction-level parallelism capability in the acceleration capability profile, generate a parallel lookup table instruction block adapted to a target instruction set architecture, assign an execution thread equal to the number of parallel computing units to the parallel lookup table instruction block, reorganize the pipeline operation sequence according to the instruction-level parallelism capability, and eliminate data dependency conflicts. The switching module is configured to load the parallel lookup table instruction block to a shared memory region of a processing unit and a field programmable gate array of a target edge device, configure topology connection logic of a light guide switch matrix based on the instruction type in the acceleration capability profile, switch a data transmission path between the processing unit and the field programmable gate array through the light guide switch matrix within a hardware instruction execution period, and complete lightweight deployment and inference execution of a large model. According to the instruction type in the acceleration capability profile, the large model weight is divided into a plurality of weight subsets by a bit grouping unit, a pre-computed vector matching the target instruction set architecture is generated by performing a pre-computation operation on each weight subset, and the pre-computed vector includes: A maximum parallel operation bit width is extracted from the instruction type in the acceleration capability profile, a preset bit grouping unit fixed length is obtained, the number of bit grouping units contained in a single weight subset is calculated by dividing the maximum parallel operation bit width by the bit grouping unit fixed length; The large model weight sequence is cut into a continuous unit sequence according to the bit grouping unit fixed length, and the continuous unit sequence is divided into a plurality of weight subsets by intercepting an equal number of continuous units according to the number of bit grouping units; For each weight subset, the complete value range of the input value is enumerated in the uniform lookup table vectorization engine, and for each input value of the weight subset in the complete value range, a set of simulated output values of the input value on all bit grouping units in the weight subset is calculated; The set of simulated output values is spliced into a vectorization result with a bit width equal to the maximum parallel operation bit width in sequence, the vectorization result is arranged in ascending order of input values, and a pre-computed vector is formed in a continuous storage region.

8. A computing device, comprising: The device comprises a processing component and a storage component, the storage component stores one or more computer instructions, the one or more computer instructions are called and executed by the processing component, and a lightweight deployment method of a large model on an edge computing device is implemented.

9. A computer storage medium, characterized in that The computer program is stored in the computer and is executed by the computer, and a lightweight deployment method of a large model on an edge computing device is implemented.

Citation Information

Patent Citations

  • Tensor processing method and device, electronic equipment and storage medium

    CN117371537A

  • In-memory computing system, method and device with dynamically adjustable precision

    CN120067043A