Feature data processing method based on edge device and related device
By replacing the exponential formula of the probability normalization function in the neural network model of the edge device with the Taylor expansion formula and dynamically selecting the order according to the size of the eigenvalue, the problems of computing resources and energy consumption of the edge device are solved, and efficient feature data processing is achieved.
Patent Information
- Application Number
- CN202511286011.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-10
AI Technical Summary
The calculation of the probability normalization function of edge devices consumes a lot of hardware resources and energy, becoming a bottleneck for feature data processing.
In the neural network model of the edge device, the exponential formula in the probability normalization function is replaced by the Taylor expansion formula, and the order of the Taylor expansion formula is dynamically selected according to the size of the eigenvalue, and the parallel computing module is used for calculation.
It significantly improves the feature data processing efficiency of edge devices, reduces computing resource requirements and energy consumption, and improves computing performance and real-time performance.
Smart Images

Figure CN120763718A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of edge computing technology, and in particular relates to a feature data processing method based on edge devices and related equipment. Background Art
[0002] With the rapid development of artificial intelligence and the Internet of Things (IoT), an increasing number of edge devices, such as smart cameras, smart sensors, and wearables, are capable of real-time data collection and preliminary intelligent processing. To improve service efficiency and ensure data security, there is a growing demand for local data processing and immediate decision-making. Neural networks, particularly convolutional neural networks (CNNs), demonstrate exceptional performance in tasks such as feature extraction, data dimensionality reduction, and classification and recognition. Consequently, they are increasingly being deployed on various edge devices, enabling on-device intelligence.
[0003] In neural networks, probabilistic normalization functions (such as the Softmax function) are widely used to normalize feature vectors and output probability distributions. However, the exponential operations in these functions consume significant computing hardware resources. Limited by factors such as the computing power, storage space, and energy consumption of edge devices, these functions become a major bottleneck for efficient feature processing on edge devices. Consequently, improving the efficiency of feature data processing on edge devices has become a pressing technical challenge. Summary of the Invention
[0004] The embodiments of the present application provide a method, apparatus, computer program product, computer-readable storage medium, and electronic device for processing feature data based on an edge device, thereby improving the processing efficiency of feature data by the edge device to a certain extent.
[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.
[0006] According to a first aspect of an embodiment of the present application, a feature data processing method based on an edge device is provided, wherein a neural network model is deployed in a processor of the edge device, wherein the neural network model includes a convolutional network and a probability normalization function, and the method includes: obtaining feature data to be processed; performing a convolution operation on the feature data based on the convolutional network to obtain a feature vector of the feature data; and performing a normalization operation on the eigenvalues in the feature vector based on the probability normalization function, while replacing the exponential formula in the probability normalization function with a Taylor expansion formula, to obtain a probability distribution corresponding to the feature data, wherein the probability distribution is used to classify the feature data.
[0007] In some embodiments of the present application, based on the aforementioned scheme, before performing normalization operation on the eigenvalues in the eigenvector based on the probability normalization function, the method also includes: determining a target order Taylor expansion formula for the probability normalization function for processing each eigenvalue in the eigenvector, the target order Taylor expansion formula is used to replace the exponential formula in the probability normalization function, and the target order is adapted to the size of each eigenvalue.
[0008] In some embodiments of the present application, based on the aforementioned solution, the target order is positively correlated with the magnitude of each eigenvalue.
[0009] In some embodiments of the present application, based on the aforementioned scheme, the method for determining a target order Taylor expansion formula for the probability normalization function for processing each eigenvalue in the eigenvector includes: determining a target operation extension instruction from a plurality of predefined operation extension instructions according to each eigenvalue in the eigenvector; and determining a corresponding target order Taylor expansion formula for the probability normalization function for processing each eigenvalue according to the target operation extension instruction.
[0010] In some embodiments of the present application, based on the aforementioned scheme, the target operation extension instruction is determined from multiple predefined operation extension instructions according to each eigenvalue in the eigenvector, including: if each eigenvalue falls into a first eigenvalue interval, the first operation extension instruction among the multiple predefined operation extension instructions is determined as the target operation extension instruction, and the target order corresponding to the first operation extension instruction is N order; if each eigenvalue falls into a second eigenvalue interval, the second operation extension instruction among the multiple predefined operation extension instructions is determined as the target operation extension instruction, and the target order corresponding to the second operation extension instruction is M order, wherein M is greater than N, and the minimum value of the second eigenvalue interval is greater than the maximum value of the first eigenvalue interval.
[0011] In some embodiments of the present application, based on the aforementioned solution, the N order is 6 orders, and the M order is 9 orders.
[0012] In some embodiments of the present application, based on the aforementioned solution, the target operation extension instruction is used to directly call a probability normalization function to perform a normalization operation on each of the eigenvalues.
[0013] In some embodiments of the present application, based on the foregoing scheme, the processor comprises a first calculation module and a second calculation module, wherein the first calculation module is configured to calculate the first order formula to the Nth order formula in the Taylor expansion formula, and the second calculation module is configured to calculate the (N+1)th order formula to the Mth order formula in the Taylor expansion formula; in the process of normalizing the feature values in the feature vector based on the probability normalization function, the method further comprises: if the target order is N, calling the first calculation module to calculate the Nth order Taylor expansion formula; if the target order is M, calling the first calculation module and the second calculation module to calculate the Mth order Taylor expansion formula.
[0014] In some embodiments of the present application, based on the foregoing scheme, the first calculation module comprises N calculation units, and the calling of the first calculation module to calculate the Nth order Taylor expansion formula comprises: calling each calculation unit in the first calculation module to calculate each order formula in the Nth order Taylor expansion formula in parallel.
[0015] In some embodiments of the present application, based on the foregoing scheme, the second calculation module comprises M-N calculation units, and the calling of the first calculation module and the second calculation module to calculate the Mth order Taylor expansion formula comprises: calling each calculation unit in the first calculation module and each calculation unit in the second calculation module to calculate each order formula in the Mth order Taylor expansion formula in parallel.
[0016] In some embodiments of the present application, based on the foregoing scheme, the normalization operation on the feature values in the feature vector based on the probability normalization function comprises: based on the probability normalization function, performing normalization operation on each feature value in the feature vector in parallel.
[0017] According to a second aspect of the embodiments of the present application, the processor of the edge device is deployed with a neural network model, the neural network model comprises a convolution network and a probability normalization function, and the device comprises: an acquisition unit configured to acquire feature data to be processed; a first operation unit configured to perform convolution operation on the feature data based on the convolution network to obtain a feature vector of the feature data; and a second operation unit configured to perform normalization operation on feature values in the feature vector based on the probability normalization function in the case of replacing an exponential formula in the probability normalization function with a Taylor expansion formula, to obtain a probability distribution corresponding to the feature data, wherein the probability distribution is used for classifying the feature data.
[0018] According to a third aspect of an embodiment of the present application, a computer program product is provided, which includes computer instructions, which are stored in a computer-readable storage medium and are suitable for being read and executed by a processor, so that a computer device having the processor executes to implement the operations performed by the method described in the first aspect above.
[0019] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which at least one computer program instruction is stored. The at least one computer program instruction is loaded and executed by a processor to implement the operations performed by the method described in the first aspect above.
[0020] According to a fifth aspect of an embodiment of the present application, an electronic device is provided, comprising one or more processors and one or more memories, wherein at least one computer program instruction is stored in the one or more memories, and the at least one computer program instruction is loaded and executed by the one or more processors to implement the operations performed by the method described in the first aspect above.
[0021] Based on the technical solution proposed in this application, replacing the exponential formula of the probability normalization function in the neural network model with the Taylor expansion formula can significantly improve the efficiency of edge devices in processing feature data. Specifically, the operation of the exponential formula has a large overhead in hardware, especially on edge devices without floating-point hardware or those that are sensitive to energy consumption, while the polynomial form of the Taylor expansion is more suitable for hardware implementation. In this way, complex exponential operations are avoided, the demand for computing resources and energy consumption can be greatly reduced, and the processing efficiency of feature data by edge devices can be improved.
[0022] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, explaining the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings: Figure 1 A flow chart of a method for processing feature data based on edge devices in an embodiment of the present application is shown; Figure 2 A block diagram of a feature data processing device based on an edge device in an embodiment of the present application is shown; Figure 3A schematic structural diagram of an electronic device in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0025] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0026] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices. It should also be noted that in the accompanying drawings, certain components that do not affect the explanation of the technical solutions of this application have been omitted for clarity.
[0027] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0028] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more.
[0029] In order to enable those skilled in the art to better understand this application, the technical concepts and application background involved in this application are first briefly explained.
[0030] Edge device: refers to a computing device or terminal deployed at the edge of the network, close to the data source or user end, and capable of performing data collection, preprocessing, analysis or preliminary decision-making locally. The edge device has limited or constrained computing resources, storage capacity and energy consumption budget, and is usually used to reduce network transmission delay, reduce bandwidth consumption, protect data privacy or meet real-time requirements. The edge devices described in this application may include but are not limited to the following device types: smart cameras, access control devices, smart home controllers, industrial sensor gateways, vehicle-mounted terminals (automotive electronic units), drone flight control and visual processing modules, wearable devices (such as smart watches, fitness trackers), mobile robots, local understanding modules of smart speakers, remote monitoring cameras, portable medical devices, and miniaturized edge computing nodes deployed on edge gateways or edge servers.
[0031] Neural network model: refers to a parameterized function model composed of several computing units (i.e., neurons or nodes) interconnected in a specific topological structure, which is used to extract, transform and map the input feature data, thereby achieving classification, regression, detection, segmentation, generation or other artificial intelligence tasks. The neural network model includes but is not limited to convolutional neural network (CNN), recurrent neural network (RNN), long short-term memory network (LSTM), gated recurrent unit (GRU), transformer, graph neural network (GNN), and combinations of the above structures or their derivative variants. In the present invention, the neural network model at least includes a convolutional network structure and a probabilistic normalization function as its partial components, which is used to extract feature vectors and map the feature vectors to probabilistic outputs or normalized feature representations.
[0032] Probability normalization function: A mathematical mapping function that converts the real-number output (usually called logits or eigenvalues) of a neural network model (or other numerical processing module) into a normalized vector with probabilistic meaning. This function transforms each element of the input vector so that each component of the output vector is non-negative and the sum of all components is 1, thus being interpretable as a discrete probability distribution. Common probability normalization functions include the Softmax function, the normalized exponential function, and variants that are normalized after the logarithmic probability transformation.
[0033] Exponential function: refers to a power function with the natural constant e (approximately equal to 2.718281828…) as its base. Its mapping to the real input z is defined as exp(z) = e z. Commonly used as a core computing unit in probability normalization functions (such as the Softmax function).
[0034] Taylor expansion (Taylor series) formula: refers to a mathematical expression that expresses a real or complex function analyzed near a point as a power series about that point, that is, an approximate representation of the real or complex function, which can have different representations depending on the order.
[0035] In neural networks, probability normalization functions are widely used to normalize feature vectors and output probability distributions. However, the exponential operations in probability normalization functions consume a lot of computing hardware resources and are limited by factors such as the computing power, storage space, and energy consumption of edge devices. This makes probability normalization functions a major bottleneck for efficient feature processing on edge devices. In this context, this application proposes a feature data processing method based on edge devices to improve the efficiency of feature data processing on edge devices.
[0036] The following describes the implementation details of the technical solution of the embodiment of the present application: Reference Figure 1 , shows a flowchart of a feature data processing method based on an edge device in an embodiment of the present application. The feature data processing method based on an edge device can be executed by a device with computing and processing functions, wherein a neural network model is deployed in the processor of the edge device, and the neural network model includes a convolutional network and a probability normalization function.
[0037] In the present application, the processor may include any one of a Field Programmable Gate Array (FPGA), a Graphics Processing Unit (GPU), a Google Tensor Processing Unit (TPU), and an Application-Specific Integrated Circuit (ASIC).
[0038] Reference Figure 1 As shown, the feature data processing method based on edge devices includes at least steps 110 to 130, which are described in detail as follows: In step 110 , feature data to be processed is obtained.
[0039] In the present application, the feature data may be image feature data, voice feature data, or sensor data collected by industrial sensors. Specifically, the present application does not impose any further restrictions on this.
[0040] In this application, acquiring feature data to be processed can involve edge devices collecting raw data in real time through their internal perception modules (e.g., cameras, sensors, microphones, etc.). For example, smart cameras can capture video frames, smart sensors can collect environmental parameters such as temperature, humidity, and pressure, and voice devices can acquire voice sampling data. The acquired data can be single-frame images, time series signals, or raw feature inputs of other dimensions, serving as input data for subsequent neural network model processing.
[0041] Reference Figure 1 In step 120, a convolution operation is performed on the feature data based on the convolutional network to obtain a feature vector of the feature data.
[0042] In this application, the processor of the edge device can perform convolution operations on the acquired feature data to be processed through a convolutional network deployed in a local neural network model.
[0043] Specifically, the edge device first receives or collects raw data (such as images, voice, or sensor signals), and then inputs the data into the convolutional network. The convolutional network extracts local spatial features from the input data through multiple convolutional layers, and converts the useful information in the data into a feature vector through a sliding window operation of the convolution kernel. This feature vector highly concentrates the main features of the original data, laying the foundation for subsequent probability normalization and classification operations. For example, in the edge camera scenario, the collected image will first be subjected to feature extraction through the convolutional layer of the convolutional neural network to generate a feature vector that can reflect the image content for subsequent intelligent recognition and decision-making. By completing this process directly on the edge device, the real-time and security of data processing can be improved, and dependence on cloud resources can be reduced.
[0044] Reference Figure 1 In step 130, when the exponential formula in the probability normalization function is replaced by the Taylor expansion formula, the eigenvalues in the eigenvector are normalized based on the probability normalization function to obtain the probability distribution corresponding to the feature data, and the probability distribution is used to classify the feature data.
[0045] In this application, for example, the probability normalization function is the Softmax function, that is, formula (1): (1) Among them, x j represents the jth eigenvalue in the eigenvector; s represents the number of eigenvalues in the eigenvector, that is, the dimension of the eigenvector.
[0046] Furthermore, when the exponential formula in the above Softmax function is replaced by the Taylor expansion formula, the Softmax function is the following formula (2): (2) Where n represents the order of Taylor expansion.
[0047] In this application, the probability distribution obtained by normalizing the eigenvalues based on the probability normalization function can be used to classify the feature data. For example, taking the example of traffic sign recognition by a smart camera, the image data collected by the smart camera is extracted by a neural network to obtain a 5-dimensional feature vector. After normalizing the eigenvalues in the feature vector, the probability distribution is: P=[0.23,0.15,0.08,0.31,0.23].
[0048] Since 0.31 is the largest, it can be determined that the image captured by the smart camera belongs to the fourth type of traffic sign.
[0049] Based on the technical solution proposed in this application, the exponential formula of the probability normalization function in the neural network model is replaced by the Taylor expansion formula, which can significantly improve the efficiency of edge devices in processing feature data. Specifically, the operation of the exponential formula is relatively expensive in hardware, especially on edge devices without floating-point hardware or those that are sensitive to energy consumption. The Taylor expansion formula consists of polynomial addition and multiplication, and the polynomial form is more suitable for hardware implementation. In this way, complex exponential operations are avoided, the computing resource requirements and energy consumption can be greatly reduced, and the processing efficiency of feature data by edge devices can be improved.
[0050] In the present application, before performing a normalization operation on the eigenvalues in the eigenvector based on the probability normalization function, the following step 131 may be further performed: Step 131, determining a target order Taylor expansion formula for the probability normalization function for processing each eigenvalue in the eigenvector, wherein the target order Taylor expansion formula is used to replace the exponential formula in the probability normalization function, and the target order is adapted to the size of each eigenvalue.
[0051] In this application, before normalizing each eigenvalue, the numerical range of each eigenvalue can be detected, and then a target-order Taylor expansion formula can be determined for the probability normalization function that processes each eigenvalue. For example, if the eigenvectors are [x1, x2, x3, x4], the target-order Taylor expansion formulas can be determined for the probability normalization functions that process x1, x2, x3, and x4, respectively.
[0052] In some embodiments of the present application, the target order is positively correlated with the magnitude of each eigenvalue.
[0053] In this application, in the probability normalization link of the neural network model, replacing the exponential formula in the probability normalization function with the Taylor expansion formula is an efficient approximation method. The order of the Taylor expansion directly determines the accuracy of the approximate calculation of the exponential formula. The larger the value of the eigenvalue, the greater the error between the low-order Taylor expansion formula and the exponential formula. Therefore, a higher order is required to ensure the accuracy of the approximation; conversely, for smaller eigenvalues, the low-order Taylor expansion is sufficient to meet the accuracy requirements. Therefore, by using the positive correlation between the target order and the size of each eigenvalue as a constraint, it is possible to optimize the allocation of computing resources while ensuring the overall calculation accuracy and improve the energy efficiency ratio on the edge device.
[0054] Specifically, we can first analyze the absolute value of each eigenvalue x, and then correspond the eigenvalue x to a predetermined interval based on the pre-designed partitioning rules. Each interval is associated with an optimal target order n. The larger the eigenvalue, the higher the assigned approximate order n, achieving a positive correlation between the target order n and the eigenvalue size.
[0055] For example, for the eigenvector x=[0.05,0.9,2.1,3.2], 0.05 can use the 4th-order Taylor expansion formula to replace the exponential formula in the probability normalization function, 0.9 can use the 6th-order Taylor expansion formula to replace the exponential formula in the probability normalization function, 2.1 can use the 9th-order Taylor expansion formula to replace the exponential formula in the probability normalization function, and 3.2 can use the 12th-order Taylor expansion formula to replace the exponential formula in the probability normalization function.
[0056] In this application, a strategy is adopted in which the target order is positively correlated with the size of each eigenvalue. The order of the Taylor expansion can be flexibly and dynamically allocated according to the actual value of the eigenvalue. For larger eigenvalues, a higher-order polynomial is used to improve the approximation accuracy, and for smaller eigenvalues, a lower-order polynomial is used to save computing resources. In this way, not only can the accuracy and stability of the probability normalization operation be effectively guaranteed, but the resource utilization and energy consumption performance of the edge device can also be significantly optimized, and the overall efficiency and real-time performance of feature data processing can be improved. It is particularly suitable for edge computing scenarios with limited computing power, sensitivity to latency and energy consumption, and the need to balance computing speed and resource constraints.
[0057] In a specific embodiment of the present application, the Taylor expansion formula for determining the target order of the probability normalization function for processing each eigenvalue in the eigenvector can be performed according to the following steps 1311 to 1312: Step 1311 : Determine a target operation extension instruction from a plurality of predefined operation extension instructions according to each eigenvalue in the eigenvector.
[0058] Step 1312: Determine a Taylor expansion formula of a target order for the probability normalization function for processing each eigenvalue according to the target operation extension instruction.
[0059] In this application, multiple predefined operation extension instructions can be designed, each instruction corresponding to a Taylor expansion formula of different orders. These extension instructions can be software functions, hardware microinstructions, or precompiled units in the operator library.
[0060] In this application, the arithmetic extension instructions may be extension instructions designed based on the RISC-V architecture. The RISC-V architecture is an open, modular, reduced instruction set computing (RISC) architecture specification. RISC-V defines a set of base instructions, optional extensions, and an ABI (Application Binary Interface) specification for software / hardware interoperability. These instructions can be implemented in processor cores, microcontrollers, systems-on-chip (SoCs), and related tool chains.
[0061] For example, for Taylor expansions of different orders, extended instructions can be designed separately, allowing the processor to perform approximate calculations on exp(x) of the corresponding order directly with a single instruction. These extended instructions can be defined in the custom instruction space of the RISCV architecture and implemented collaboratively by software and hardware. They support both software calls and efficient execution at the hardware level, greatly improving instruction dispatch and pipeline efficiency. By dynamically selecting different RISCV extended instructions, the processor can flexibly adapt to the computational requirements of each eigenvalue, significantly reducing the overall computational load and energy consumption while ensuring approximate accuracy.
[0062] In this embodiment, the step of determining the target operation extension instruction from a plurality of predefined operation extension instructions according to each eigenvalue in the eigenvector may be performed according to the following steps 13111 to 13112: Step 13111: If each of the eigenvalues falls within the first eigenvalue interval, a first operation extension instruction among a plurality of predefined operation extension instructions is determined as the target operation extension instruction, and the target order corresponding to the first operation extension instruction is N order.
[0063] Step 13112: If each of the eigenvalues falls into the second eigenvalue interval, the second operation extension instruction among the multiple pre-defined operation extension instructions is determined as the target operation extension instruction, and the target order corresponding to the second operation extension instruction is M order, where M is greater than N, and the minimum value of the second eigenvalue interval is greater than the maximum value of the first eigenvalue interval.
[0064] In the present application, the actual distribution characteristics of the eigenvalues can be divided into two or more non-overlapping eigenvalue intervals. For example, the first eigenvalue interval is [a, b]; the second eigenvalue interval is (b, c], where a, b, and c are real numbers, and the minimum value of the second interval is greater than the maximum value of the first interval, ensuring no overlap.
[0065] For each eigenvalue interval, a pre-defined hardware / software level operation expansion instruction can be defined. For example, Table 1.
[0066] Table 1: Pre-defined operation expansion instructions In the present application, as shown in Table 1 above, the first operation expansion instruction (such as RISCV custom instruction or software specific function, etc.) corresponds to the N-order Taylor expansion formula with lower precision requirement. The second operation expansion instruction corresponds to the M-order Taylor expansion formula (M>N) with higher precision to compensate for the approximation error caused by large eigenvalue interval. For example, the eigenvalue x=3.1 falls into the first eigenvalue interval (0, 5], and the first operation expansion instruction (such as VSOFTMAX) is called; the eigenvalue x=6.6 falls into the second eigenvalue interval (5, 10], and the second operation expansion instruction (such as VSOFTMAX6) is called.
[0067] In the present application, by dynamically mapping the eigenvalue interval and the order of Taylor expansion formula, combined with multiple pre-defined operation expansion instructions, each eigenvalue can adopt the most suitable approximation calculation path, which can significantly improve the adaptability and calculation efficiency of the probability normalization function. This scheme can intelligently select Taylor expansion formula with lower or higher order according to the actual amplitude of the eigenvalue, optimize resource utilization and improve energy efficiency, while considering numerical accuracy to avoid approximation error caused by large eigenvalue. At the same time, this technical scheme has good software and hardware adaptability, can be directly mapped to the expansion instruction in RISCV architecture or realized through software scheduling, is suitable for various computing platforms, and effectively improves the overall performance and practical value of the probability normalization operation in the power limited scene of edge devices, chip end, etc.
[0068] In the present application, the target operation expansion instruction can be used to directly call the probability normalization function to perform normalization operation on each eigenvalue.
[0069] The target operation extension instructions proposed in this application are not just simple arithmetic operation instructions, but dedicated hardware units that can be directly mapped and integrated into processors (such as the RISC-V architecture). These instructions are designed as highly complex single instructions that can complete all the core operation steps required for the probability normalization function in one go, such as addition, subtraction, multiplication, and division, thereby achieving efficient normalization processing for each input eigenvalue.
[0070] Traditional normalization operations typically rely on the step-by-step calling of high-level software library functions. For each eigenvalue, the software must sequentially call multiple arithmetic functions (such as calculating exp(x), accumulating the sum, and then performing normalized division). This involves multiple parameter passing, instruction dispatching, and memory access operations, which can easily lead to system call delays, pipeline blockages, and reduced computational efficiency. Frequent function jumps can become a performance bottleneck, especially in edge devices or embedded systems with limited computing power.
[0071] In contrast, this application integrates the core processes of probabilistic normalization (such as efficient approximation of exp(x), multi-step accumulation, and gradual normalization) into a single "full-featured hardware instruction." When executing this instruction, the processor no longer relies on software library functions, but instead directly calls the underlying hardware units to perform all necessary operations, including addition, subtraction, multiplication, and division. This significantly reduces the overhead associated with instruction calls and context switches, enabling fast normalization calculations in a single or even a few cycles. By directly mapping the target operation extension instructions to the processor hardware units, the latency of the normalization calculation is significantly reduced, avoiding the performance bottlenecks associated with traditional library function calls. This enables extremely fast data processing and throughput, and improves the efficiency of feature data processing on edge devices. Furthermore, this solution simplifies the software stack and code maintenance, freeing developers from having to focus on the details of the underlying normalization implementation, thereby enhancing application simplicity and robustness. Leveraging hardware instruction-level parallelism and pipeline scheduling, it also maximizes the resource utilization of the computing units, improving energy efficiency and computational density. In addition, this instruction has good platform adaptability and scalability, and can be flexibly integrated into various embedded devices, AI chips and edge computing platforms to meet the needs of efficient normalized computing in diversified scenarios.
[0072] In the present application, the processor may include a first computing module and a second computing module, wherein the first computing module is used to perform operations on the 1st to Nth order equations in the Taylor expansion equation, and the second computing module is used to perform operations on the N+1th to Mth order equations in the Taylor expansion equation.
[0073] Furthermore, in the present application, in the process of performing normalization operation on the eigenvalues in the eigenvector based on the probability normalization function, the following steps 132 to 133 may be performed: Step 132: If the target order is N, call the first calculation module to calculate the N-order Taylor expansion formula.
[0074] Step 133: If the target order is M, call the first calculation module and the second calculation module to calculate the M-order Taylor expansion formula.
[0075] In this application, to improve the adaptability and computational efficiency of the probability normalization function, a first computation module and a second computation module specifically designed for Taylor expansion calculations can be introduced within the processor. This architectural design allows for dynamic selection of Taylor expansions of different orders based on the magnitude or accuracy requirements of the actual eigenvalues, enabling the rational scheduling of computational resources and performance optimization.
[0076] Specifically, the first calculation module is used to efficiently process Taylor expansion calculations from the 1st to the Nth order, which generally covers the approximate accuracy requirements of most common input data during normalization. For the case where the target order is N, only the first calculation module needs to be called to complete the entire Taylor expansion operation, ensuring calculation accuracy while minimizing hardware resources and energy consumption.
[0077] When the normalization operation requires higher precision and requires expansion to an M-order (M>N) Taylor expansion, both the first and second computation modules can be called simultaneously. The first computation module handles parallel or pipelined processing of expressions from order 1 to order N, while the second computation module specializes in supplementary computation of expressions from order N+1 to order M. This division of labor not only highly parallelizes the overall computation process, significantly shortening the overall time required for high-order expansions, but also allows for on-demand activation and dynamic allocation of hardware resources, avoiding wasted high-order computing power in low-precision scenarios.
[0078] In addition, the above solution also implements efficient hardware reuse design to further improve resource utilization and reduce hardware costs. Specifically, for the Taylor expansion calculation unit, the calculation of the 1st to Nth order formulas is uniformly implemented using shared hardware units, and there is no need to design independent calculation paths for each order. When higher-precision normalization operations are required, it is only necessary to add a set of dedicated calculation paths for the N+1th to Mth order formulas based on the original shared units. In this way, the basic computing hardware can be reused, and only limited hardware resources are invested in the high-order part. Compared with the method of expanding the calculation unit separately for each order, it can save about LUT hardware resources, effectively reduce the overall chip area and power consumption, and help improve the scalability and cost performance of the processor.
[0079] In the present application, the first computing module may include N computing units, and the second computing module may include MN computing units.
[0080] In this application, in order to further optimize the hardware implementation efficiency of the Taylor expansion formula, the first computing module and the second computing module in the processor both adopt a highly parallel structural design, thereby significantly improving the throughput and response speed of the normalization operation.
[0081] Furthermore, the calling of the first calculation module to calculate the N-order Taylor expansion formula may be performed according to the following step 1321: Step 1321: Call each computing unit in the first computing module to perform operations on each order of the N-order Taylor expansion equation in parallel.
[0082] Specifically, the first computation module consists of N computational units, each corresponding to a term in the Taylor expansion. When normalization is required for an N-order Taylor expansion, each term is assigned to the corresponding N computational units, allowing all units to execute their respective computations in parallel within the same clock cycle.
[0083] For example, for the 6th order Taylor expansion formula "1+x+x² / 2!+x³ / 3!+x 4 / 4!+x 5 / 5!+x 6 / 6!", the six calculation units in the first calculation module can be called to calculate the formulas "x", "x² / 2!", "x³ / 3!", "x 4 / 4!","x 5 / 5!、"x 6 / 6!" Parallel computing. This parallel processing mechanism effectively avoids the latency bottleneck caused by serial computing and greatly improves the overall computing efficiency of the N-order expansion.
[0084] Furthermore, the calling of the first calculation module and the second calculation module to calculate the M-order Taylor expansion formula may be performed according to the following step 1331: Step 1331: Call each calculation unit in the first calculation module and each calculation unit in the second calculation module to perform calculations on each order of the M-order Taylor expansion formula in parallel.
[0085] When the target computational order is further increased to M (M>N), in addition to the first computational module, a second computational module is also activated. The second computational module consists of MN computational units and is specifically responsible for processing the N+1 to M-order Taylor expansion equations. Similarly, all computational units can operate in parallel, collaborating with the first computational module to complete all order calculations of the M-order Taylor expansion equation. That is, the N computational units in the first computational module and the MN computational units in the second computational module are called simultaneously, each calculating the order terms they are responsible for in parallel, ultimately completing the normalization operation of the M-order Taylor expansion equation quickly.
[0086] Through the aforementioned parallel hardware architecture design, this solution can flexibly schedule computing resources based on the required Taylor expansion order, effectively balancing computing performance and hardware utilization. When the order of calculation is low, simply enabling the first computing module is sufficient; in high-precision scenarios, the coordinated parallel operation of the two-level modules ensures both high efficiency and ease of integration and expansion. Furthermore, the independent calculation of each order of the formula facilitates subsequent hardware optimization, fault isolation, and dynamic management, further improving the reliability and maintainability of feature data processing.
[0087] In the present application, the normalization operation on the eigenvalues in the eigenvector based on the probability normalization function may also be performed according to the following step 134: Step 134 : performing normalization operations on the eigenvalues in the eigenvector in parallel based on the probability normalization function.
[0088] In the present application, further, in order to improve the overall normalization operation efficiency, a parallel processing mechanism can be used for the normalization operation of the feature vector.
[0089] Specifically, when probabilistic normalization is required for each eigenvalue in a eigenvector, not only is highly parallel computation achieved between the various orders of the Taylor expansion, but a parallel computation mechanism can also be introduced between the various components of the eigenvector. That is, for each eigenvalue in the input eigenvector, normalization operations can be performed in parallel based on a probabilistic normalization function (such as Softmax). In this way, each component of the eigenvector can be assigned to a separate processing unit or computational path, with each processing unit independently completing its corresponding normalization computation task, thereby achieving synchronized normalized output for the entire vector.
[0090] This dual parallelization design, while simultaneously executing Taylor expansion operations at all orders through a large number of parallel computing units, also enables multi-path concurrent processing at the eigenvector component level. This significantly improves the throughput of the overall normalization task and effectively reduces latency in batch data processing scenarios. Furthermore, this design is well-suited for hardware acceleration. For example, when deployed on FPGAs, ASICs, or dedicated AI chips, it can fully utilize multiple data paths to achieve large-scale, low-latency parallel normalization processing.
[0091] In addition, this solution makes the normalization module extremely scalable. Regardless of how the dimension of the feature vector is expanded, the normalization hardware can horizontally expand the number of processing units on demand, flexibly adapting to various practical application needs, including neural network reasoning, signal processing, data mining and other scenarios, greatly improving the performance and versatility of the overall processing system.
[0092] In order to enable those skilled in the art to better understand the present application, the present application will be explained below in conjunction with a specific example and its test data.
[0093] In a specific embodiment of the present application, the specific values of the characteristic value interval and the target order can be designed according to the following Table 2: Table 2 Examples of predefined operation extension instructions In this application, according to the design shown in Table 2, 512 sets of feature vectors are normalized and the test results shown in Table 3 are obtained: Table 3 Test results In this application, the test results in Table 3 show that the 6th-order Taylor expansion scheme has a delay of only 0.0000025 seconds, an energy efficiency of 0.5 μJ / inference, and an accuracy of 98.8%. The 9th-order Taylor expansion scheme has a slightly higher delay (0.0000028 seconds), but its accuracy is improved to 99.1%, and its energy efficiency is only 0.7 μJ / inference. Compared with traditional exponential formulas, both the 6th-order and 9th-order Taylor expansions are significantly better than the exponential scheme. While ensuring high accuracy, the delay is reduced by more than 40%, and energy consumption is saved by 40% to 60%.
[0094] This solution achieves hardware acceleration of normalization operations through RISCV extended instructions, greatly simplifying the calling method and eliminating multiple layers of function encapsulation, thereby improving development efficiency by approximately 40%. It supports dynamic switching between 6th-order and 9th-order Taylor expansions. While ensuring controllable accuracy (the 9th-order mode has an error of less than 0.3%, and the 6th-order mode meets the accuracy standard within a small range), the 6th-order mode saves approximately 25% of LUT resources compared to the fixed 9th-order solution, flexibly adapting to different power consumption requirements. At the same time, a single hardware instruction can process four operations (addition, subtraction, multiplication, and division) in parallel, speeding up by three times compared to traditional loop solutions, with the minimum latency reduced to three cycles. By reusing 0th- to 6th-order computing units, the high-order mode only supplements the 7th- to 9th-order paths, further saving approximately 20% of hardware resources and achieving a perfect combination of high efficiency and high flexibility.
[0095] In this application, it is possible to support static setting of the default order by software through the order configuration register, so as to facilitate the use of optimal resource configuration in fixed scenarios such as speech recognition.
[0096] In this application, in terms of alternative algorithms, the Taylor expansion scheme can be replaced by a 6th-order Chebyshev approximation scheme. Although it can achieve an accuracy comparable to that of the Taylor expansion, it will increase the hardware multiplier resource consumption by approximately 15%; In this application, the switching of 2nd / 4th / 8th order Taylor expansion formulas can be supported by the segmented index lookup table method, which is highly flexible, but the BRAM resource usage is about 40% higher than that of this solution.
[0097] The following describes an embodiment of the device of the present application, which can be used to execute the feature data processing method based on an edge device in the above-mentioned embodiment of the present application. A neural network model is deployed in the processor of the edge device, and the neural network model includes a convolutional network and a probabilistic normalization function. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the feature data processing method based on an edge device in the above-mentioned embodiment of the present application.
[0098] See also Figure 3 , shows a block diagram of a feature data processing device based on edge devices in an embodiment of the present application.
[0099] like Figure 3 As shown, the feature data processing device 200 based on the edge device according to an embodiment of the present application includes: an acquisition unit 201, a first operation unit 202 and a second operation unit 203.
[0100] Among them, the acquisition unit 201 is used to obtain the feature data to be processed; the first operation unit 202 is used to perform a convolution operation on the feature data based on the convolutional network to obtain a feature vector of the feature data; the second operation unit 203 is used to replace the exponential formula in the probability normalization function with a Taylor expansion formula, and perform a normalization operation on the eigenvalues in the feature vector based on the probability normalization function to obtain the probability distribution corresponding to the feature data, and the probability distribution is used to classify the feature data.
[0101] In some embodiments of the present application, based on the aforementioned scheme, the device further includes: a determination unit, for determining a target order Taylor expansion formula for the probability normalization function for processing each eigenvalue in the eigenvector before performing a normalization operation on the eigenvalue in the eigenvector based on the probability normalization function, wherein the target order Taylor expansion formula is used to replace the exponential formula in the probability normalization function, and the target order is adapted to the size of each eigenvalue.
[0102] In some embodiments of the present application, based on the aforementioned solution, the target order is positively correlated with the magnitude of each eigenvalue.
[0103] In some embodiments of the present application, based on the aforementioned scheme, the determination unit is configured to: determine a target operation extension instruction from a plurality of predefined operation extension instructions according to each eigenvalue in the eigenvector; and determine a Taylor expansion formula of a corresponding target order for the probability normalization function for processing each eigenvalue according to the target operation extension instruction.
[0104] In some embodiments of the present application, based on the above-mentioned scheme, the determination unit is configured as follows: if each of the eigenvalues falls into the first eigenvalue interval, the first operation extension instruction among multiple predefined operation extension instructions is determined as the target operation extension instruction, and the target order corresponding to the first operation extension instruction is N order; if each of the eigenvalues falls into the second eigenvalue interval, the second operation extension instruction among multiple predefined operation extension instructions is determined as the target operation extension instruction, and the target order corresponding to the second operation extension instruction is M order, wherein M is greater than N, and the minimum value of the second eigenvalue interval is greater than the maximum value of the first eigenvalue interval.
[0105] In some embodiments of the present application, based on the aforementioned solution, the N order is 6 orders, and the M order is 9 orders.
[0106] In some embodiments of the present application, based on the aforementioned solution, the target operation extension instruction is used to directly call a probability normalization function to perform a normalization operation on each of the eigenvalues.
[0107] In some embodiments of the present application, based on the foregoing scheme, the processor comprises a first calculation module and a second calculation module, wherein the first calculation module is configured to calculate the first order formula to the Nth order formula in the Taylor expansion formula, and the second calculation module is configured to calculate the N+1th order formula to the Mth order formula in the Taylor expansion formula, and the second operation unit 203 is configured to: in the process of normalizing the feature values in the feature vector based on the probability normalization function, if the target order is N, the first calculation module is called to calculate the Nth order Taylor expansion formula; if the target order is M, the first calculation module and the second calculation module are called to calculate the Mth order Taylor expansion formula.
[0108] In some embodiments of the present application, based on the foregoing scheme, the first calculation module comprises N calculation units, and the second operation unit 203 is configured to call each calculation unit in the first calculation module and calculate each order formula in the Nth order Taylor expansion formula in parallel.
[0109] In some embodiments of the present application, based on the foregoing scheme, the second calculation module comprises M-N calculation units, and the second operation unit 203 is configured to call each calculation unit in the first calculation module and each calculation unit in the second calculation module and calculate each order formula in the Mth order Taylor expansion formula in parallel.
[0110] In some embodiments of the present application, based on the foregoing scheme, the second operation unit 203 is configured to normalize each feature value in the feature vector in parallel based on the probability normalization function.
[0111] Based on the same inventive concept, the embodiments of the present application provide a computer program product, which comprises computer instructions stored in a computer readable storage medium and adapted to be read and executed by a processor to enable a computer device having the processor to perform the operations performed by the edge device-based feature data processing method as described above.
[0112] Based on the same inventive concept, the embodiments of the present application provide a computer readable storage medium, which stores at least one computer program instruction, and the at least one computer program instruction is loaded and executed by a processor to enable the computer readable storage medium to perform the operations performed by the edge device-based feature data processing method as described above.
[0113] Based on the same inventive concept, the embodiments of the present application further provide an electronic device, which refers to Figure 3, shows a structural diagram of an electronic device in an embodiment of the present application, wherein the electronic device includes one or more memories 304, one or more processors 302, and at least one computer program (computer program instruction) stored in the memory 304 and executable on the processor 302. When the processor 302 executes the computer program, the feature data processing method based on the edge device as described above is implemented.
[0114] Among them, Figure 3 In the present invention, a bus architecture (represented by bus 300) is shown. Bus 300 may include any number of interconnected buses and bridges. Bus 300 links various circuits, including one or more processors represented by processor 302 and memory represented by memory 304. Bus 300 may also link various other circuits, such as peripherals, voltage regulators, and power management circuits, all of which are well known in the art and, therefore, will not be described further herein. Bus interface 305 provides an interface between bus 300 and receiver 301 and transmitter 303. Receiver 301 and transmitter 303 may be the same component, namely a transceiver, which provides a means for communicating with various other devices over a transmission medium. Processor 302 is responsible for managing bus 300 and general processing, while memory 304 may be used to store data used by processor 302 when performing operations.
[0115] The functions described herein may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored as one or more instructions or codes on or transmitted via a computer-readable medium. Other examples and implementations are within the scope and spirit of this application and the appended claims. For example, due to the nature of software, the functions described above may be implemented using software executed by a processor, hardware, firmware, hardwiring, or a combination of any of these. Furthermore, the functional units may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0116] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0117] The units described as separate components may or may not be physically separate, and the components of the control device may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0118] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store computer program instructions.
[0119] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of the claims of the present application.
Claims
1. A feature data processing method based on edge devices, characterized in that: A neural network model is deployed in a processor of the edge device, the neural network model includes a convolutional network and a probability normalization function, and the method includes: Obtain feature data to be processed; Performing a convolution operation on the feature data based on the convolutional network to obtain a feature vector of the feature data; When the exponential formula in the probability normalization function is replaced by a Taylor expansion formula, the eigenvalues in the eigenvector are normalized based on the probability normalization function to obtain the probability distribution corresponding to the feature data, and the probability distribution is used to classify the feature data.
2. The method according to claim 1, characterized in that Before performing a normalization operation on the eigenvalues in the eigenvector based on the probability normalization function, the method further includes: A Taylor expansion formula of a target order is determined for the probability normalization function for processing each eigenvalue in the eigenvector. The Taylor expansion formula of the target order is used to replace the exponential formula in the probability normalization function. The target order is adapted to the size of each eigenvalue.
3. The method according to claim 2, characterized in that The target order is positively correlated with the magnitude of each eigenvalue.
4. The method according to claim 2, characterized in that The Taylor expansion formula for determining the target order of the probability normalization function for processing each eigenvalue in the eigenvector includes: determining a target arithmetic extension instruction from a plurality of predefined arithmetic extension instructions according to each eigenvalue in the eigenvector; According to the target operation extension instruction, a Taylor expansion formula corresponding to the target order is determined for the probability normalization function for processing each eigenvalue.
5. The method according to claim 4, characterized in that The step of determining a target operation extension instruction from a plurality of predefined operation extension instructions according to each eigenvalue in the eigenvector includes: If each of the characteristic values falls within a first characteristic value interval, a first operation extension instruction among a plurality of predefined operation extension instructions is determined as the target operation extension instruction, and the target order corresponding to the first operation extension instruction is N order; If each of the eigenvalues falls into the second eigenvalue interval, the second operation extension instruction among the multiple pre-defined operation extension instructions is determined as the target operation extension instruction, and the target order corresponding to the second operation extension instruction is M order, where M is greater than N, and the minimum value of the second eigenvalue interval is greater than the maximum value of the first eigenvalue interval.
6. The method according to claim 5, characterized in that The N order is 6 orders, and the M order is 9 orders.
7. The method according to claim 5, characterized in that The target operation extension instruction is used to directly call the probability normalization function to perform a normalization operation on each of the eigenvalues.
8. The method according to claim 5, characterized in that The processor includes a first calculation module and a second calculation module, wherein the first calculation module is used to calculate the 1st order formula to the Nth order formula in the Taylor expansion formula, and the second calculation module is used to calculate the N+1th order formula to the Mth order formula in the Taylor expansion formula. In the process of normalizing the eigenvalues in the eigenvector based on the probability normalization function, the method further includes: If the target order is N, calling the first calculation module to calculate the N-order Taylor expansion formula; If the target order is M, the first calculation module and the second calculation module are called to calculate the M-order Taylor expansion formula.
9. The method according to claim 8, characterized in that The first calculation module includes N calculation units, and calling the first calculation module to calculate the N-order Taylor expansion formula includes: The various calculation units in the first calculation module are called to perform calculations on the various order formulas in the N-order Taylor expansion formula in parallel.
10. The method according to claim 9, characterized in that The second calculation module includes MN calculation units, and calling the first calculation module and the second calculation module to calculate the M-order Taylor expansion formula includes: Each calculation unit in the first calculation module and each calculation unit in the second calculation module are called to perform calculations on each order formula in the M-order Taylor expansion formula in parallel.
11. The method according to claim 1, wherein The performing a normalization operation on the eigenvalues in the eigenvector based on the probability normalization function includes: Based on the probability normalization function, normalization operations are performed on each eigenvalue in the eigenvector in parallel.
12. A feature data processing device based on edge devices, characterized in that: A neural network model is deployed in the processor of the edge device, and the neural network model includes a convolutional network and a probability normalization function. The apparatus includes: An acquisition unit, used for acquiring feature data to be processed; A first operation unit is configured to perform a convolution operation on the feature data based on the convolutional network to obtain a feature vector of the feature data; The second operation unit is used to perform a normalization operation on the eigenvalues in the eigenvector based on the probability normalization function, when the exponential formula in the probability normalization function is replaced by a Taylor expansion formula, to obtain a probability distribution corresponding to the feature data, and the probability distribution is used to classify the feature data.
13. A computer program product, characterized in that The computer program product includes computer instructions stored in a computer-readable storage medium and adapted to be read and executed by a processor, so as to enable a computer device having the processor to perform the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one program code, and the at least one program code is loaded and executed by a processor to implement the operations performed by the method according to any one of claims 1 to 11.
15. An electronic device, characterized in that: The electronic device includes one or more processors and one or more memories, wherein at least one program code is stored in the one or more memories, and the at least one program code is loaded and executed by the one or more processors to implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Data processing method for neural network, accelerator and electronic equipment
CN118709729A
Grounding grid corrosion defect identification method based on edge enhancement model, medium and equipment
CN119559510A
Normalized exponential operation approximation method and neural network applying same
CN120030270A
Unmanned aerial vehicle real-time vegetation classification system and method based on lightweight AI model
CN120526334A
Method for quantizing ultra-light network model, and low-power inference accelerator applying same
WO2025127188A1