Rapid calibration method, device, equipment and system for edge end equipment
By only adjusting the weight of the digital memory layer in the edge computing device to compensate for the write error of the analog memory layer, combined with the advantages of analog and digital memory, the problems of weight writing error and high power consumption in the edge computing device are solved, and efficient and low-power calculation effects are achieved.
Patent Information
- Application Number
- CN202411877143.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-06
AI Technical Summary
In edge computing devices, the analog memory layer weight writing error is large and the adjustment is difficult, and traditional architectures consume high power and lack flexibility.
By using the digital memory layer of the neural network model in the edge-end device for rapid calibration, only the weight of the digital memory layer is adjusted to compensate for the write error and drift effect in the analog memory process, combining the advantages of analog and digital memory, reducing power consumption and improving computing efficiency.
It improves the accuracy of inference and the stability and reliability of the system, reduces the overall training time and the use of computing resources, achieves a balance between low power consumption and high precision, and improves the personalized adaptation ability and response speed of the equipment.
Smart Images

Figure CN119940436A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and system for rapid calibration of edge devices. Background Art
[0002] With the rapid development of artificial intelligence technology, the demand for low-power, high-computing-efficiency edge AI devices is growing. As a distributed computing model, edge computing effectively reduces latency, improves response speed, and reduces dependence on central servers by processing data close to the data source. This is especially important for application scenarios with high real-time requirements, such as smart homes, industrial monitoring, medical equipment, and autonomous driving. In these applications, edge computing devices must balance the relationship between power consumption and computing performance. The traditional von Neumann architecture has limitations in power consumption and efficiency due to its inherent memory wall problem. Therefore, it is necessary to explore new storage and computing technologies to meet the needs of fast response and high efficiency in edge computing.
[0003] RRAM (Resistive Random Access Memory) has the characteristics of high density, low power consumption, non-volatility and good scalability, making it an important choice for realizing analog and digital storage and computing. It is suitable for application scenarios that require fast response and high efficiency in edge computing. RRAM arrays can be used as weight storage units in neural networks or as core units for matrix operations. Compared with traditional memory, RRAM can use its storage characteristics to directly complete multiplication and addition operations when performing matrix operations, greatly improving computing efficiency. Therefore, it is very suitable for resource-constrained situations in edge computing environments.
[0004] Although RRAM brings many advantages to edge computing, RRAM chips have some defects in practical applications, such as weight writing errors, retention performance issues, weight drift and relaxation effects, which seriously affect the calculation accuracy of RRAM. Weight writing error: Due to the writing mechanism of RRAM, the weight may deviate during the writing process, thus affecting the reasoning accuracy of the neural network. Retention problem: The weight of RRAM may decay after long-term storage, resulting in reduced accuracy of stored information. The retention time is closely related to environmental conditions. Weight drift and relaxation effect: The weight may change during storage and reading, especially under the influence of environmental factors such as temperature changes, the weight value will drift, resulting in the accumulation of calculation errors. In the application of digital storage and computing, the accurate digital value of the RRAM chip can be read out by sacrificing a certain number of bits and passed into the register for calculation. However, this will sacrifice the high efficiency brought by analog storage and computing.
[0005] In the current specific technology, analog RRAM is used as the carrier of edge computing. After the weights are trained by the end-side device or cloud server, they are written into the RRAM chip. However, due to the noise problem of RRAM, the accuracy of its model cannot be guaranteed, and the self-learning ability of the RRAM edge is limited, and only simple task updates can be performed. At the same time, the large-scale weight writing speed of RRAM is slow, so the flexibility of this technical solution is also poor. Summary of the invention
[0006] The present application provides a method, apparatus, device and system for rapid calibration of edge devices to solve the problems of large error in writing weights of analog storage and computing layers in edge computing devices in related technologies, difficulty in adjustment, high power consumption and insufficient flexibility of traditional architectures, etc.
[0007] The first aspect of the present application provides a method for rapid calibration of an edge device, comprising the following steps: obtaining user data collected by the edge device; training a digital storage and computing layer of a neural network model based on the user data, wherein the analog storage and computing weights corresponding to the analog storage and computing layer of the neural network model remain unchanged during the training process, wherein the digital storage and computing layer uses the digital storage and computing part of the storage and computing integrated architecture for calculation, and the analog storage and computing layer uses the analog storage and computing part of the storage and computing integrated architecture for calculation; and correcting the digital storage and computing weights corresponding to the digital storage and computing layer based on the training results.
[0008] Optionally, in one embodiment of the present application, the digital storage layer includes an expansion layer and a reduction layer, wherein the digital storage part of the expansion layer expands the number of input channels, and the reduction layer compresses the number of input channels based on the digital storage part.
[0009] Optionally, in one embodiment of the present application, a SA (Sense Amplifier) reads the digital storage weight and puts it into a digital computing module on the edge device chip, and the digital computing module completes the digital storage and outputs it.
[0010] Optionally, in one embodiment of the present application, the simulation storage and computing layer includes a computing layer, wherein the computing layer processes the expanded input channel based on the simulation storage and computing part.
[0011] Optionally, in one embodiment of the present application, after the analog storage and computing part completes the matrix multiplication and addition operation of the analog storage and computing, under the control of the row and column switches, the intermediate result of the neural network model calculation is obtained by current accumulation and sampling by the back-end analog and digital converter.
[0012] Optionally, in one embodiment of the present application, the storage and computing integrated architecture includes a memristor, a phase change memory, or a magnetoresistive memory.
[0013] Optionally, in one embodiment of the present application, after correcting the digital storage computing weights corresponding to the digital storage computing layer according to the training results, it also includes: writing the corrected digital storage computing weights into the edge device.
[0014] The second aspect of the present application provides a fast calibration device for an edge device, including: an acquisition module for acquiring user data collected by the edge device; a training module for training a digital storage and computing layer of a neural network model based on the user data, wherein the analog storage and computing weights corresponding to the analog storage and computing layer of the neural network model remain unchanged during the training process, wherein the digital storage and computing layer uses the digital storage and computing part of the storage and computing integrated architecture for calculation, and the analog storage and computing layer uses the analog storage and computing part of the storage and computing integrated architecture for calculation; and a correction module for correcting the digital storage and computing weights corresponding to the digital storage and computing layer according to the training results.
[0015] The third aspect of the present application provides an edge device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the above-mentioned methods for rapid calibration of edge devices.
[0016] The fourth aspect of the present application provides a rapid calibration system for edge devices, including: an edge device, the edge device is deployed with a neural network model and adopts an integrated storage and computing architecture that combines analog and digital, wherein the neural network model includes a digital storage and computing layer and an analog storage and computing layer, the digital storage and computing layer uses the digital storage and computing part of the integrated storage and computing architecture for calculation, and the analog storage and computing layer uses the analog storage and computing part of the integrated storage and computing architecture for calculation; an end-side device, the end-side device obtains user data collected by the edge device, trains the digital storage and computing layer of the neural network model according to the user data, the analog storage and computing weights corresponding to the analog storage and computing layer of the neural network model remain unchanged during the training process, and the digital storage and computing weights corresponding to the digital storage and computing layer are corrected according to the training results.
[0017] Therefore, this application includes the following beneficial effects:
[0018] The embodiment of the present application effectively compensates for the write error and drift effect that may occur in the analog storage process by only adjusting the weight of the digital storage layer, improves the inference accuracy, and ensures the stability and reliability of the model under different environmental conditions; only the digital storage layer is trained, which reduces the overall training time and the use of computing resources, and avoids the complexity and uncertainty caused by large-scale weight updates; the expansion layer increases the number of input channels to capture more feature information, and the reduction layer compresses the number of channels to reduce the processing complexity, which not only enhances the model's expressiveness but also ensures high efficiency; combining the advantages of analog and digital storage, reduces data handling overhead, improves computing efficiency, and achieves a balance between low power consumption and high precision; The physical properties of the RRAM array are used to directly complete multiplication and addition operations, greatly improving computing efficiency; the results of analog storage and calculation are read through the high-precision sensing amplifier SA, ensuring the accuracy of the calculation results and improving the stability and reliability of the system; using memristors, phase change memories or magnetoresistive memories as components of the storage and computing integrated architecture not only achieves the advantages of low power consumption and high speed, but also increases the flexibility of system design; rewriting the optimized digital storage and computing weights into the edge device allows the device to immediately apply the latest adjustments, enhancing the real-time and response speed of the system, and quickly adapting to new application scenarios or changes in user needs without the need for long-term retraining of the entire network, greatly improving user experience and service quality.
[0019] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0021] Figure 1 A schematic diagram of the architecture of a fast calibration system for an edge device provided according to an embodiment of the present application;
[0022] Figure 2 A schematic diagram of a process for quickly calibrating an edge device according to an embodiment of the present application;
[0023] Figure 3 A schematic diagram of a storage-computing integrated array based on a crossbar switch matrix structure provided according to an embodiment of the present application;
[0024] Figure 4 A neural network model provided according to an embodiment of the present application;
[0025] Figure 5 A block diagram of a fast calibration device for an edge terminal device provided according to an embodiment of the present application;
[0026] Figure 6 A schematic diagram of the structure of a terminal device provided according to an embodiment of the present application;
[0027] Figure 7 A system block diagram of a rapid calibration system for edge devices provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0028] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0029] With the rapid development of artificial intelligence technology, the demand for low-power, high-computing-efficiency edge AI devices is increasing. Edge computing refers to data processing close to the data source, which can reduce latency, improve response speed, and reduce dependence on central servers. In edge computing devices, low-power and high-efficiency storage and computing technologies are crucial. At the same time, in the process of applying AI models on edge devices, due to differences in different users, usage environments or usage scenarios, AI models often need to be quickly calibrated to ensure their optimal performance.
[0030] As devices close to the user group, edge devices are often used in a variety of scenarios, such as smart homes, industrial monitoring, medical equipment, and autonomous driving. In these different scenarios, the distribution of data collected by the device will vary depending on environmental conditions (such as temperature, humidity, and light) and sensor characteristics (such as cameras, microphones, and pressure sensors). This change in data distribution may lead to a decrease in the prediction accuracy of the model, so the AI model needs to quickly adapt to the data distribution in different environments. In addition, each user may have significant differences in behavioral characteristics, operating methods, or health conditions. For example, in wearable devices, there are individual differences in physiological parameters (such as heart rate and gait) among different users, and the model needs to be fine-tuned for specific users to improve the accuracy of health monitoring. Similarly, there are large differences in speech recognition models between different voice features and accents, which requires edge AI models to be able to quickly adjust according to the characteristics of the user.
[0031] However, the related technologies have problems such as low efficiency, high power consumption, low precision, long time consumption, and low flexibility, as follows:
[0032] 1. Compared with edge computing using traditional computing architecture, this solution has low efficiency, high power consumption, and insufficient computing power and battery life: In edge computing, the use of traditional computer architecture (such as CPU, GPU) leads to low computing efficiency and high power consumption due to the limitations of instruction set and memory bottlenecks. This architecture cannot meet the edge device's requirements for high efficiency, low power consumption, and high computing power, limiting the performance and battery life of the device.
[0033] 2. Compared with the analog in-memory computing fully implemented with RRAM, the accuracy of this solution is affected by weight error: Although RRAM has the advantages of high efficiency and low power consumption in edge computing, in actual applications, weight writing errors, weight drift and retention performance issues will lead to reduced computing accuracy. The noise introduced by weight error reduces the accuracy of model reasoning and affects the reliability of the system.
[0034] 3. Compared with the analog in-memory computing fully implemented with RRAM, this solution takes a long time to write weights and lacks flexibility: The large-scale weight writing process takes a long time in RRAM, resulting in slower model updates and deployment. Due to the long writing time, it is difficult to adjust according to different users and usage scenarios in a timely manner, and it cannot meet the needs of personalization and rapid response.
[0035] 4. In addition, even compared with the ASIC (Application-Specific Integrated Circuit) solution, this solution has high design complexity and low flexibility: Although the use of dedicated integrated circuits can improve computing efficiency and reduce power consumption, the design complexity is high and the R&D cycle is long. The fixed characteristics of ASIC lead to its poor flexibility and inability to adapt to changing application requirements and environmental conditions.
[0036] Edge devices are often limited by computing resources and storage capacity, making it difficult to run large models or perform complex model training. This means that during deployment, efficient fast calibration methods need to be developed to quickly tune the model with limited resources and avoid the high computing costs and time consumption caused by retraining the model. Fast calibration is generally achieved by adjusting some weights in the network.
[0037] In order to adapt to the personalized needs of edge devices while compensating for the noise introduced by weight errors in RRAM analog storage and calculation, an embodiment of the present application proposes a fast calibration method for edge devices. During the calibration process, the terminal-side smart device (such as a mobile phone, PC, etc.) can be connected to an external terminal-side smart device (such as a mobile phone, PC, etc.) through modules such as Bluetooth and WiFi for training. Before the start of the fast calibration, the edge device will check its own status and send information related to the weight deployment to the terminal-side device. The fast calibration process includes three steps: data collection, training, and edge-side writing: the data collected by the edge device is first sent to the terminal-side device, and then the terminal-side device uses the back propagation algorithm to assist in training the expansion layer and reduction layer of the network based on the actual status of the edge device and the collected data. The calculation layer does not need to be trained and adjusted because the number of weights accounts for a large proportion, thereby reducing the parameters that need to be adjusted and improving the training efficiency. After the weight optimization of the digital storage and calculation part is completed, it is sent back to the edge device by communication, and the new RRAM digital storage and calculation weights are written.
[0038] like Figure 1 As shown, the user's personalized domain data is sent to a high-performance mobile phone or server on the end side, and combined with the analog storage weights written on the chip for auxiliary training to update the digital storage weights of the expansion layer and the reduction layer. This fast calibration training process can also make up for the weight noise brought by the RRAM computing layer, because this weight noise is essentially a special domain, so it is equivalent to adaptive calibration of the analog storage hardware. Specifically, the real noisy weights in the RRAM analog storage part will be carried during training, and only the digital storage part will be trained to quickly adjust the weights of the digital layer. This training method not only completes the rapid calibration of the domain, but also adapts to the weight writing noise of analog storage and compensates for the loss of precision in analog computing. Through this on-chip training or external training, edge devices can quickly switch between different application environments or users, adapt to new data distributions, and improve the inference accuracy of the model.
[0039] Through this architectural design, the present application can effectively combine the advantages of analog and digital storage and computing technologies, achieving both efficient computing with low power consumption and ensuring high-precision output of the system. In addition, by dividing the network into different structures, the depth and complexity of the network can be flexibly adjusted to meet the needs of different application scenarios, thereby improving the adaptability and scalability of the system.
[0040] As a result, it is achieved that the problem of reduced accuracy caused by weight errors in RRAM analog memory calculations is alleviated under the condition of the lowest possible power consumption, the problem of fast calibration flexibility for different users, usage environments and scenarios is solved, and the personalized adaptation ability of the device is improved. By adjusting the digital storage layer, the present application can use a small amount of computing resources to complete the rapid fine-tuning of the model, thereby significantly shortening the adaptation time of the system. This fast calibration method is particularly suitable for edge devices, because the edge usually has limited computing and storage resources, and cannot perform complex model training and parameter adjustment like cloud devices. Through the method of the present application, the model can be quickly adjusted to adapt to different application environments without changing the analog storage part, such as switching from a home environment to an industrial environment, or switching from a specific user to another user, thereby maintaining high reasoning accuracy and reliability in a wide range of application scenarios.
[0041] Specifically, Figure 2 A flowchart of a rapid calibration method for edge devices provided in an embodiment of the present application.
[0042] like Figure 2 As shown, the fast calibration method of the edge device includes the following steps:
[0043] In step S201, user data collected by the edge device is obtained.
[0044] It can be understood that the embodiments of the present application ensure that the training and adjustment processes are based on the most realistic usage scenarios by obtaining user data collected by edge devices from the actual environment, which not only improves the adaptability of the model to specific environments or users, but also enables the model to better capture individual differences, thereby improving the quality of personalized services.
[0045] In step S202, the digital storage and computing layer of the neural network model is trained according to the user data. During the training process, the analog storage and computing weights corresponding to the analog storage and computing layer of the neural network model remain unchanged. The digital storage and computing layer uses the digital storage and computing part of the storage and computing integrated architecture for calculation, and the analog storage and computing layer uses the analog storage and computing part of the storage and computing integrated architecture for calculation.
[0046] It can be understood that the embodiment of the present application only trains the digital storage and computing layer, while keeping the weights of the analog storage and computing layer unchanged, so that fast calibration can be achieved without affecting the existing analog calculations. Since the number of weights in the analog storage and computing layer is large and the writing time is long, fixing this part of the weights can significantly reduce training time and resource consumption, while avoiding the complexity and uncertainty brought about by large-scale weight updates.
[0047] In an embodiment of the present application, the digital storage layer includes an expansion layer and a reduction layer.
[0048] Among them, the digital storage part of the expansion layer expands the number of input channels, and the reduction layer compresses the number of input channels based on the digital storage part; the expansion layer uses 1×1 convolution or full connection operation to expand the number of input channels from n to k*n to increase the computing capacity. The expansion layer is implemented through digital in-memory calculation, and the output will be used as the input of the calculation layer to participate in further calculations; the reduction layer uses 1×1 convolution or full connection operation to compress the number of input channels from k*n back to n. The reduction layer is implemented through digital in-memory calculation, and the output passes through operations such as activation functions and then enters the next layer of the network.
[0049] It can be understood that the digital storage layer of the embodiment of the present application includes an expansion layer and a reduction layer. The expansion layer increases the number of input channels to capture more feature information, and the reduction layer compresses the number of channels to reduce the complexity of subsequent processing. The existence of the expansion layer and the reduction layer not only enhances the expressiveness of the model, but also ensures high efficiency. Especially for edge devices, this lightweight structure reduces the computing burden and helps to provide more accurate services under limited resources.
[0050] Specifically, the expansion layer and the reduction layer use digital storage and computing. For digital storage and computing integration, in this mode, the RRAM array stores the weight information required for calculation, and the array is surrounded by a digital computing module for performing further numerical calculations. Digital storage and computing directly reads the weight data in the RRAM array into the register in a serial accelerated manner, and then performs subsequent calculations and processing through the register. At the same time, the digital computing circuit is a noise-free computing mode, which can compensate for the errors introduced in the analog storage and computing, and flexibly adjust the weights based on the scenario, thereby improving the overall calculation accuracy of the model in different states; in addition, low-bit-width data, such as 8bit or 4bit, can be used for digital calculations to save more energy.
[0051] In the embodiment of the present application, SA reads the digital storage and calculation weights and puts them into the digital calculation module on the edge device chip, and the digital calculation module completes the digital storage and calculation and then outputs it.
[0052] It can be understood that the embodiment of the present application can effectively compensate for the noise and errors introduced in the analog storage and calculation process by accurately reading the digital storage and calculation weights in RRAM through SA, and passing these weights to the on-chip digital computing module for processing. The digital computing module can perform calculations in a noise-free mode, ensuring high-precision result output, thereby improving the accuracy of model reasoning. In addition, SA directly reads the weights from RRAM and transmits them to the on-chip digital computing module, avoiding the frequent movement of data between computing and storage in traditional architectures, further reducing overall power consumption.
[0053] In the embodiment of the present application, the simulation storage and computing layer includes a computing layer.
[0054] Among them, the computing layer processes the expanded input channels based on the analog storage and computing part. The computing layer uses 3×3 or 5×5 convolution to process the expanded channels. It mainly implements analog storage and computing through the RRAM array. The output of the calculation will be stored in the register and participate in the calculation of the reduction layer.
[0055] It can be understood that the simulation layer of the embodiment of the present application includes a computing layer, which is responsible for processing the expanded input channels, using the advantages of analog storage to perform efficient matrix operations, and providing powerful computing support for the neural network.
[0056] In an embodiment of the present application, after the analog storage and computing part completes the matrix multiplication and addition operation of the analog storage and computing, under the control of the row and column switches, the intermediate result of the neural network model calculation is obtained by current accumulation and sampling by the analog and digital converter at the back end.
[0057] It can be understood that after completing the matrix multiplication and addition operation of the analog storage and calculation part, the embodiment of the present application obtains the intermediate result by current accumulation and sampling by the analog and digital converter at the back end under the control of the row and column switches, making full use of the physical characteristics of the RRAM array and the advantages of analog storage and calculation to achieve high-efficiency matrix operations: RRAM not only serves as a weight storage unit, but also can directly complete multiplication and addition operations in the analog domain, greatly improving the computing efficiency, reducing energy consumption by reducing data movement, and supporting highly parallel current accumulation operations, significantly improving the computing speed. The precise sampling of the analog and digital converter ensures the acquisition of high-quality intermediate results. This flexible design can adapt to different types of neural network layers and task requirements, thereby improving the performance, energy efficiency and response speed of neural network reasoning as a whole.
[0058] It should be noted that the CIM (Computing-in-Memory) crossbar array (abbreviated as analog memory and computing integrated array) is an efficient analog computing device that can be used to realize large-scale matrix-vector multiplication and addition operations, and is usually widely used in AI (Artificial Intelligence) and neural network calculations. The basic computing unit of its circuit is usually a conductance or charge modulated circuit device, such as a memristor. By gating the row and column switches, the current is accumulated and sampled in the analog domain to achieve efficient matrix multiplication operations.
[0059] The computing layer uses RRAM analog storage and computing, and can use the computing characteristics of its memory array to directly complete matrix multiplication and addition operations. Under the control of row and column switches, the intermediate results of neural network calculations are obtained by current accumulation and sampling by the back-end ADC (Analog-to-Digital Converter). This analog storage and computing method not only has efficient computing power, but also can greatly reduce energy consumption.
[0060] For example, Figure 3 As shown in (a), taking the RRAM analog storage and computing integrated array for neural network reasoning (matrix multiplication) as an example, the weight information is usually stored in the analog storage and computing integrated device in the form of conductance (G), and the row and column switches are selected by using the word line (WL), and different levels of voltage are applied to the bit line (BL) using the digital-to-analog conversion device DAC to obtain the current (Itot). After the current is collected at the source line (SL) end and sampled by the analog-to-digital conversion device (ADC), the results of multiple sets of matrix multiplication and addition operations can be obtained in parallel.
[0061] The specific calculation process is as follows Figure 3 (b) shows in detail Figure 3 (a) shows a row in the storage-computation integrated array. The weight values of the neural network are quantized to integers and stored in the RRAM unit. WL controls all transistor switches in the row circuit. BL is the input voltage of each RRAM, and the voltage value (V1-Vn) is adjusted to correspond to different input signals. The input signal is a neuron or feature map quantized to an integer. According to Kirchhoff's current law, the current passing through each RRAM is V×G. All currents flowing through the RRAM will converge on SL and be sampled as digital signals by the back-end ADC to complete a multiplication and addition operation in the analog domain. The current formula in the analog domain is as follows:
[0062] I tot =V1×G1+V2×G2+...+V n ×G n
[0063] Among them, I tot [A] is the total current, V i(i=1,2,...,n) [V] is the voltage applied to each memristor, G i(i=1,2,...,n) [S] is the conductance value of each memristor, and the corresponding dimension is in square brackets. When performing neural network inference, the input data X and weight value W need to be mapped to the input voltage V and conductance value G respectively.
[0064] In an embodiment of the present application, the storage and computing integrated architecture includes a memristor, a phase change memory or a magnetoresistive memory.
[0065] It can be understood that in addition to memristors, the storage and computing integrated architecture of the embodiment of the present application can also use PCM (Phase-Change Memory) and MRAM (Magnetoresistive Random Access Memory) as alternatives for analog storage and computing. These alternative devices can also achieve efficient matrix operations and reduce write errors to a certain extent. In some scenarios, they may show better performance or cost-effectiveness, increasing the flexibility of system design.
[0066] The following is a detailed description of the basic computing modules in the neural network model through a specific embodiment. Figure 4 As shown in the figure, the basic computing module of the network is divided into expansion layer, computing layer and reduction layer.
[0067] The embodiment of the present application can use this layered architecture, and the part with large amount of calculation is implemented by RRAM analog storage and computing, ensuring that the low power consumption characteristics of analog storage and computing are used as efficiently as possible; while the expansion layer and the reduction layer, which are used for adjustment and correction, are implemented by RRAM digital storage and computing, and sufficient adjustable parameters are retained in the digital computing process that reduces the relatively high energy consumption as much as possible to ensure flexibility and high precision. This combination enables the entire network to achieve efficient computing while also having good error compensation capabilities, thereby improving the accuracy and reliability of reasoning. Among them, all weights are written on the same RRAM array, the digital storage part is read by the SA and then placed on the chip The digital computing module calculates and outputs it, and the analog storage part directly uses the RRAM array to calculate based on Kirchhoff's law, and the result is directly output after passing through the ADC.
[0068] In addition, based on this application, the order of the calculation layer, expansion layer, and reduction layer of the block structure can be adjusted, or different convolution kernel sizes can be used to adapt to more diverse computing needs. Although these variations are different in specific implementation methods, the overall technical principles remain the same, and efficient computing and rapid adaptability are still achieved by combining analog storage and digital storage.
[0069] In step S203, the digital storage weights corresponding to the digital storage layer are corrected according to the training results.
[0070] It can be understood that the embodiments of the present application can effectively compensate for the errors that may occur in the analog storage and computing process, such as the write error and drift effect of RRAM, by correcting the weights of the digital storage and computing layer. This method not only improves the accuracy of reasoning, but also enhances the reliability and robustness of the system, especially when facing changes in different environmental conditions.
[0071] In an embodiment of the present application, after correcting the digital storage computing weights corresponding to the digital storage computing layer according to the training results, it also includes: writing the corrected digital storage computing weights into the edge device.
[0072] It can be understood that the embodiment of the present application rewrites the corrected digital storage weights into the edge device, so that the device can immediately apply the latest adjustments, enhancing the real-time and response speed of the system. This method allows edge devices to quickly adapt to new application scenarios or changes in user needs without waiting for a long time for full network retraining, greatly improving user experience and service quality.
[0073] According to the fast calibration method of the edge device proposed in the embodiment of the present application, the combination of the expansion layer, the computing layer and the shrinking layer in the neural network is realized by adopting an integrated storage and computing architecture combining analog and digital, and the advantages of analog storage and computing are used as much as possible under the premise of ensuring efficient computing, reducing energy consumption and improving the response speed of the system. Specifically, a digital storage and computing layer is inserted between the RRAM analog storage and computing layers, and the calculation results of the RRAM are dynamically calibrated, which effectively compensates for the noise introduced by the weight error, significantly improves the accuracy of the model reasoning, and enhances the reliability of the system. In addition, by only updating the weight of the digital storage and computing layer (accounting for less than 5% of the total weight), the characteristics of fast training and updating are realized, so that the edge device can quickly fine-tune the digital adaptation layer according to different uses, environments and usage scenarios, meet personalized needs, and improve user experience. Finally, the present application can be lightweight calibrated by combining the high efficiency of RRAM analog storage and computing and the high precision of digital storage and computing without significantly increasing the amount of calculation and power consumption, so as to achieve low power consumption, high efficiency and high computing power of edge devices, and improve the endurance and overall performance of the device.
[0074] Next, a rapid calibration device for edge devices according to an embodiment of the present application will be described with reference to the accompanying drawings.
[0075] Figure 5 It is a block diagram of a rapid calibration device for an edge device according to an embodiment of the present application.
[0076] like Figure 5 As shown, the fast calibration device 50 of the edge device includes: an acquisition module 510, a training module 520 and a correction module 530.
[0077] Among them, the acquisition module 510 is used to acquire user data collected by the edge device; the training module 520 is used to train the digital storage and computing layer of the neural network model according to the user data. During the training process, the analog storage and computing weights corresponding to the analog storage and computing layer of the neural network model remain unchanged. The digital storage and computing layer uses the digital storage and computing part of the storage and computing integrated architecture for calculation, and the analog storage and computing layer uses the analog storage and computing part of the storage and computing integrated architecture for calculation; the correction module 530 is used to correct the digital storage and computing weights corresponding to the digital storage and computing layer according to the training results.
[0078] It should be noted that the aforementioned explanation of the embodiment of the rapid calibration method for edge devices is also applicable to the rapid calibration device for edge devices of this embodiment, and will not be repeated here.
[0079] According to the fast calibration device for edge devices proposed in the embodiment of the present application, the acquisition module is responsible for collecting user data from the edge device to ensure that personalized needs and specific characteristics of application scenarios can be fully taken into account; the training module is specifically trained for the digital storage and computing layer based on the user data. During this process, the weight of the analog storage and computing layer remains unchanged, and its efficient analog computing capability is used to handle most of the computing tasks, while the weight of the digital storage and computing layer, which only accounts for less than 5% of the total, is updated, which not only greatly reduces the amount of computing and power consumption required for training, but also ensures that the system can quickly adapt to different usage environments and scenarios; finally, the correction module accurately adjusts the weight of the digital storage and computing layer based on the training results, effectively compensating for the noise and errors that may be introduced in the RRAM analog storage and computing process, and significantly improving the accuracy and reliability of model reasoning.
[0080] Figure 6 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application. The terminal device may include:
[0081] A memory 601 , a processor 602 , and a computer program stored in the memory 601 and executable on the processor 602 .
[0082] When the processor 602 executes the program, the fast calibration method of the edge device provided in the above embodiment is implemented.
[0083] Furthermore, the device also includes:
[0084] The communication interface 603 is used for communication between the memory 601 and the processor 602 .
[0085] The memory 601 is used to store computer programs that can be executed on the processor 602 .
[0086] The memory 601 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.
[0087] If the memory 601, the processor 602 and the communication interface 603 are implemented independently, the communication interface 603, the memory 601 and the processor 602 can be connected to each other through a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0088] Optionally, in a specific implementation, if the memory 601, the processor 602 and the communication interface 603 are integrated on a chip, the memory 601, the processor 602 and the communication interface 603 can communicate with each other through an internal interface.
[0089] The processor 602 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.
[0090] Figure 7 A system block diagram of a rapid calibration system for edge devices provided in an embodiment of the present application.
[0091] like Figure 7 As shown, the fast calibration system 700 of the edge device includes: an edge device 710 and a terminal device 720. Among them, the edge device 710 is deployed with a neural network model and adopts a storage and computing integrated architecture combining analog and digital. The neural network model includes a digital storage and computing layer and an analog storage and computing layer. The digital storage and computing layer uses the digital storage and computing part of the storage and computing integrated architecture for calculation, and the analog storage and computing layer uses the analog storage and computing part of the storage and computing integrated architecture for calculation; the terminal device 720 obtains the user data collected by the edge device, and trains the digital storage and computing layer of the neural network model according to the user data. During the training process, the analog storage and computing weights corresponding to the analog storage and computing layer of the neural network model remain unchanged, and the digital storage and computing weights corresponding to the digital storage and computing layer are corrected according to the training results.
[0092] It can be understood that the embodiments of the present application, through the collaboration between edge devices and terminal devices, can quickly fine-tune the digital storage and computing part while compensating for the weight writing accuracy problems of the analog memory calculation and quickly calibrate to adapt to the changing needs of different users and environments.
[0093] According to the fast calibration system of edge devices proposed in the embodiment of the present application, by combining the advantages of analog and digital storage and computing, an effective balance between low power consumption and high accuracy is achieved in edge devices. The analog storage and computing part provides efficient matrix computing capabilities, ensuring low power consumption and high parallelism of the computing layer, while the digital storage and computing part compensates for errors in analog calculations through precise weight adjustment, enabling the system to provide high-precision reasoning results under resource-constrained conditions, and can be used for flexible adjustment of fast calibration, so that the system can flexibly adapt to different application scenarios and user needs, significantly improving the intelligence level and user experience of edge computing devices.
[0094] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0095] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0096] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0097] It should be understood that the various parts of the present application can be implemented in hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, the steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array, a field programmable gate array, etc.
[0098] A person of ordinary skill in the art may understand that all or part of the steps carried by the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the above-mentioned program may be stored in a computer-readable storage medium, which, when executed, includes one of the steps of the method embodiment or a combination thereof.
[0099] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A fast calibration method for an edge device, characterized in that: The edge device is deployed with a neural network model and adopts a storage and computing integrated architecture combining analog and digital, wherein the method comprises the following steps: Obtain user data collected by edge devices; The digital storage and computing layer of the neural network model is trained according to the user data, and the analog storage and computing weights corresponding to the analog storage and computing layer of the neural network model remain unchanged during the training process, wherein the digital storage and computing layer uses the digital storage and computing part of the storage and computing integrated architecture for calculation, and the analog storage and computing layer uses the analog storage and computing part of the storage and computing integrated architecture for calculation; The digital storage weights corresponding to the digital storage layer are corrected according to the training results.
2. The fast calibration method for edge devices according to claim 1, characterized in that: The digital storage layer includes an expansion layer and a reduction layer, wherein the digital storage part of the expansion layer expands the number of input channels, and the reduction layer compresses the number of input channels based on the digital storage part.
3. The fast calibration method for edge devices according to claim 1 or 2, characterized in that: SA reads the digital storage weight and puts it into the digital computing module on the edge device chip, and the digital computing module completes the digital storage and outputs it.
4. The fast calibration method for edge devices according to claim 1, characterized in that: The simulation storage and computing layer includes a computing layer, wherein the computing layer processes the expanded input channel based on the simulation storage and computing part.
5. The fast calibration method for edge devices according to claim 1 or 4, characterized in that: After the analog storage part completes the matrix multiplication and addition operation of the analog storage, it obtains the intermediate result of the neural network model calculation by accumulating the current and sampling by the analog and digital converter at the back end under the control of the row and column switches.
6. The fast calibration method for edge devices according to claim 1, characterized in that: The storage-computing integrated architecture includes a memristor, a phase change memory or a magnetoresistive memory.
7. The fast calibration method for edge devices according to claim 1, characterized in that: After correcting the digital storage weights corresponding to the digital storage layer according to the training results, the method further includes: The corrected digital storage weight is written into the edge device.
8. A fast calibration device for edge devices, characterized in that: The edge device is deployed with a neural network model and adopts a storage and computing integrated architecture combining analog and digital, wherein the device includes: Acquisition module: used to obtain user data collected by edge devices; Training module: used to train the digital storage and computing layer of the neural network model according to the user data, and the analog storage and computing weights corresponding to the analog storage and computing layer of the neural network model remain unchanged during the training process, wherein the digital storage and computing layer uses the digital storage and computing part of the storage and computing integrated architecture for calculation, and the analog storage and computing layer uses the analog storage and computing part of the storage and computing integrated architecture for calculation; Correction module: used to correct the digital storage weights corresponding to the digital storage layer according to the training results.
9. A terminal side device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for rapid calibration of an edge device as described in any one of claims 1 to 7.
10. A fast calibration system for edge devices, characterized in that: include: An edge device, wherein the edge device is deployed with a neural network model and adopts a storage-computing integrated architecture combining analog and digital, wherein the neural network model includes a digital storage-computing layer and an analog storage-computing layer, the digital storage-computing layer uses the digital storage-computing part of the storage-computing integrated architecture for calculation, and the analog storage-computing layer uses the analog storage-computing part of the storage-computing integrated architecture for calculation; The end-side device obtains user data collected by the edge-side device, trains the digital storage computing layer of the neural network model according to the user data, the analog storage computing weights corresponding to the analog storage computing layer of the neural network model remain unchanged during the training process, and corrects the digital storage computing weights corresponding to the digital storage computing layer according to the training results.