Feature cache set construction method and device for edge-oriented side image classification, and authenticable cache inference method

CN122821233APending Publication Date: 2026-09-25ZHEJIANG UNIV CITY COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611018068.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

当查询样本位于分类决策边界附近时,经验阈值可能导致错误标签被静默复用,从而造成不可控的分类精度下降

Benefits of technology

本申请采用高准确率主分类模型和满足Lipschitz约束的守卫模型的双模型结构,在不修改主分类模型内部结构的前提下,将守卫模型作为边缘侧图像分类的缓存命中判断分支;通过Lipschitz约束、分类间隔和谱范数计算每个缓存样本独立的安全复用半径,使缓存命中判断从经验阈值转化为具有明确几何边界的样本级复用规则,提高了缓存复用决策的可靠性和可解释性。此外,所提出的面向边缘侧图像分类的可认证缓存推理方案,在利用特征缓存集合实现前置推理的同时,通过未命中回退至高准确率主分类模型,保证系统在无法安全复用时仍能获得可靠推理结果,并通过特征缓存减少主分类模型调用次数,降低边缘侧或边缘云协同场景下的推理延迟、网络压力和计算开销。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821233A_ABST
    Figure CN122821233A_ABST
Patent Text Reader

Abstract

The application discloses a feature cache set construction method and an authenticable cache reasoning method and device for edge side image classification, relates to the field of computer recognition, and is used for supporting high-precision edge side cache reasoning. The application obtains a to-be-deployed main classification model, reference data set, label set and cache deployment information, constructs a guard model and trains the same, obtains a spectral norm of an affine classification head, calculates low-dimensional features of each sample data, guard model prediction labels, cache load labels and classification intervals, calculates a safe reuse radius of each sample data based on the classification intervals and the spectral norm, forms a candidate cache item, screens out cache items to obtain a feature cache set, performs matching query on a query image in the set, and according to a query result, reuses a cache load label or falls back to the main classification model for classification identification. The application can reduce the number of times of calling a high-accuracy main classification model and improve the reliability and interpretability of cache reuse decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision, image classification and edge intelligent inference technology, and in particular to a method for constructing a feature cache set for authenticable pre-inference for edge-side image classification, an authenticable cache inference method for edge-side image classification, and an authenticable cache inference device for edge-side image classification. Background Technology

[0002] With the development of edge intelligence, more and more image classification services are being deployed on edge devices such as cameras, industrial inspection terminals, mobile devices, and unmanned systems. These edge devices typically need to classify and recognize continuous image streams in a short period of time. However, high-accuracy deep neural network models often have high computational, memory usage, and energy consumption requirements. Running them directly on resource-constrained edge devices can lead to high latency, low throughput, or excessive energy consumption.

[0003] One known solution is to use edge acquisition and cloud inference, or edge server inference. This involves edge devices acquiring images and sending them to a remote, high-performance server, where a high-accuracy master classification model is run for inference. However, in high-throughput scenarios such as industrial sites, park monitoring, mobile sensing, and augmented reality, numerous sensors may continuously generate similar or repetitive image queries. If each query image calls a remote master classification model, it will not only put pressure on network transmission and server load but also increase end-to-end latency.

[0004] Caching-assisted inference can leverage the visual repetition present in edge workloads to store historical inference results in a cache and reuse the cached results when similar queries arrive, thereby reducing the number of calls to the main classification model.

[0005] However, traditional exact match caching requires completely consistent input and struggles to handle compression, cropping, lighting variations, or minor changes in viewpoint. While semantic caching expands the reusable range, it often relies on empirical similarity thresholds or global distance thresholds, lacking interpretable safety boundaries. When a query sample is near the classification decision boundary, empirical thresholds may cause incorrect labels to be silently reused, resulting in uncontrollable degradation of classification accuracy.

[0006] Therefore, there is an urgent need for a supporting scheme or solution that can support fast caching inference at the edge without modifying the internal structure of the existing high-accuracy main classification model, so as to support high-precision edge-side caching inference or achieve high-precision edge-side caching inference. Summary of the Invention

[0007] The purpose of this invention is to provide a method for constructing a feature cache set to support authenticable high-precision cache inference for edge-side image classification, addressing all or part of the aforementioned problems; and to provide an authenticable cache inference method and apparatus for edge-side image classification, thereby improving the reliability and interpretability of cache reuse decisions while reducing the number of calls to the high-accuracy main classification model.

[0008] The technical solution adopted in this invention is as follows: A method for constructing a feature cache set, the feature cache set being used for authenticable pre-inference for edge-side image classification; the construction method includes: S1. Obtain the main classification model to be deployed, the reference dataset, the label set, and the cache deployment information; the cache deployment information includes the cache budget and cache selection strategy. S2. Construct a guard model, train the guard model using the main classification model and the reference dataset to obtain a feature mapping network and an affine classification head that satisfy the Lipschitz constraint; the feature mapping network calculates the low-dimensional features of the input data, and the affine classification head calculates the predicted label of the guard model from the low-dimensional features. S3. Obtain the spectral norm of the affine classification head; S4. Traverse the sample data of the reference dataset and calculate the low-dimensional features, guard model predicted label, cache load label and classification margin for each sample data. S5. Based on the classification interval and the spectral norm, calculate the safe reuse radius of each sample data to form a candidate cache item. The candidate cache item includes low-dimensional features, cache load label and safe reuse radius. S6. According to the cache selection strategy, under the cache budget constraint, filter cache items from the candidate cache items to obtain a feature cache set.

[0009] Optionally, training the guard model using the main classification model and the reference dataset includes: Freeze the main classification model as the teacher model; The sample data in the reference dataset are input into the main classification model and the guard model respectively to obtain the first logits and the guard model logits, as well as the low-dimensional features extracted by the guard model. Calculate the cross-entropy loss of the guard model based on the cache load label of the sample data; Based on the first logits and the guard model logits, calculate the distillation loss; Based on the claimed guard model logits, calculate the boundary margin loss; The inter-class margin loss is calculated based on the margin between the predicted class and the competing class in the guard model logits. The total loss is obtained by weighted summation of at least one of the cross-entropy loss, distillation loss, boundary margin loss, class center compaction loss, and inter-class margin loss. The total loss is used to update the model parameters of the guard model.

[0010] Optionally, the convolutional layers of the feature mapping network of the guard model are normalized using operator norm; the linear layers of the affine classification head of the guard model are normalized using spectral normalization.

[0011] Optionally, the cache deployment information may also include radius calculation rules; Based on the classification interval and the spectral norm, the safe reuse radius of each sample data is calculated, including: Based on the radius calculation rules, at least one of the conservative radius calculation method and the class-by-class compact radius calculation method is used to calculate the safe reuse radius.

[0012] Optionally, the method for calculating the safe reuse radius using the conservative radius calculation method is as follows: ; in, Representing sample data The safe reuse radius; For sample data The classification interval; This represents the spectral norm of the affine classification head.

[0013] Optionally, methods for calculating the safe reuse radius using the class-by-class compact radius calculation method include: For each competitive category, calculate the difference between the output logit corresponding to the guard model's predicted label and the output logit corresponding to the competitive category, and divide it by the L2 norm of the difference between the weight vector corresponding to the guard model's predicted label and the weight vector corresponding to the competitive category to obtain the boundary distance of the corresponding competitive category. The minimum value among all the boundary distances is selected as the safe reuse radius.

[0014] Optionally, the cache selection strategy includes: Cache items are selected based on the principle of prioritizing secure reuse radius; Alternatively, at least one of the following can be used to select cache items: the safety reuse radius priority principle, the maximum coverage principle, the clustering screening principle, and the feature redundancy removal screening principle.

[0015] Optionally, after forming candidate cache items, the following also includes: For each sample data, determine whether the guard model prediction label of the guard model is consistent with the cache load label of the sample data; if they are consistent, retain the candidate cache item corresponding to the sample data; if they are inconsistent, discard the candidate cache item corresponding to the sample data.

[0016] In a second aspect, this application also provides an authenticable cached inference method for edge-side image classification, the inference method comprising: Obtain the feature cache set constructed using the above-described feature cache set construction method; that is, the inference method can directly reuse the above-described S1~S6 stages. S7. Receive the query image, extract the low-dimensional features of the query image using the guard model as query features, and perform matching queries with each cached item in the feature cache set. S8. Based on the query results, determine whether the query feature falls into the safe reuse area of ​​the hit cache item; if it falls into the area, reuse the cache load label of the hit cache item as the classification result; if it does not fall into the area, fall back to the main classification model for classification and recognition.

[0017] In a third aspect, this application also provides an authenticable cached inference apparatus for edge-side image classification, the inference apparatus including a processor and a storage medium; the storage medium stores a computer program, and the processor runs the computer program to execute the above-described authenticable cached inference method for edge-side image classification.

[0018] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: This application employs a dual-model structure: a high-accuracy main classification model and a guard model satisfying Lipschitz constraints. Without modifying the internal structure of the main classification model, the guard model serves as the cache hit determination branch for edge-side image classification. By calculating the independent safe reuse radius for each cached sample through Lipschitz constraints, classification margin, and spectral norm, the cache hit determination is transformed from an empirical threshold into a sample-level reuse rule with clear geometric boundaries, improving the reliability and interpretability of cache reuse decisions. Furthermore, the proposed authenticable cache inference scheme for edge-side image classification utilizes a feature cache set for pre-inference while simultaneously backtracking to the high-accuracy main classification model upon a miss. This ensures reliable inference results even when safe reuse is not possible. The feature cache also reduces the number of main classification model calls, lowering inference latency, network pressure, and computational overhead in edge-side or edge-cloud collaborative scenarios. Attached Figure Description

[0019] The present invention will be described by way of example and with reference to the accompanying drawings, wherein: Figure 1 This is a flowchart of the multi-stage implementation of the feature cache set construction method.

[0020] Figure 2 This is a network construction diagram of the guard model in one embodiment.

[0021] Figure 3 This is a flowchart of the guard model training process.

[0022] Figure 4 This is a flowchart illustrating the implementation of an authentic cached inference method for edge-side image classification.

[0023] Figure 5 This is an online inference flowchart for the query image.

[0024] Figure 6 This is a diagram illustrating the construction of an authenticated cached inference device for edge-side image classification. Detailed Implementation

[0025] All features disclosed in this specification, or all steps in all disclosed methods or processes, may be combined in any way, except for mutually exclusive features and / or steps.

[0026] Any feature disclosed in this specification (including any appended claims and abstract) may be replaced by other equivalent or similar features, unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is merely one example of a series of equivalent or similar features.

[0027] This application provides an authenticable cached inference method for edge-side image classification, which is used to perform cached inference classification and recognition on query images.

[0028] Parameter explanation: logits: Logical values, referring to the raw scores calculated directly for each class in the last layer of a multi-class classification model, without softmax or sigmoid processing. Each row of logits represents the raw score for a class. Although not normalized, a higher logit value indicates that the model considers that class to be the correct prediction.

[0029] like Figure 1 As shown, in one optional implementation, the reasoning method includes the following multiple stages: S1. Obtain the high-accuracy main classification model to be deployed, the reference dataset, the label set, and cache deployment information.

[0030] (High accuracy) main classification model with MThis indicates that it is a pre-trained image classifier model capable of providing reliable classification results. While ensuring accurate and reliable classification results, any model architecture can be selected for training. For example, the main classification model M can employ high-precision image classification models such as residual networks, convolutional neural networks, or visual Transformers. The main classification model M processes the input data... The first logits calculated are represented as ,based on The first predicted label is obtained by taking the largest output vector for one category. Main classification model M It can be deployed on edge servers, cloud servers, or high-performance local devices. However, its parameter scale is relatively large compared to the lightweight model, and its operating load and computing power requirements are high, making it inconvenient to deploy on edge devices.

[0031] Reference dataset with It indicates that it contains multiple sample data (image data), represented as , i This serves as an index for sample data. It can be the training set, validation set, historical query sample set, or a representative sample set that has been manually labeled.

[0032] The label set, denoted by Y, contains all categories for image classification; that is, it is the set of cached labels for each sample data, represented as... , ,in, Indicates the total number of categories. Indicates the first Category labels, This represents the center of the k-th category.

[0033] Cache deployment information constrains the parameters or conditions of the cache. It includes at least one of the following: (1) Cache budget That is, the maximum number of cached items.

[0034] (2) Feature Dimension This refers to the dimension of the image features. If the model structure is already determined, this parameter does not need to be specified separately.

[0035] (3) Radius Calculation Rules This parameter is used to indicate / determine the calculation method for the safe reuse radius. In one alternative implementation, the calculation method for the safe reuse radius includes a conservative radius calculation method and a class-by-class compact radius calculation method. This parameter is not required if there is only one rule, or if only one rule is considered.

[0036] (4) Cache selection strategy This is used to indicate a method for filtering cache items from candidate cache items. In one optional implementation, cache items are filtered based on a safety reuse radius priority principle; in another optional implementation, at least one of the following is used to filter cache items: a safety reuse radius priority principle, a maximum coverage principle, a clustering filtering principle, and a feature redundancy removal filtering principle. For example, based on the safety reuse radius priority principle, cache items are further filtered by combining the maximum coverage principle / clustering filtering principle (either of the latter two principles).

[0037] (5) Source of cache load tags. Cache load tags are the category tags returned when a cached item is hit in the cache. The cache load label source is used to indicate the cache load label. The source. In an alternative implementation, the cache load labels are derived from the classification results of the main classification model M; for example, for sample data. The first logits calculated by the main classification model M are The first predicted label obtained from this That is, the sample data. Cache load tags In another alternative implementation, the cached payload label is derived from the actual label (classification label) of the sample data, that is, for the sample data... Its cache load label Set its true label. This parameter is not needed if there is only one source, or if only one source is considered.

[0038] (6) Whether to enable central consistency constraints. This indicates whether to validate the guard model's predicted labels (in words). (This indicates the cache load label returned from the cached item when the cache is hit) Consistency. This can be indicated by a central consistency check flag. For the sake of solely considering authenticated cache inference, this parameter is optional.

[0039] (7) Enable write-back on misses. This indicates whether to write the query image to the cache when a cache item is missed (when the cache item set is empty). This parameter is not needed for the construction of the feature cache set. This parameter is optional for the purpose of authentication cache inference.

[0040] S2. Construct a lightweight guard model, train the guard model using the master classification model and the reference dataset, and obtain a feature mapping network and an affine classification head that satisfy the Lipschitz constraint.

[0041] The lightweight guard model GuardNet is denoted by G. Compared to the high-accuracy main classification model M, this guard model G has the characteristics of smaller parameter size and lower computational complexity, making it suitable for deployment on edge devices to perform fast pre-judgment on query images and provide authentication basis for subsequent cache reuse.

[0042] In a preferred embodiment, the guard model G employs a GS2Pixel lightweight convolutional network structure, such as... Figure 2 As shown, the GS2Pixel lightweight convolutional network structure includes, in sequence, a first convolutional feature extraction unit (containing a first convolutional layer and a first activation layer), a first information-preserving downsampling and channel mixing unit (containing a first spatial rearrangement downsampling layer, a first channel mixing layer, and a second activation layer), a second convolutional feature extraction unit (containing a second convolutional layer and a third activation layer), a second information-preserving downsampling and channel mixing unit (containing a second spatial rearrangement downsampling layer, a second channel mixing layer, and a fourth activation layer), a third convolutional feature extraction unit (containing a third convolutional layer, a fifth activation layer, and a flattening layer), a low-dimensional feature mapping layer, and an affine classification head.

[0043] Specifically, the input image first undergoes initial feature extraction through a first convolutional layer and is processed by the GroupSort-2 activation function. Then, the local spatial information is rearranged to the channel dimension using the PixelUnshuffle(2) spatial rearrangement downsampling operator, and the rearranged channel information is compressed and fused through a 1×1 convolutional mixing layer, thus forming the first information-preserving downsampling and channel mixing unit. Afterward, the features sequentially pass through a second convolutional layer, a second GroupSort-2 activation function, a second PixelUnshuffle(2) spatial rearrangement downsampling operator, and a second 1×1 convolutional mixing layer to further extract features and retain local discriminative information as much as possible while reducing spatial resolution. The processed features are then input into a third convolutional layer for deep feature extraction, flattened after processing by the GroupSort-2 activation function, and input into a low-dimensional feature mapping layer to obtain a low-dimensional feature vector used for cache matching and secure reuse determination. Finally, the affine classification head outputs the classification logits corresponding to the guard model.

[0044] The PixelUnshuffle(2) spatial rearrangement downsampling operator is used to map adjacent local pixel blocks to the channel dimension in a rearranged manner during the downsampling process. Compared with average pooling or max pooling, it can retain local pixel information more completely while reducing the feature map spatial size. The 1×1 convolutional mixing layer is used to compress information and linearly mix the channels after rearrangement. The GroupSort-2 activation function is a non-expanding activation function, which is beneficial to satisfy the Lipschitz constraint requirements of the guard model. Thus, the GS2Pixel lightweight convolutional network structure can extract low-dimensional discriminative features suitable for cache reuse authentication while maintaining a low parameter scale and computational complexity.

[0045] In this embodiment, the guard model can be represented as a concatenation of a feature mapping network and an affine classification head, wherein the feature mapping network is denoted as... Affine classification head is denoted as For the input sample data Its low-dimensional features Represented as: The guard model's output logits are: ; Where W represents the affine classification head The weight matrix, where b represents the bias term. Representing sample data The corresponding guard model logits vector. Based on the class corresponding to the largest component in the guard model logits, the sample data is obtained. The guard model predicts the label, denoted as .

[0046] In one alternative implementation, such as Figure 3 As shown, during the training phase, the main classification model M is frozen as the teacher model, the guard model G is used as the student model, and the guard model is trained using a reference dataset or a subset thereof.

[0047] Specifically, the sample data from the reference dataset are input into the main classification model M and the guard model G, respectively, to obtain the first logits and the guard model logits. Based on the true labels of the sample data, the cross-entropy loss of the guard model is calculated. Optionally, the distillation loss is calculated based on the first logits and the guard model logits. Optionally, a boundary margin loss is constructed based on the difference between the highest and second-highest class scores in the guard model's logits. Optionally, the inter-class margin loss can also be calculated based on the guard model's logits. And / or calculate the class center compaction loss based on the distance between low-dimensional features and class centers. Here, the class center refers to the central vector representing the overall distribution position of samples of the corresponding class in the low-dimensional feature space of the guardian model. The class center can be initialized based on the mean of the low-dimensional features of each class sample, and / or jointly updated as a learnable parameter during training.

[0048] The inter-class margin loss is calculated based on the logits of the guard model. Let the sample data... The guard model logits vector obtained from the guard model is: C, as defined above, represents the total number of categories, and the guard model predicts the label as... For samples that are correctly predicted, i.e., satisfying... For the sample, calculate the difference between the guard model logit corresponding to the predicted label and the guard model logit corresponding to each outlier category. The difference between the correct-classified samples and all their out-of-class categories is averaged and then negatively incremented to obtain the inter-class margin loss. This loss term increases the guard model logit margin of the correct class relative to all out-of-class categories, thereby enhancing the separability between different classes.

[0049] The class center compaction loss is calculated based on the distance between low-dimensional features and class centers. Let the sample data... The low-dimensional features obtained by the feature mapping network are ,category The corresponding category center is For sample data Take the category center corresponding to its true category. And based on low-dimensional features With Category Center The square Euclidean distance between them forms the loss term, which is preferably constructed by the method of constructing the loss term. The batch average form is used to obtain the class center compaction loss. This loss term is used to constrain samples of the same class to cluster towards their respective class centers in the low-dimensional feature space, thereby improving intra-class compactness.

[0050] The total loss is obtained by weighting and summing at least one of the following: cross-entropy loss, distillation loss, boundary margin loss, inter-class margin loss, and class center compaction loss. And using the total loss... Update the model parameters of the guard model.

[0051] Among them, cross-entropy loss Used to ensure the guard model has basic classification capabilities; distillation loss Used to enable the guard model G to learn the output distribution of the main classification model M; boundary margin loss. The guard model is trained using a reference dataset to increase the margin between the highest and second-highest class scores of correctly classified samples, thereby facilitating the subsequent increase of the sample-level safe reuse radius; the inter-class margin loss is used to enhance the separation between different classes; and the class center compaction loss is used to improve the clustering of similar samples in the low-dimensional feature space. Based on the set maximum number of iterations, the guard model is trained until the model converges or the maximum number of iterations is reached.

[0052] In addition, to provide a verifiable theoretical basis for the subsequent calculation of the safe reuse radius, Lipschitz constraints are applied to the guard model during the forward computation process corresponding to the training and inference of the guard model.

[0053] In one optional implementation, the convolutional layers in the feature mapping network of the guard model are normalized using operator norm, while the linear layers in the low-dimensional feature mapping layer and the affine classification head are normalized using spectral normalization. Simultaneously, a non-expanding activation function and a downsampling operator are used to ensure that the overall Lipschitz constant of the guard model is no greater than 1 or no greater than a preset threshold. In this embodiment, the non-expanding activation function is preferably the GroupSort-2 activation function, and the downsampling operator is preferably a spatial rearrangement downsampling operator.

[0054] The training and parameter constraint process of the guard model can be summarized as follows: Initialize the lightweight guard model; freeze the main classification model M; for each training batch, input the image into the main classification model and the guard model respectively to obtain the teacher logits (i.e., the first logits), low-dimensional features, and guard model logits; calculate the cross-entropy loss based on the true labels, calculate the distillation loss based on the teacher logits and guard model logits, and optionally calculate the boundary margin loss, inter-class margin loss, and class center compaction loss; perform a weighted summation of each loss term to obtain the total loss; update the guard model parameters according to the total loss; and perform operator norm normalization and spectral normalization in the forward computation of the convolutional layer and the linear layer respectively, finally obtaining the feature mapping network and affine classification head that satisfy the Lipschitz constraint. The training and parameter constraint method of the guard model is shown in Algorithm 1.

[0055]

[0056] S3. Obtain the spectral norm of the affine classification head. Determine the radius calculation rules (if applicable).

[0057] After training the guard model, obtain the affine classification head. weight matrix (From the affine classification head) Classification transformation (obtained from) the weight matrix. spectral norm Safe reuse radius From the weight matrix spectral norm The local classification interval of the sample data is determined.

[0058] When feature mapping network Affine classification head when Lipschitz constraints are satisfied spectral norm It can be used to limit the rate of change of the logits of the guard model in the feature space. Spectral norm. The larger the value, the higher the affine classification head. The more sensitive the feature is to perturbations, the smaller the safe reuse radius should be used for the same classification interval. A larger classification margin indicates that the sample data is further away from the decision boundary, and a larger safe reuse radius can be used under the same spectral norm. .

[0059] S4, Traverse the reference dataset Sample data in Calculate each sample data low-dimensional features Guardian model predicts labels Cache load tags and classification interval .

[0060] For the reference dataset Each sample data in Utilizing the feature mapping network of the guard model Low-dimensional features of calculator And calculate the guard model logits The class with the highest score in the guard model's logits is used as the guard model's predicted label. .

[0061] Cache load tags It refers to the category label returned when a cached item is hit during a cache query. As mentioned earlier, its source can be the first predicted label from the inference of the main classification model M. It can also be its actual label. Once the source of the cache load label is determined, it can be uniquely identified.

[0062] Classification interval This is used to measure how close a sample data point is to the decision boundary in the classification space of the guard model. For sample data... The calculated guard model logits, in order to To indicate the category of the largest logit, use... This indicates the largest logit, with To indicate the category of the second largest logit, use... This indicates the second largest logit. Category interval. The difference between the first and second largest logit, i.e. .

[0063] In one alternative implementation, for each sample data, the guard model's predicted label can be further determined. Is it related to cache load tag? If they are consistent, the corresponding sample data can be discarded in the subsequent candidate cache item selection stage.

[0064] S5, Based on classification interval Spectral norm Calculate each sample data (Sample-level) secure reuse radius This forms candidate cache entries.

[0065] As mentioned above, in some optional implementations, radius calculation rules are configured in the cache deployment information. This constrains the method for calculating the safe reuse radius. The radius calculation rule may include a conservative radius calculation method or a class-by-class compact radius calculation method.

[0066] For calculating sample data using the conservative radius calculation method safe reuse radius In one embodiment, the calculation includes: .

[0067] The reason for using the conservative radius calculation method is that: if the query feature (Low-dimensional features of the image to be classified) and sample data low-dimensional features The 2-norm distance satisfies , This indicates that calculating the L2 norm, under the Lipschitz constraint of the affine classification head, involves querying features. The corresponding guard model's logits change will not exceed the classification margin. The allowed safety range ensures that the guard model can predict labels. Within this local area, the decision boundary is not crossed.

[0068] For calculating sample data using the class-by-class compact radius calculation method safe reuse radius In one embodiment, the calculation method includes: Let the guard model predict the label as Weight matrix The Middle c The row vector corresponding to the class is denoted as For each competitive category (i.e., the category other than the label predicted by the guard model), calculate the sample data separately. To the corresponding competitive category j Boundary distance : ; In the formula, This represents the predicted label in the guard model's logits. The output logit, Weight matrix Medium category The corresponding row vector; In the guard model's logits, the corresponding competition category The output logit, Weight matrix Competition Category The corresponding row vector.

[0069] Finally, for all competitive categories Calculated boundary distance ,Pick The minimum safe reuse radius is selected. As a safe reuse radius.

[0070] Compared to the conservative radius calculation method, the class-by-class tight radius calculation method calculates the boundary distance for each competing category separately, which can obtain a reuse area that is no less than the conservative radius.

[0071] After calculating the safe reuse radius of each sample data, candidate cache items can be constructed accordingly. In one optional implementation, for each sample data... If its safe reuse radius is greater than 0, then the sample data can be reused. Construct candidate cache entries. Candidate cache entries must include at least: low-dimensional features corresponding to the sample data. Cache load tags and safe reuse radius Optionally, it may also include guard model prediction labels. Main classification model predicts labels Authentic Labels At least one piece of auxiliary information.

[0072] In an alternative implementation, a central consistency constraint can also be enabled to perform consistency filtering on candidate cache items. If the guard model predicts the label... With cache load tags If there is an inconsistency, it indicates that the local consistency certificate of the candidate cache item is not aligned with its cached return label, and it can be discarded in subsequent filtering processes. Furthermore, in a more conservative implementation, the guard model's predicted labels can be retained with priority. Main classification model predicts labels and real labels Candidate cache entries with consistent sample data are used to improve the reliability of cache reuse.

[0073] By calculating a safe reuse distance for each sample data point, unlike methods that determine whether to reuse cached payload labels using a fixed global distance threshold, this approach calculates a reusable radius for each sample data point to determine whether its cached payload label should be reused. For sample data with a large classification margin, a larger reuse area is allowed; while for sample data near the classification boundary or with a small classification margin, a smaller reuse area is used or the sample data is discarded directly, thereby reducing the risk of silent misclassification near the boundary.

[0074] It should be noted that the secure reuse radius directly authenticates the local label consistency on the GuardNet side. When the cached payload label matches the GuardNet predicted label, the cached item's returned label is directly aligned with this local consistency certificate. When the cached payload label comes from the main classification model or the true label, the reliability of the cached payload label can be improved through training alignment, central consistency verification filtering, and experimental evaluation.

[0075] S6. Based on the cache selection strategy of the cache deployment parameters, and under the cache budget constraint, filter the cache items from the candidate cache items to obtain the feature cache set.

[0076] Each cached item in the feature cache set must include at least low-dimensional features. Cache load tags and safe reuse radius .

[0077] Based on the cache budget configured in the cache deployment information and caching selection strategy The specified filtering method selects items from the candidate cache items that do not exceed the cache budget. Number of cached items.

[0078] For the method of constructing the feature cache set LipCache, please refer to Algorithm 2.

[0079]

[0080] In one alternative implementation, based on the security reuse radius Size, prioritize screening for safe reuse radius Larger candidate cache entries. That is, from all candidate cache entries, select according to the safe reuse radius. Filter by size from largest to smallest, without exceeding the cache budget. Number of candidate cache items.

[0081] In another alternative implementation, within the safe reuse radius Building upon this foundation, a feature space diversity constraint is introduced for candidate cache item selection. For example, a maximum coverage principle (maximum low-dimensional feature coverage) or a clustering selection principle (adaptive feature clustering of all candidate cache items, followed by selection of a predetermined proportion / number of candidate cache items within each cluster) can be employed. This ensures that the selected candidate cache items are distributed across different feature regions, thereby improving cache coverage. A feature redundancy removal selection principle can be further implemented: candidate cache items are selected sequentially from largest to smallest based on their safe reuse radius. The low-dimensional feature distance or feature similarity between the current candidate cache item and the selected cache items is calculated. If the distance is less than a preset distance threshold or the similarity is greater than a preset similarity threshold, the item is considered redundant and discarded; otherwise, it is added to the feature cache set. For multiple candidate cache items deemed redundant, priority is given to retaining those with larger safe reuse radii. If the safe reuse radii are the same, those with larger classification intervals or higher guard model confidence are retained to avoid excessive concentration of cache items in local feature regions and to improve cache coverage and cache resource utilization efficiency.

[0082] The above S1-S6 stages form a method for constructing a feature cache set. Using this method, an authenticable pre-inference feature cache set for edge-side image classification can be constructed. The guard model is deployed on edge devices, receives query images, extracts their low-dimensional features, and performs fast pre-queries in the feature cache set.

[0083] In the authenticable cached inference method for edge-side image classification, based on the feature cache set constructed using the aforementioned feature cache set construction method, such as Figure 4 As shown, it also includes the following stages: S7. Receive the query image, use the guard model to extract the low-dimensional features of the query image as query features, and perform matching queries with each cached item in the feature cache set.

[0084] The query feature matching query in the cached items (set) calculates the similarity between the query feature and the low-dimensional features of each cached item to determine whether a cache hit occurs. If a cached item is matched, it indicates a cache hit; otherwise, it indicates a cache hit. If the cached item set is empty, it means that no cached item was hit.

[0085] like Figure 5 As shown, the edge device receives the query image. Utilizing the feature mapping network of the deployed guard model Calculate its query features Calculate the query features. With each cache item in the feature cache set low-dimensional features L2 distance between .

[0086] For any cache item ,like If the query feature falls within the safe reuse region of the cached item, it means the cached item has been hit, and the cached item is added to the hit set. (Includes all matched cached items). This hit set It can be represented as .

[0087] By comparing the secure reuse distance with adaptive sample data, if ,For example (Similarly for class-wise tight radius), the change in the output of the affine classification head is insufficient to reverse the logit difference between the guard model's predicted label and any competing class. Therefore, the guard model's predicted label remains unchanged in this local area, which can ensure the reliability of label reuse discrimination based on the safe reuse distance of each sample data.

[0088] S8. Based on the query results, determine whether to reuse the cache load label of the hit cache item or fall back to the main classification model for classification and identification.

[0089] Specifically, based on the query results, it is determined whether the query features fall within the safe reuse area of ​​the cached item; if they do, the cache load label of the cached item is reused as the classification result; if they do not fall within, it is reverted to the main classification model for classification and recognition.

[0090] If a cache item is hit, that is, if the set is hit. If not empty, then collect from the set. The cache load label of the cached item with the closest similarity distance is selected as the classification result for the query image. The selected cached item is represented as... The cache load label for this cached item is represented as ,Should This is the output classification result. If no cached item is hit, it means the set has been hit. If the value is empty, it means that the current query image does not fall within the safe reuse area of ​​any cached item. In this case, the main classification model M is called to perform a complete category identification on the query image and calculate the first logits. Get category tags .

[0091] Furthermore, when a cached item is missed, if the cache deployment information is configured with a parameter indicating whether to enable miss write-back, and this parameter indicates an implementation that enables miss write-back, then the query image is calculated using the same method. The safe reuse radius of the query image Query features Safe reuse radius and classification labels Write it into the feature cache set according to the cache item construction format.

[0092] For the method of authenticating cached inference for query images, please refer to the implementation of Algorithm 3.

[0093]

[0094] Since the lightweight guard model only handles feature extraction and cache item hit determination, and the high-accuracy main classification model remains the final fallback model in case of a miss, this application can reduce the number of calls to the main classification model while maintaining high classification accuracy. For query images with significant classification intervals and falling into the adaptive safe reuse region, the system can directly return cached load labels, thereby reducing inference latency; for query samples that do not meet the safe reuse conditions, the system falls back to the main classification model to avoid excessive reuse caused by empirical thresholds.

[0095] Based on the ideas of this application, this application also proposes an authenticated cached inference device for edge-side image classification, such as... Figure 6 As shown, the inference device includes a processor and a storage medium, which can be one or more combinations of read-only memory, random access memory, flash memory, hard disk, solid-state drive, or other media capable of storing computer programs. The storage medium stores a computer program, which can be stored in a single storage medium or in fragments across multiple storage media. The processor, running the computer program, can independently execute the authenticated cached inference method for edge-side image classification according to any of the above embodiments. The processor may include a single processing unit or consist of multiple processing units. For example, a processing unit deployed on an edge device, or a processing unit deployed on an edge device, an edge server, or a cloud server.

[0096] To verify the effectiveness of the proposed solution, effectiveness tests were also conducted on the full CIFAR-10 test set in the embodiments of this application. The test results are shown in Table 1. Based on the experimental results of the full CIFAR-10 test set, the comprehensive performance of the authenticable cached inference scheme proposed in this invention in terms of cache hit rate, end-to-end classification accuracy, authentication reliability, and edge-side inference acceleration effect can be further illustrated.

[0097] Table 1. CAFAR-10 Full Quantity Test Table

[0098] Furthermore, in real-world edge deployment experiments, the proposed solution still achieved a cache hit rate of 45.8% and an end-to-end classification accuracy of 89.8%, while achieving approximately 1.70 times faster end-to-end inference, demonstrating that the proposed solution not only has theoretical verifiability but also good engineering deployment value.

[0099] In summary, the feature cache set construction method based on LipCache for edge-side image classification proposed in this application can significantly reduce the frequency of high-accuracy main classification model calls while maintaining high classification accuracy, cache hit rate and edge-side inference efficiency, thus having good theoretical value and practical application prospects.

[0100] Furthermore, compared to empirical cache reuse schemes using global nearest neighbor matching, although such empirical schemes can achieve higher apparent cache hit rates under certain configurations, their end-to-end classification accuracy is significantly reduced, and the self-consistent authentication reuse rate is only 17.10%, posing a significant risk of erroneous reuse. In contrast, the scheme in this application adopts more prudent theoretical constraints on hit rate, but provides clear theoretical authentication boundaries and achieves a better balance between classification accuracy, reuse reliability, and actual deployment availability.

[0101] As shown in Table 1, under the conditions of using ResNet50 as the main classification model, GS2Pixel lightweight architecture as the guard model, high-priority sample selection after deduplication as the cache selection strategy, and a cache capacity of 300 per class, the proposed solution can achieve a cache hit rate of approximately 44%, an end-to-end classification accuracy of approximately 90%, and a self-consistent authentication reuse rate under the theoretical safety radius constraint. This demonstrates that the proposed solution can stably obtain the performance benefits brought by cache reuse while ensuring reuse security.

[0102] This invention is not limited to the specific embodiments described above. The invention extends to any new feature or combination disclosed in this specification, as well as any new method or process step or combination disclosed herein.

Claims

1. A method for constructing a feature cache set, wherein the feature cache set is used for authenticable pre-inference for edge-side image classification; characterized in that, The construction method includes: S1. Obtain the main classification model to be deployed, the reference dataset, the label set, and the cache deployment information; the cache deployment information includes the cache budget and cache selection strategy. S2. Construct a guard model, train the guard model using the main classification model and the reference dataset to obtain a feature mapping network and an affine classification head that satisfy the Lipschitz constraint; the feature mapping network calculates the low-dimensional features of the input data, and the affine classification head calculates the predicted label of the guard model from the low-dimensional features. S3. Obtain the spectral norm of the affine classification head; S4. Traverse the sample data of the reference dataset and calculate the low-dimensional features, guard model predicted label, cache load label and classification margin for each sample data. S5. Based on the classification interval and the spectral norm, calculate the safe reuse radius of each sample data to form a candidate cache item. The candidate cache item includes low-dimensional features, cache load label and safe reuse radius. S6. According to the cache selection strategy, under the cache budget constraint, filter cache items from the candidate cache items to obtain a feature cache set.

2. The feature cache set construction method as described in claim 1, characterized in that, Training the guard model using the main classification model and the reference dataset includes: Freeze the main classification model as the teacher model; The sample data in the reference dataset are input into the main classification model and the guard model respectively to obtain the first logits and the guard model logits, as well as the low-dimensional features extracted by the guard model. Calculate the cross-entropy loss of the guard model based on the cache load label of the sample data; Based on the first logits and the guard model logits, calculate the distillation loss; The boundary margin loss is calculated based on the guard model logits; The class center compaction loss is calculated based on the distance between the low-dimensional features and the class centers. The total loss is obtained by weighted summation of at least one of the cross-entropy loss, distillation loss, boundary margin loss, class center compaction loss, and inter-class margin loss. The total loss is used to update the model parameters of the guard model.

3. The feature cache set construction method as described in claim 2, characterized in that, The convolutional layers of the feature mapping network of the guard model are normalized using operator norm; the linear layers of the affine classification head of the guard model are normalized using spectral normalization.

4. The feature cache set construction method as described in claim 1, characterized in that, The cache deployment information also includes radius calculation rules; Based on the classification interval and the spectral norm, the safe reuse radius of each sample data is calculated, including: Based on the radius calculation rules, at least one of the conservative radius calculation method and the class-by-class compact radius calculation method is used to calculate the safe reuse radius.

5. The feature cache set construction method as described in claim 4, characterized in that, The method for calculating the safe reuse radius using the conservative radius calculation method is as follows: ; in, Representing sample data The safe reuse radius; For sample data The classification interval; This represents the spectral norm of the affine classification head.

6. The feature cache set construction method as described in claim 4, characterized in that, Methods for calculating the safe reuse radius using the class-by-class compact radius calculation method include: For each competitive category, calculate the difference between the output logit corresponding to the guard model's predicted label and the output logit corresponding to the competitive category, and divide it by the L2 norm of the difference between the weight vector corresponding to the guard model's predicted label and the weight vector corresponding to the competitive category to obtain the boundary distance of the corresponding competitive category. The minimum value among all the boundary distances is selected as the safe reuse radius.

7. The feature cache set construction method as described in claim 1, characterized in that, The cache selection strategy includes: Cache items are selected based on the principle of prioritizing secure reuse radius; Alternatively, at least one of the following principles can be used to select cache items: the principle of prioritizing safe reuse radius, the principle of maximum coverage, the principle of clustering selection, and the principle of feature redundancy removal selection.

8. The method for constructing a feature cache set as described in any one of claims 1-7, characterized in that, After forming candidate cache entries, the following is also included: For each sample data, determine whether the predicted label of the guard model is consistent with the cache load label of the sample data; if they are consistent, retain the corresponding candidate cache item; if they are inconsistent, discard the corresponding candidate cache item.

9. A method for authenticated cached inference for edge-side image classification, characterized in that, include: Obtain the feature cache set constructed using the feature cache set construction method as described in any one of claims 1-8; S7. Receive the query image, extract the low-dimensional features of the query image using the guard model as query features, and perform matching queries with each cached item in the feature cache set. S8. Based on the query results, determine whether the query feature falls into the safe reuse area of ​​the hit cache item; if it falls into the area, reuse the cache load label of the hit cache item as the classification result; if it does not fall into the area, fall back to the main classification model for classification and recognition.

10. An authenticated cached inference apparatus for edge-side image classification, characterized in that, It includes a processor and a storage medium; the storage medium stores a computer program, and the processor runs the computer program to perform the authenticated cached inference method for edge-side image classification as described in claim 9.