Class incremental learning method for edge scene image recognition and electronic equipment
By adopting a class incremental learning method based on width learning in edge scenarios, using Gramma matrix and feature fusion technology, the problem of large memory and computing overhead in edge scenarios is solved, and efficient image recognition performance is achieved.
Patent Information
- Application Number
- CN202510204353.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-24
AI Technical Summary
In the prior art, the class incremental learning method based on deep neural networks cannot effectively support image recognition tasks due to high memory and computing overhead in edge scenarios.
A class incremental learning method based on width learning is adopted, and the input samples are upgraded through the random mapping module, and a new class sample is judged using the Gram matrix, feature fusion and output layer weight fine-tuning are performed to achieve edge scene image recognition.
This method can quickly and accurately perform class incremental learning under low memory and low computing overhead, significantly improving the performance of edge scene image recognition.
Smart Images

Figure CN120125968A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and in particular to a quasi-incremental learning method and electronic device for edge scene image recognition. Background Art
[0002] With the rapid development of information technology and the widespread application of Internet of Things technology, the application of image recognition in edge scenarios is gradually becoming a key force for intelligent upgrading. "Edge scenarios" usually refer to scenarios or applications that occur at the "edge" location in the network architecture. The "edge" here refers to the edge of the network close to the data source or end user, rather than a centralized cloud server. In such a scenario, data processing and analysis will be carried out as close as possible to where the data is generated, which can reduce the delay in data transmission and reduce the burden on the central server. Combining image recognition technology with edge scenarios will not only help improve the speed and efficiency of image processing, but also provide technical support for multiple development strategies such as intelligent monitoring, intelligent manufacturing, and smart cities.
[0003] Image recognition systems are gradually expanding to more complex and diverse application scenarios, and the data challenges they face are also increasing. In edge scenarios, facing the emergence of new types of objects, image recognition systems must have strong real-time analysis and adaptation capabilities. This requires not only that the system can quickly capture and analyze the feature information of new objects, but also that it must accurately learn and recognize new images and retain the ability to recognize previous images. However, most traditional image recognition models often rely on obtaining all training data at one time during the training process, which means that in the initial training stage, the system may not be able to cover all potential image recognition content, which greatly limits the practical application effect of image recognition technology, especially in edge scenarios where data streams often distribute a large number of new types of images. Nowadays, by deploying advanced sensor networks, we can collect a large amount of image data at the edge. How to efficiently use this data to achieve rapid recognition and make accurate judgments, while continuously learning new image features, is a core issue that needs to be solved in the field of image recognition in edge scenarios. Therefore, studying the ability of image recognition of new types of images in edge scenarios is of great significance to improving the environmental adaptability of the system. Introducing incremental learning technology in image recognition to achieve real-time learning and model updating of new class features is crucial to improving image recognition performance in edge computing environments.
[0004] The current mainstream class-incremental learning methods are as follows: 1) Regularization methods, which aim to restrict the key parameters of previous tasks from deviating during the training of new samples, while promoting adaptation to new data during the training process. 2) Sample replay methods, which aim to enable the model to view some old data when training new data. 3) Dynamic network structure adaptation methods, which focus on adapting to new class data by dynamically expanding the network structure. Although these algorithms have shown good results, these algorithms are all based on deep neural networks, which brings a problem that the deep class-incremental learning process requires a large amount of time and memory resources. As mentioned before, the edge scenario is the network edge close to the data source or the end user, rather than the centralized cloud server. Therefore, these deep class-incremental learning technologies are not competent for the image recognition tasks in the edge scenario, because the memory and computing speed of the devices at the edge are limited and often cannot bear the high memory and computing overhead of the deep class-incremental learning method. Therefore, a low-consumption and low-memory class-incremental method for image recognition in the edge scenario and an electronic device are needed. Summary of the Invention
[0005] To at least partly solve one of the technical problems existing in the prior art, an object of the present invention is to provide a width learning-based class-incremental learning method, an electronic device, and a storage medium for image recognition in an edge scenario.
[0006] The first technical solution adopted by the present invention is:
[0007] A class-incremental learning method for image recognition in an edge scenario includes the following steps:
[0008] Collect real-time data collected by sensors in the edge scenario as input samples;
[0009] Perform feature dimensionality increase on the input samples through a random mapping module of width learning, map the features to a higher-dimensional feature representation to filter out irrelevant feature information in the samples, and expand the expression ability of useful feature information;
[0010] Judge whether the input sample belongs to a new class sample or an old class sample according to the Gram matrix. If it is a new class sample, perform amplification processing on the sample label;
[0011] Perform feature fusion on the Gram matrix of the new class sample and the Gram matrix of the previous sample;
[0012] Use the Gram matrix after feature fusion to fine-tune the output layer weights of the width learning model to obtain an edge scenario image recognition model after class-incremental learning.
[0013] Further, the performing feature dimensionality increase on the input samples through a random mapping module of width learning includes:
[0014] The input sample is randomly mapped through width learning, and the mapping result is stored in the feature nodes;
[0015] Through random mapping, the feature nodes are mapped into enhancement nodes.
[0016] Furthermore, the random mapping process from the input sample X to the feature nodes Z is as follows:
[0017]
[0018] where Z i is the i-th group of feature nodes, represents the activation function of the i-th group of feature nodes, W ei is the random mapping weight of the i-th group of feature nodes, β ei is the weight bias of the i-th group of feature nodes;
[0019] The mapping process of the enhancement nodes is as follows:
[0020] H j = ξ j (Z n W hj + β hj )
[0021] where ξ j represents the activation function of the j-th enhancement node, W hj is the random mapping weight of the j-th group of enhancement nodes, β hj is the weight bias of the j-th group of enhancement nodes, Z n is the information of n groups of feature nodes obtained in the random mapping stage of the feature nodes; the feature matrix A of the input sample finally formed by combining the feature nodes and the enhancement nodes is A = [Z n , H n .
[0022] Furthermore, according to the Gram matrix, it is judged whether the input sample belongs to a new class sample or an old class sample. If it is a new class sample, the sample label is amplified, including:
[0023] Pre-store the Gram matrix of each previous known type for the feature identifier representing this type;
[0024] Obtain the Gram matrix of the input sample, and calculate the included angle between the Gram matrix of the input sample and the pre-stored Gram matrix;
[0025] If the included angle is greater than the preset threshold, it is determined that the input sample is a new class sample; otherwise, it is determined that the input sample is an old class sample;
[0026] If it is determined that the input sample is a new class sample, the sample label is amplified.
[0027] Furthermore, the included angle τ k is calculated as follows:
[0028]
[0029] In the formula, is the Gram matrix of the input samples;
[0030] The expression for amplifying the sample labels is:
[0031]
[0032] In the formula, Y n-1 represents the label matrix before seeing the new class samples, M n is the matrix for measuring the similarity between new class samples and old class samples, L n is the new class label information amplification matrix; m represents the sample size of this class incremental task, and n represents the total number of classes.
[0033] Furthermore, the expression for feature fusion is:
[0034]
[0035] In the formula, β n-1 is the Gram matrix of the previous stage; A n represents the features of the new class samples after width learning random mapping, T is the transpose; λ is a hyperparameter, I is the identity matrix, A i is the sample data of the i-th batch entering, and n is the incremental task of the n-th stage.
[0036] Furthermore, fine-tuning the output layer weights of the width learning model using the Gram matrix after feature fusion includes:
[0037] The calculation formula for the output layer weights of the width learning model is:
[0038]
[0039] In the formula, β n is the Gram matrix, Y is the training sample label matrix, and A represents the sample feature matrix of the samples after width learning random mapping;
[0040] Perform Woodbury optimization on the inverse of the Gram matrix;
[0041] Adopt Cholesky decomposition to optimize the inverse expansion result to obtain the Cholesky factor;
[0042] Update the output layer weights of the width learning model according to the obtained Cholesky factor.
[0043] Furthermore, the Woodbury optimization of the inverse of the Gram matrix includes:
[0044] Split the β in the output layer weight calculation process n into Optimize it using the Woodbury expansion formula to obtain the following expansion:
[0045]
[0046] In the formula, is the inverse Gram matrix obtained in the previous stage. Directly save the result in the previous incremental process for update;
[0047] The Cholesky decomposition optimization of the inverse expansion result to obtain the Cholesky factor includes:
[0048] Introduce Cholesky decomposition and decompose into a lower triangular matrix P n-1 , and through Q n =A n P n-1 realize new class feature learning for the lower triangular matrix P n-1 , and substitute it into formula (1) to get:
[0049]
[0050] For in formula (2), perform another Cholesky decomposition to obtain a lower triangular matrix G n-1 , and further obtain:
[0051]
[0052] where P n-1 , G n-1 are both obtained through and A n obtained.
[0053] Furthermore, update the output layer weight through one of the following two algorithms:
[0054] Algorithm 1:
[0055] Algorithm 2:
[0056] Among them, W n represents the weight parameter matrix of the model after the completion of the nth incremental task; M n is a matrix used to measure the similarity between new class samples and old class samples, L nIt is a new class label information amplification matrix;
[0057] Select different algorithms according to the sample size and feature complexity of the input model to achieve faster and more accurate model class incremental learning.
[0058] The second technical solution adopted by the present invention is:
[0059] An electronic device, the electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above-mentioned method for class incremental learning for edge scenario image recognition.
[0060] The third technical solution adopted by the present invention is:
[0061] A computer-readable storage medium, and at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the above-mentioned method for class incremental learning for edge scenario image recognition.
[0062] The fourth technical solution adopted by the present invention is:
[0063] A computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned method for class incremental learning for edge scenario image recognition.
[0064] The beneficial effects of the present invention are: by introducing the Gram matrix operation, the present invention realizes the feature fusion of the new and old class samples after width mapping. In addition, considering the complexity and real-time problems of the data stream, the Woodbury inverse expansion formula is introduced to optimize the inverse calculation of the Gram matrix. At the same time, aiming at the instability existing in the model class incremental process in the data stream, it is proposed that the Cholesky decomposition can be used to replace the inverse process. It can make the model calculation faster and more accurate and the model memory consumption smaller, can greatly improve the model class incremental learning performance, and greatly promote the technical development and practical implementation in the field of edge scenario image recognition. Description of the Drawings
[0065] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings of the relevant technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions of the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0066] Figure 1 is the width learning framework diagram;
[0067] Figure 2 is a schematic diagram of the width learning model architecture and the details of implementing weight parameter fine-tuning in the embodiments of the present invention;
[0068] Figure 3 is the step flow chart of a class incremental learning method for edge scenario image recognition in the embodiments of the present invention. Detailed implementation manners
[0069] The embodiments of the present invention are described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0070] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc., is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.
[0071] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is two or more, greater than, less than, exceeding, etc. are understood as not including the present number, and above, below, within, etc. are understood as including the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.
[0072] In the description of the present invention, unless otherwise clearly defined, terms such as "set", "installed", "connected", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in the present invention in combination with the specific content of the technical solution.
[0073] Technical Explanation:
[0074] (1) Wide Learning System
[0075] As Figure 1 shown, Figure 1 is the framework diagram of wide learning. In this model, the input data is randomly mapped through BLS and the results are stored in the feature nodes. Subsequently, through random mapping again, the feature nodes are mapped to the enhancement nodes. Finally, all the feature nodes and enhancement nodes are combined into a single-layer network model, and the output weights are determined through ridge regression. Specifically, given a batch of training samples and the corresponding training sample label matrix N represents the number of training samples, M represents the number of features of each sample, and C is the number of categories. The random mapping process from the sample data X to the feature nodes Z is as follows:
[0076]
[0077] where represents the activation function of the i-th group of feature nodes, W ei is the random mapping weight of the i-th group of feature nodes, and β ei is the weight bias of the i-th group of feature nodes. Similarly, the mapping process of the enhancement nodes is as follows:
[0078]
[0079] where ξ j represents the activation function of the i-th enhancement node, W hj is the random mapping weight of the j-th group of enhancement nodes, and β hj is the weight bias of the j-th group of enhancement nodes. Finally, the single-layer network structure A formed by the feature nodes and the enhancement nodes is A = [Z n , H n . Therefore, the BLS model can be finally designed as the following optimization problem:
[0080]
[0081] where W is the output layer weight that the model needs to calculate, and λ is a hyperparameter. According to the above optimization problem, the final output layer weight of the BLS model can be calculated by the following formula:
[0082] W = (λI + A T A) -1 AT Y (3)
[0083] where I is an identity matrix. Above is the basic working principle of the width learning system.
[0084] (2) Woodbury Matrix Identity
[0085] The Woodbury Matrix Identity is an important identity for calculating the inverse of a matrix and has important significance in linear algebra and numerical calculations. The Woodbury Matrix Identity is widely used in inverse matrix updates, matrix approximations, calculating the pseudoinverse of a positive semi - definite matrix, and various algorithms to ensure numerical stability and accuracy. Specifically, the calculation form of the Woodbury Matrix Identity is as follows:
[0086] (A + UCV) -1 = A -1 - A -1 U(C -1 + VA -1 U) -1 VA -1 (4)
[0087] where In some matrix inverse operations, when it is necessary to calculate (A + UCV) -1 and A -1 has already been calculated, this identity is very useful. For example, in the least squares method, only (C -1 + VA -1 U) -1 needs to be found. When the matrix dimension of C is much smaller than that of A, the calculation result is more efficient than directly calculating (A + UCV) -1 . Moreover, avoiding the inversion of a larger matrix can further improve the accuracy of the result. At the same time, the Woodbury Matrix Identity can be further optimized into the following formula:
[0088] (A + UCV) -1 = A -1 - A -1 U(I + CVA -1 U) -1 CVA -1 (5)
[0089] This form of the inverse expansion will be more stable.
[0090] (3) Cholesky decomposition and Gram Matrix A
[0091] In linear algebra, the Cholesky decomposition is to decompose a Hermitian positive definite matrix into the product of a lower triangular matrix and its conjugate transpose, which is useful for efficient numerical solutions. The specific form is as follows:
[0092] A = LL T
[0093] where is a positive semi - definite matrix, is a lower triangular matrix. The Cholesky decomposition has a wide range of applications in matrix inversion, least - squares estimation, and numerical calculations. It not only improves the computational efficiency but also provides a concise and stable way to solve practical problems.
[0094] In linear algebra, the Gram matrix is the result of the inner products of a set of vectors <v 1 , v 2 , …, v n . The entries in the Gram matrix are given by the inner product G i,j <v i , v j . The specific representation is as follows:
[0095]
[0096] The Gram matrix, also known as the style transfer matrix, is used to quantify the mutual influence between vectors. This concept has been widely applied in the field of deep learning and can be used to measure the degree of mutual influence of sample features.
[0097] Example 1
[0098] Refer to Figure 2 and Figure 3 , this example provides a class - incremental learning method for edge - scenario image recognition, which specifically includes the following steps:
[0099] S1. Collect the real - time data collected by edge - scenario sensors as input samples; preferably, the real - time data here includes image information of new objects and numerical information of new objects, etc.
[0100] S2. Perform feature dimensionality elevation on the input samples through the random mapping module of width learning, map the features to a higher - dimensional feature representation to filter out irrelevant feature information in the samples and expand the expression ability of useful feature information;
[0101] S3. Judge whether the input sample belongs to a new - class sample or an old - class sample according to the Gram matrix. If it is a new - class sample, perform amplification processing on the sample label;
[0102] S4. Perform feature fusion on the Gram matrix of the new - class sample and the Gram matrix of the previous samples;
[0103] S5. Fine-tune the weights of the output layer of the width learning model using the Gram matrix after feature fusion to obtain an edge scene image recognition model after class incremental learning.
[0104] In this embodiment, first, the sample features of the real-time data collected during the edge scene image recognition process are dimensionally elevated through the random mapping module of width learning. Then, different label matrix amplification strategies are adopted according to whether the input sample is a new class sample or an old class sample. Next, the Gram matrix of the new class sample and the Gram matrix of the previous sample are feature-fused. Then, the weights of the output layer can be fine-tuned using the Gram matrix after feature fusion.
[0105] The above method will be explained in detail below in conjunction with specific embodiments.
[0106] (1) Feature mapping of the input sample
[0107] During the edge scene image recognition process, relying on an advanced sensor array, the tabular numerical information or specific pictures of new class target objects appearing in the detection data stream can be obtained in real time. The information collected by these sensors can be used as input samples and input into the system. As an alternative embodiment, public data sets (such as CIFAR, ImageNet, KEE, etc.) are used to experiment with the performance of the system. Then, in this embodiment, the input sample is first dimensionally elevated using the random feature mapping module of width learning (formulas (1) and (2)). After non-linear transformation, the features are mapped to a higher-dimensional feature representation. This non-linear mapping can filter out the irrelevant feature information in the sample and expand the expression ability of the useful feature information, which helps to achieve more accurate feature fusion subsequently and helps the model obtain better recognition ability.
[0108] (2) Learning of sample label information
[0109] After dimensional elevation of the features, it is necessary to confirm whether the input sample information belongs to a new class or an old class. The system will pre-store the Gram matrix of each previously known class to represent the feature identifier of this class. The Gram matrix of the input sample is used to measure the angle in the feature space with the Gram matrix of the old class. If the spatial angle exceeds a threshold, it is determined to be a new class. The measurement index is as follows:
[0110]
[0111] If it is determined to be a new class, the sample label needs to be amplified. The formula is as follows:
[0112]
[0113] where Yn-1 Denote the label matrix before seeing new-class samples as M n is a matrix used to measure the similarity between new-class samples and old-class samples. Specifically L n is the new-class label information amplification matrix. Specifically where m represents the sample size of this class increment, and n represents the total number of classes.
[0114] A visual explanation of sample label information learning is as follows Figure 2 shown. By constructing a new-class label matrix, the original label space is amplified. Then, a new column of nodes is added in the output layer, and the mapping relationship between new-class labels and new-class sample features is learned by fine-tuning the parameters.
[0115] (3) Feature fusion for new-class samples
[0116] Each batch of samples entering the training has a corresponding Gram information matrix. In the embodiment of the present invention, a global Gram matrix β n is maintained to achieve the fusion of new-class sample features and old-class sample features. The specific feature fusion process is as follows:
[0117]
[0118] where β n-1 is the Gram matrix of the previous stage, which contains all the sample feature information trained by the previous model, and A n represents the features of new-class samples. Through we obtain the Gram matrix of new-class samples, that is, the feature relationship in new-class samples. λ is a hyperparameter, and I is the identity matrix. The feature fusion of new and old-class samples is implemented by adding the β n-1 and matrices.
[0119] (4) Woodbury optimization by taking the inverse of the Gram matrix
[0120] Since the output layer weight update formula W requires taking the inverse of the Gram matrix, which involves the process, the present invention splits the β n in the calculation process into and optimizes it using the Woodbury expansion formula to obtain the following expansion formula:
[0121]
[0122] where is the inverse Gram matrix obtained in the previous stage. Therefore, we do not need to calculate again and can directly save the result in the previous increment process for Update to achieve computational acceleration. Also, in Equation (6), the most computationally expensive part is still the inversion process. However, at this time, the scale of inversion changes from the original to That is, the original computational cost of matrix inversion depends on the total number of feature nodes. At this time, after the Woodbury inversion expansion, the computational cost depends on the number of training samples in this batch. When the number of samples is small and the number of feature nodes is large during each class increment process, Equation (6) can achieve significant computational acceleration.
[0123] (5) Optimize the inversion expansion result using Cholesky decomposition
[0124] In the previous stage, the model saved the result. Introduce Cholesky decomposition to decompose into a lower triangular matrix P n-1 . At the same time, through Q n = A n P n-1 realize the new class feature learning for the lower triangular matrix P n-1 . Substitute these into Equation (6) to get:
[0125]
[0126] Among them, for in Equation (7), another Cholesky decomposition can be performed to obtain a lower triangular matrix G n-1 . It can be further obtained that:
[0127]
[0128] where P n-1 , G n-1 can both be obtained through and A n .
[0129] (6) Update the model weight parameters
[0130] Based on the width-based weight formula: W = (λI + A T A) -1 A T Y, use the Cholesky factors obtained in Step 5 for updating the weights of the model output layer. Through formula derivation, this patent can provide two different algorithms for updating the weights of the output layer.
[0131] Algorithm 1:
[0132] Algorithm 2:
[0133] Selecting different algorithms according to the sample size and feature complexity of the input model can achieve faster and more accurate model class incremental learning.
[0134] In summary, compared with the prior art, the advantages of the present invention lie in expanding the class incremental learning ability of width learning. The Gram matrix operation is introduced to realize the feature information fusion of new and old class samples. At the same time, considering the complexity and real-time nature of the data stream, the Woodbury inversion expansion formula is introduced to optimize the inversion calculation of the Gram matrix. At the same time, aiming at the instability existing in the model class incremental process in the data stream, it is proposed that the Cholesky decomposition can be used to replace the inversion process. Introducing the above three methods into the width learning architecture can make the model calculation faster and more accurate and the model memory consumption smaller, and can greatly improve the model class incremental learning performance. The lightweight class incremental width learning system proposed by the present invention can greatly promote the technical development and practical implementation in the field of edge scenario image recognition.
[0135] Embodiment 2
[0136] The embodiment of the present invention further provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement Figure 3 a class incremental learning method for edge scenario image recognition as shown.
[0137] It can be understood that the memory may include a random access memory (RAM), and may also include a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.
[0138] The processor may include one or more processing cores. The processor connects various parts within the entire server using various interfaces and circuits, and performs various functions of the server and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by invoking data stored in the memory. Optionally, the processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate a combination of one or more of a central processing unit (CPU) and a modem, etc. Among them, the CPU mainly processes the operating system and application programs, etc.; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor and may be implemented separately by a single chip.
[0139] Since this electronic device is an electronic device corresponding to a class incremental learning method for edge scenario image recognition in an embodiment of the present invention, and the principle of how this electronic device solves problems is similar to that of this method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and repeated parts will not be elaborated.
[0140] Embodiment 3
[0141] An embodiment of the present invention further provides a computer-readable storage medium, in which at least one instruction, at least one program segment, a code set, or an instruction set is stored, and the at least one instruction, the at least one program segment, the code set, or the instruction set is loaded and executed by a processor to implement Figure 3 a class incremental learning method for edge scenario image recognition as shown.
[0142] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium, which includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium that can be used to carry or store data.
[0143] Since this storage medium is a storage medium corresponding to a class incremental learning method for edge scenario image recognition in an embodiment of the present invention, and the principle of solving problems by this storage medium is similar to that of this method, the implementation of this storage medium can refer to the implementation process of the above method embodiment, and the repeated parts will not be elaborated.
[0144] Embodiment 4
[0145] In some possible implementation manners, various aspects of the method in the embodiment of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a class incremental learning method for edge scenario image recognition according to various exemplary implementation manners described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0146] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0147] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0148] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered within the protection scope of the present invention.
Claims
1. A quasi-incremental learning method for edge scene image recognition, characterized in that: The following steps are involved: Real-time data collected by sensors in edge scenarios is used as input samples; The input samples are upgraded in dimension through the random mapping module of width learning, and the features are mapped to higher-dimensional feature representations. According to the Gram matrix, determine whether the input sample belongs to a new class sample or an old class sample. If it is a new class sample, amplify the sample label; Perform feature fusion on the Gram matrix of the new class sample and the Gram matrix of the previous sample; The Gram matrix after feature fusion is used to fine-tune the output layer weights of the width learning model to obtain an edge scene image recognition model after incremental learning.
2. The incremental learning method for edge scene image recognition according to claim 1, characterized in that: The random mapping module through width learning performs feature dimension increase on the input sample, including: The input samples are randomly mapped through width learning, and the mapping results are stored in the feature nodes; Through random mapping, feature nodes are mapped to enhanced nodes.
3. The incremental learning method for edge scene image recognition according to claim 2, characterized in that: The random mapping process of input sample X to feature node Z is as follows: In the formula, Z i is the i-th group of feature nodes, represents the activation function of the i-th group of feature nodes, W ei is the random mapping weight of the i-th group of feature nodes, β ei is the weight bias of the i-th group of feature nodes; The mapping process of the enhanced node is as follows: H j =ξ j (Z n W hj +b hj ) In the formula, ξ j represents the activation function of the j-th enhancement node, W hj is the random mapping weight of the j-th group of enhanced nodes, β hj is the weight bias of the jth group of enhanced nodes, Z n It is the n groups of feature node information obtained in the feature node random mapping stage; The feature matrix A of the input sample formed by the combination of the final feature node and the enhanced node is [Z n ,H n ].
4. The incremental learning method for edge scene image recognition according to claim 1, characterized in that: According to the Gram matrix, the input sample is judged to be a new class sample or an old class sample. If it is a new class sample, the sample label is amplified, including: Pre-store the Gram matrix of each previously known type The characteristic identifier used to represent this type; Obtaining a Gram matrix of an input sample, and calculating an angle between the Gram matrix of the input sample and a pre-stored Gram matrix; If the angle is greater than the preset threshold, the input sample is determined to be a new class sample; otherwise, the input sample is determined to be an old class sample; If the input sample is determined to be a new class sample, the sample label is amplified.
5. The incremental learning method for edge scene image recognition according to claim 4, characterized in that: The angle τ k The calculation formula is as follows: In the formula, is the Gram matrix of the input sample; The expression for amplifying sample labels is: Where Y n-1 represents the label matrix before seeing new class samples, M n It is a matrix used to measure the similarity between new class samples and old class samples. n is the new class label information amplification matrix; m represents the sample size of this class increment, and n represents the total number of categories.
6. The incremental learning method for edge scene image recognition according to claim 1, characterized in that: The expression of the feature fusion is: In the formula, β n-1 is the Gram matrix of the previous stage; A n It represents the features of new class samples after random mapping through width learning, and T is the transposition; λ is a hyperparameter, I is the identity matrix, and A i is the sample data entering in the i-th batch, and n is the incremental task in the n-th stage.
7. The incremental learning method for edge scene image recognition according to claim 1, characterized in that: The method of fine-tuning the output layer weight of the width learning model using the Gram matrix after feature fusion includes: The calculation formula for the output layer weight of the width learning model is: In the formula, β n is the Gram matrix, Y is the training sample label matrix, and A represents the sample feature matrix after random mapping through width learning; Woodbury optimization of the Gram matrix inversion; The Cholesky factor is obtained by optimizing the inverse expansion result using Cholesky decomposition; Update the output layer weights of the width learning model according to the obtained Cholesky factor.
8. The incremental learning method for edge scene image recognition according to claim 7, characterized in that: The Woodbury optimization of the Gram matrix inversion includes: The β of the output layer weight calculation process n Split into Using the Woodbury expansion to optimize it, we get the following expansion: In the formula, is the inverse Gram matrix obtained in the previous stage, directly replace the previous incremental process The results are saved renew; The Cholesky factor is obtained by optimizing the inverse expansion result using the Cholesky decomposition, including: Introducing the Cholesky decomposition, Decompose into the lower triangular matrix P n-1 , via Q n =A n P n-1 The implementation uses the lower triangular matrix P n-1 The new class feature learning is introduced into formula (1) to obtain: For formula (2) Perform another Cholesky decomposition to obtain the lower triangular matrix G n-1 , and further obtain: Where P n-1 , G n-1 All passed and A n To obtain.
9. The incremental learning method for edge scene image recognition according to claim 8, characterized in that: The output layer weights are updated using one of the following two algorithms: Algorithm 1: Algorithm 2: Among them, W n represents the weight parameter matrix of the model after the nth incremental task is completed; M n It is a matrix used to measure the similarity between new class samples and old class samples. n is the new class label information augmentation matrix; Select different algorithms based on the sample size and feature complexity of the input model to achieve faster and more accurate model-class incremental learning.
10. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Class incremental learning method and system based on analogy learning
CN115879533A
Image classification method and device, electronic equipment and readable storage medium
CN118097283A
Robust data increment width learning method
CN118195025A
Wild bird image data stream online identification method and system fusing deep learning and width learning
CN118887704A
Prepositional phrase attachment over word embedding products
US20180260381A1
Cited By
Parallel increment width learning method and system based on updated triangular matrix
CN120430377A
A Parallel Incremental Width Learning Method and System Based on Updating Triangular Matrices
CN120430377B