Image recognition method, system and equipment based on large and small model fusion and medium
Through an image recognition method based on the fusion of large and small models, and by utilizing large model training and small model collaboration, the problems of high computing resources and insufficient accuracy in image recognition are solved, and efficient and low-resource recognition effects are achieved.
Patent Information
- Application Number
- CN202510815714.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-12
AI Technical Summary
In existing technologies, large models have high computing resource requirements in image recognition and are unable to meet real-time and low power consumption requirements, while small models have insufficient recognition accuracy in complex scenes and are unable to meet high accuracy requirements.
By converting raw image data into feature vectors, training a large model and extracting knowledge from it, training a small model, and deploying the small model on the end device, combined with the fusion strategy of the application scenario, the large and small models work together for recognition.
It reduces hardware resource requirements while ensuring recognition accuracy, improves recognition efficiency, is suitable for resource-constrained scenarios, and solves the problem of high computing resources for large models and insufficient accuracy for small models.
Smart Images

Figure CN120635651A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to an image recognition method, system, device and medium based on large and small model fusion. Background Art
[0002] With the rapid development of artificial intelligence (AI), intelligent recognition technology has penetrated deeply into fields such as image recognition, speech recognition, and natural language processing, becoming a core driver of intelligent upgrades across various industries. However, current mainstream intelligent recognition methods face significant technical bottlenecks in practical applications. Specifically, recognition solutions based on large models, with their large number of parameters and complex network structures, are able to capture deep semantic features and complex relationships in data, demonstrating excellent recognition accuracy across a wide range of tasks. However, the sheer size of large models, coupled with the numerous matrix operations and multi-layer nonlinear transformations involved in their computation, places significant demands on the hardware's computing power, memory capacity, and energy efficiency. For example, in edge computing scenarios, a single inference of a large model can take hundreds of milliseconds and require several gigabytes of storage, making it difficult to meet the real-time and low-power requirements of applications. In contrast, recognition solutions based on small models, through techniques such as network compression and parameter quantization, can reduce model size to megabytes, accelerating inference speeds to milliseconds, making them suitable for resource-constrained environments such as mobile phones and embedded devices. However, small models are limited by the number of parameters and network depth, and their feature extraction capabilities are limited. In complex scenarios, the recognition accuracy drops significantly, making it difficult to meet the needs of some application scenarios with high accuracy requirements.
[0003] Therefore, it is urgent to propose an image recognition method, system, device and medium based on large and small model fusion to solve the above problems. Summary of the Invention
[0004] In view of this, the present invention proposes an image recognition system, device and medium based on large and small model fusion, which can improve image recognition efficiency and reduce the requirements for hardware resources while ensuring image recognition accuracy.
[0005] Based on the above objectives, an embodiment of the present invention provides an image recognition method based on large and small model fusion, which specifically includes the following steps: Convert the original image dataset into feature vectors, train a preset large model based on the feature vectors, obtain the target large model and deploy it on the cloud platform; Extracting target knowledge from the middle layer of the target large model, training a preset small model based on the target knowledge, obtaining the target small model and deploying it on the end-side device; If the client device and the cloud platform are not interoperable, calling the target small model to identify the image data to be identified; If the terminal device is interoperable with the cloud platform, the image data to be identified is identified based on the preset fusion strategy corresponding to the application scenario, the target large model and the target small model to obtain a corresponding target recognition result.
[0006] In some embodiments, the step of training a preset large model based on the feature vector to obtain a target large model includes: Inputting the feature vector into the preset large model, and mapping the feature vector to obtain an embedding vector based on the embedding dimension corresponding to the embedding layer in the preset large model; adding a position code to the embedding vector, inputting the embedding vector and its position code into an encoder of the preset large model, and training the preset large model based on the encoder and first preset training parameters; Obtaining a first recognition result corresponding to the trained preset large model, and determining whether an accuracy rate of the first recognition result meets a first preset value; If not satisfied, adjusting the first preset training parameters, and returning to the step of training the preset large model based on the encoder and the first preset training parameters; If satisfied, the trained preset large model is used as the target large model.
[0007] In some embodiments, the step of extracting target knowledge from the middle layer of the target large model and training a preset small model based on the target knowledge to obtain the target small model includes: Extracting a feature map and network parameters based on the middle layer of the target large model, and using the feature map and the network parameters as the target knowledge; Training the preset small model based on knowledge distillation technology, second preset training parameters and the target knowledge; Obtaining a second recognition result corresponding to the trained preset small model, and determining whether the accuracy of the second recognition result meets a second preset value; If not, adjust the second preset training parameters, and return to the step of training the preset small model based on the knowledge distillation technology, the preset training parameters and the target knowledge; If satisfied, the trained preset small model is used as the target small model.
[0008] In some embodiments, the step of identifying the image data to be identified based on a preset fusion strategy corresponding to the application scenario, the target large model, and the target small model to obtain a corresponding target recognition result includes: Determine whether the preset fusion strategy corresponding to the application scenario is weighted fusion; If the preset fusion strategy is weighted fusion, a first weight coefficient is assigned to the target large model and a second weight coefficient is assigned to the target small model according to the accuracy of the first recognition result corresponding to the target large model and the accuracy of the second recognition result corresponding to the target small model; Inputting the image data to be recognized into the target large model to obtain a third recognition result, and inputting the image data to be recognized into the target small model to obtain a fourth recognition result; The third recognition result and the fourth recognition result are weighted and summed based on the first weight coefficient and the second weight coefficient to obtain the target recognition result.
[0009] In some embodiments, the image recognition method based on large and small model fusion further includes: If the preset fusion strategy is not weighted fusion, determining whether the preset fusion strategy is feature fusion; If the preset fusion strategy is feature fusion, the image data to be identified is input into the target large model to extract the first feature, and the image data to be identified is input into the target small model to extract the second feature; splicing the first feature and the second feature to obtain a comprehensive feature, and determining the target recognition result based on the comprehensive feature; If the preset fusion strategy is not feature fusion, the image data to be recognized is input into the target large model to obtain a fifth recognition result, and the image data to be recognized is input into the target small model to obtain a sixth recognition result; Based on a preset rule, one of the fifth recognition result and the sixth recognition result is selected as the target recognition result.
[0010] In some embodiments, the step of converting the original image dataset into a feature vector includes: Preprocessing each original image data in the original image data set based on preset preprocessing parameters to obtain standard image data; Preset types of image features are extracted based on the standard image data, and the feature vector is obtained based on conversion of all the image features.
[0011] In some embodiments, the image recognition method based on large and small model fusion further includes: Determining whether the target recognition result meets the preset user requirements; If not, adjust the preset preprocessing parameters, and return to the step of converting the original image data set into a feature vector, and training a preset large model based on the feature vector to obtain a target large model.
[0012] Another aspect of the present invention provides an image recognition system based on large and small model fusion, including: A large model training unit is configured to convert the original image data set into a feature vector, and train a preset large model based on the feature vector to obtain a target large model and deploy it on the cloud platform; A small model training unit is configured to extract target knowledge from the middle layer of the target large model, train a preset small model based on the target knowledge, obtain a target small model, and deploy it on a terminal device; The identification unit is configured to call the target small model to identify the image data to be identified if the terminal device and the cloud platform are not interoperable; if the terminal device and the cloud platform are interoperable, the image data to be identified is identified based on the preset fusion strategy corresponding to the application scenario, the target large model and the target small model to obtain the corresponding target recognition result.
[0013] According to another aspect of an embodiment of the present invention, a computer device is provided, comprising: at least one CPU processor and a GPU processor; and a memory, wherein the memory stores a computer program that can be run on the CPU processor and the GPU processor, and the computer program implements the steps of the above method when executed by the CPU processor and the GPU processor.
[0014] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided, which stores a computer program that implements the above method steps when executed by a processor.
[0015] The present invention has at least the following beneficial technical effects: The image recognition method based on large and small model fusion of the present invention converts the original image data set into feature vectors and trains a preset large model based on them. It can eliminate noise interference in the image and compress the data dimension through the feature extraction step, so that the large model focuses on learning key features, thereby obtaining a high-precision target large model. Target knowledge containing deep feature representations is extracted from the middle layer of the target large model to train a preset small model. With the help of knowledge distillation technology, the small model can inherit the core recognition capabilities of the large model while maintaining a lightweight architecture, achieving fast reasoning. Based on the preset fusion strategy, the target large model and the target small model are operated in coordination, which can not only use the high precision of the large model to compensate for the recognition defects of the small model in complex scenarios, but also reduce the overall computing overhead through the high efficiency of the small model. Ultimately, in the processing of the image data to be recognized, the combined effect of recognition accuracy close to that of the large model, high recognition efficiency and significantly reduced hardware resource usage is achieved. This effectively solves the industry pain points of the traditional large model's strong dependence on computing power and the insufficient accuracy of the small model, and provides an intelligent recognition solution that balances performance and efficiency for resource-constrained scenarios such as edge computing and real-time monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 A block diagram of an embodiment of an image recognition method based on large and small model fusion provided by the present invention; Figure 2 A flowchart of an embodiment of the processing flow of the image recognition method based on large and small model fusion provided by the present invention; Figure 3 A schematic diagram of an embodiment of an image recognition system based on large and small model fusion provided by the present invention; Figure 4 A schematic structural diagram of an embodiment of a computer device provided by the present invention; Figure 5 This is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION
[0018] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the embodiments of the present invention are further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0019] It should be noted that all expressions using "first" and "second" in the embodiments of the present invention are for distinguishing two non-identical entities with the same name or non-identical parameters. It can be seen that "first" and "second" are only for the convenience of expression and should not be understood as limitations on the embodiments of the present invention. Subsequent embodiments will not explain this one by one.
[0020] Based on the above objectives, the first aspect of the embodiment of the present invention provides an embodiment of an image recognition method based on large and small model fusion. Figure 1 As shown, it includes the following steps: Step S100: converting the original image dataset into feature vectors, and training a preset large model based on the feature vectors to obtain a target large model and deploy it on the cloud platform; Step S200: extract target knowledge from the middle layer of the target large model, train a preset small model based on the target knowledge, obtain the target small model, and deploy it on the end-side device; Step S300: If the client device and the cloud platform are not interoperable, the target small model is called to recognize the image data to be recognized; In step S400, if the terminal device is interoperable with the cloud platform, the image data to be identified is identified based on the preset fusion strategy, target large model and target small model corresponding to the application scenario to obtain the corresponding target recognition result.
[0021] In some embodiments, the image recognition method based on large and small model fusion of the present invention also includes: determining whether the target recognition result meets the preset user requirements; if not, adjusting the preset preprocessing parameters, and returning to the step of converting the original image data set into a feature vector, and training the preset large model based on the feature vector to obtain the target large model.
[0022] In some embodiments, as Figure 2 As shown, the original image dataset is first preprocessed and converted into feature vectors. Based on these feature vectors, a pre-set large model is trained. The Vision Transformer is used as the pre-set large model, which has powerful feature extraction and learning capabilities. During training, by learning the feature vectors, the large model can grasp the complex features and deep semantic information of the data, thereby obtaining a target large model with high recognition accuracy. Target knowledge is extracted from the intermediate layers of the target large model. This target knowledge includes key features and semantic information learned by the large model during data processing. Then, based on this target knowledge, a pre-set small model is trained. The small model can adopt a lightweight neural network architecture, such as a mobile network or a compressed neural network. Through methods such as knowledge distillation, the small model inherits the recognition capabilities of the large model as much as possible while maintaining a small model size, resulting in a target small model that achieves fast and efficient recognition. Based on a pre-set fusion strategy corresponding to the application scenario, the target large model, and the target small model, the target recognition results corresponding to the image data to be recognized are obtained. Pre-set fusion strategies include, but are not limited to, weighted fusion, feature fusion, and decision fusion. Weighted fusion combines the outputs of the large and small models based on their confidence levels or weights to produce the final recognition result. Feature fusion combines or fuses the features extracted by the large and small models to form a richer feature representation, which is then input into a classifier for recognition. Decision fusion comprehensively considers the classification results of the large and small models and determines the final recognition category through decision-making rules or voting mechanisms. In practical applications, the image data to be recognized is simultaneously input into the large and small models. After processing and calculation by each model, the final target recognition result is obtained based on the selected fusion strategy.
[0023] In some embodiments, after obtaining the target recognition results, it is necessary to continuously collect feedback on the recognition results and, based on data changes in actual application scenarios, determine whether the target recognition results meet the preset user requirements. If not, the preset preprocessing parameters are adjusted, such as changing the image scaling, adjusting the cropped area, modifying the grayscale threshold or the denoising intensity, etc. The process then returns to the step of converting the original image dataset into feature vectors and training a preset large model based on the feature vectors to obtain the target large model. The entire process is restarted to continuously improve the model's recognition performance and adaptability, so that it can better meet the preset user requirements.
[0024] The image recognition method based on the fusion of large and small models of the present invention can utilize the high precision of the large model to compensate for the recognition defects of the small model in complex scenarios, and can also reduce the overall computing overhead through the high efficiency of the small model. Ultimately, in the processing of the image data to be recognized, it achieves the combined effect of recognition accuracy close to that of the large model, high recognition efficiency, and significantly reduced hardware resource usage. It effectively solves the industry pain points of the traditional large model's strong reliance on computing power and the insufficient accuracy of the small model, and provides an intelligent recognition solution that balances performance and efficiency for resource-constrained scenarios such as edge computing and real-time monitoring. It can also be widely used in other intelligent recognition fields, such as intelligent speech recognition, image recognition, fault diagnosis, etc., and has broad application prospects and market value.
[0025] In some embodiments, the step of training a preset large model based on a feature vector to obtain a target large model includes: inputting a feature vector into the preset large model, mapping the feature vector to obtain an embedding vector based on the embedding dimension corresponding to the embedding layer in the preset large model; adding a position code to the embedding vector, inputting the embedding vector and its position code into the encoder of the preset large model, and training the preset large model based on the encoder and the first preset training parameters; obtaining a first recognition result corresponding to the trained preset large model, and determining whether the accuracy of the first recognition result meets the first preset value; if not, adjusting the first preset training parameters, and returning to the step of training the preset large model based on the encoder and the first preset training parameters; if satisfied, using the trained preset large model as the target large model.
[0026] In some embodiments, the first preset training parameters include a learning rate, number of iterations, etc. The feature vector generated after preprocessing is input into the preset large model, and based on the embedding dimension corresponding to the embedding layer in the preset large model, the feature vector is mapped to an embedding vector through linear projection. Next, a position code is added to the embedding vector to retain the spatial position information of the feature, and the embedding vector and its position code are input into the encoder of the preset large model. Then, the preset large model is trained based on the network structure of the encoder and the first preset training parameters. During the training process, the first recognition result of the trained preset large model for the sample data is obtained, and it is determined whether the accuracy of the recognition result meets the first preset value (such as 95%). If not, the first preset training parameters are adjusted, such as reducing the learning rate or increasing the number of iterations, and the training is continued based on the encoder and the adjusted training parameters. If it is satisfied, the trained preset large model is used as the target large model, which has the ability to learn complex features and deep semantics of the data.
[0027] In some embodiments, the Vision Transformer (an image classification model) is selected as the pre-set master model. When training pre-processed raw image data, the following specific operations are performed: First, the raw image data is converted into feature vectors. Based on the dimensions of the Vision Transformer embedding layer (e.g., 768 dimensions), the feature vectors are mapped into embedding vectors of corresponding dimensions via linear projection. Simultaneously, the image is segmented into patches of a fixed size (e.g., 16×16 pixels). Each patch is mapped via the embedding layer and then superimposed with a positional encoding to form a sequence of embedding vectors containing spatial information. Next, the encoder network structure is configured, such as a 12-layer Transformer (a deep learning model architecture) encoder with a 12-head attention mechanism. The first preset training parameters are configured with a learning rate of 5e-5 and a number of iterations of 300. The embedding vector sequence is input into the encoder. A multi-head self-attention mechanism and a feedforward neural network are used to extract features layer by layer, enabling the model to learn edge, texture, and semantic information from the raw image data, such as the outline features of pedestrians and vehicles. During training, the model's recognition accuracy for objects (pedestrians, vehicles, animals, etc.) in the image is monitored in real time. When the model recognition accuracy reaches 96.45%, it is determined to meet the first preset value. If the accuracy does not reach the threshold during training, for example, after 100 iterations, the accuracy is only 85%, the learning rate is adjusted to 1e-5 and the number of iterations is increased to 500 until the accuracy reaches the target, and the target large model is finally obtained.
[0028] The image recognition method based on large and small model fusion of the present invention, through the standardized processing of feature vectors, enables the large model to have stronger generalization ability for input data of different sizes and formats. Especially in resource-constrained scenarios, the feature vectors can reduce the dependence of the large model on high-end hardware, while maintaining high accuracy, and reduce the computing power requirement for a single recognition, laying an efficient foundation for subsequent small model distillation and model fusion.
[0029] In some embodiments, target knowledge is extracted from the middle layer of the target large model, and a preset small model is trained based on the target knowledge to obtain the target small model, including: extracting feature maps and network parameters based on the middle layer of the target large model, and using the feature maps and network parameters as target knowledge; training the preset small model based on knowledge distillation technology, second preset training parameters and target knowledge; obtaining a second recognition result corresponding to the trained preset small model, and determining whether the accuracy of the second recognition result meets a second preset value; if not, adjusting the second preset training parameters, and returning to the step of training the preset small model based on knowledge distillation technology, preset training parameters and knowledge; if satisfied, using the trained preset small model as the target small model.
[0030] In some embodiments, the second preset training parameters include a learning rate, training rounds, and other parameters. The steps for extracting target knowledge from the intermediate layers of the target large model and training the preset small model to obtain the target small model are as follows: First, feature maps and network parameters are extracted from the intermediate layers of the target large model and used as the target knowledge. The feature maps contain semantic information at different levels learned by the large model during data processing, while the network parameters carry the training experience of the large model. Next, the preset small model is trained based on knowledge distillation techniques, the second preset training parameters, and the target knowledge. During training, the complex knowledge of the large model is transferred to the small model by minimizing the difference between the outputs of the small model and the large model. During training, a second recognition result of the trained preset small model on the sample data is obtained, and a determination is made as to whether the accuracy of the recognition result meets a second preset value (e.g., 85%). If not, the second preset training parameters are adjusted, such as by reducing the learning rate or increasing the number of training rounds, and the training process is resumed. If so, the trained preset small model is used as the target small model. This small model inherits the recognition capabilities of the large model while maintaining a smaller model size.
[0031] In some embodiments, the target large model is Vision Transformer, and the preset small model is a lightweight MobileNet (a lightweight convolutional neural network) as an example. First, feature maps are extracted from the middle layer of Vision Transformer, such as the shallow edge feature map and the deep semantic feature map, and network parameters, such as the weight matrix of the attention layer and the feedforward layer, are extracted and used as target knowledge. When training MobileNet, based on the knowledge distillation technology, the learning rate in the second preset training parameter is set to 1e-4, and the training rounds are set to 200 rounds. By calculating the differences in feature maps and output probability distributions between the small model and the large model, such as using mean square error to measure the difference in feature maps or using cross entropy loss to measure the difference in probability distribution, the parameters of MobileNet are continuously adjusted to make the output of the small model as close as possible to the large model. During the training process, the accuracy of the second recognition result of MobileNet is monitored in real time. When the model recognition accuracy reaches 86.7% and meets the second preset value, the training is completed. If the accuracy does not reach the threshold during training, for example, the accuracy is only 75% after 100 rounds of training, the learning rate is adjusted to 5e-5 and the number of training rounds is increased to 300 rounds until the accuracy reaches the target, and the target small model is finally obtained.
[0032] The image recognition method based on the fusion of large and small models of the present invention obtains the target knowledge of the middle layer of the target large model to train a preset small model. The target small model finally trained can effectively inherit the recognition ability of the large model while maintaining a fast reasoning speed, and can achieve fast feature extraction and classification recognition, which is suitable for resource-constrained edge device scenarios.
[0033] In some embodiments, based on the preset fusion strategy, target large model and target small model corresponding to the application scenario, the image data to be identified is identified to obtain the corresponding target recognition result, including: determining whether the preset fusion strategy corresponding to the application scenario is weighted fusion; if the preset fusion strategy is weighted fusion, according to the accuracy of the first recognition result corresponding to the target large model and the accuracy of the second recognition result corresponding to the target small model, assigning a first weight coefficient to the target large model and a second weight coefficient to the target small model; inputting the image data to be identified into the target large model to obtain a third recognition result, and inputting the image data to be identified into the target small model to obtain a fourth recognition result; performing weighted summation on the third recognition result and the fourth recognition result based on the first weight coefficient and the second weight coefficient to obtain the target recognition result.
[0034] In some embodiments, the image recognition method based on large and small model fusion of the present invention also includes: if the preset fusion strategy is not weighted fusion, determining whether the preset fusion strategy is feature fusion; if the preset fusion strategy is feature fusion, inputting the image data to be identified into the target large model to extract the first feature, and inputting the image data to be identified into the target small model to extract the second feature; splicing the first feature and the second feature to obtain a comprehensive feature, and determining the target recognition result based on the comprehensive feature; if the preset fusion strategy is not feature fusion, inputting the image data to be identified into the target large model to obtain a fifth recognition result, and inputting the image data to be identified into the target small model to obtain a sixth recognition result; based on the preset rules, selecting one from the fifth recognition result and the sixth recognition result as the target recognition result.
[0035] In some embodiments, based on the preset fusion strategy, target large model and target small model corresponding to the application scenario, the image data to be identified is identified to obtain the corresponding target recognition result. The specific steps are as follows: First, the type of the preset fusion strategy corresponding to the application scenario is determined. If the preset fusion strategy corresponding to the application scenario is weighted fusion, then according to the accuracy of the first recognition result corresponding to the target large model and the accuracy of the second recognition result corresponding to the target small model, the first weight coefficient is assigned to the target large model and the second weight coefficient is assigned to the target small model respectively to obtain the target recognition result by weighted summation. If it is not weighted fusion, it is further determined whether it is feature fusion. If it is feature fusion, the first feature and the second feature are extracted from the target large model and the target small model respectively, and the comprehensive feature is obtained by splicing, and the target recognition result is determined based on the comprehensive feature. If it is not feature fusion, the decision fusion strategy is adopted by default, and the image data to be identified is input into the two models respectively to obtain the fifth recognition result and the sixth recognition result, and one of them is selected as the target recognition result based on the preset rules. The preset fusion strategy includes but is not limited to the above three strategies. Specifically, for general application scenarios that require a balance between accuracy and speed, such as industrial quality inspection and daily security monitoring, it is necessary to ensure a certain recognition accuracy rate while requiring the system to respond quickly. In this case, the preset fusion strategy is weighted fusion. For complex application scenarios that prioritize high precision, such as medical imaging diagnosis and remote sensing image analysis, extremely high recognition accuracy is required, and the model must simultaneously capture detailed features and deep semantics. In this case, the preset fusion strategy is feature fusion. For application scenarios that prioritize efficiency, with strong real-time or resource-constrained requirements, such as autonomous driving and smart home devices, data processing must meet real-time requirements and the device computing resources are limited. In this case, the preset fusion strategy is decision fusion.
[0036] In some embodiments, when the preset fusion strategy corresponding to the application scenario is weighted fusion, a first weight coefficient of 0.7 is assigned to the target large model and a second weight coefficient of 0.3 is assigned to the target small model based on the 96.45% accuracy of the target large model and the 83.7% accuracy of the target small model on the validation set. The image data to be recognized is then fed into the target large model to obtain a third recognition result, for example, with a class probability distribution of [cat: 0.8, dog: 0.1, other: 0.1]. The image data is then fed into the target small model to obtain a fourth recognition result, for example, with a class probability distribution of [cat: 0.6, dog: 0.3, other: 0.1]. The third and fourth recognition results are weighted and summed based on the first and second weight coefficients, such as [cat: 0.7×0.8+0.3×0.6=0.74, dog: 0.7×0.1+0.3×0.3=0.16, others: 0.7×0.1+0.3×0.1=0.1]. The final target recognition result is "cat." This result retains the high precision of the large model while leveraging the fast inference capabilities of the small model, effectively improving overall recognition performance.
[0037] In some embodiments, when the preset fusion strategy corresponding to the application scenario is feature fusion, the image data to be identified is input into the target large model to extract the first feature containing deep semantics, such as the abstract feature of the object category. The image data to be identified is input into the target small model to extract the second feature of shallow details, such as edge and texture features. For example, for an animal image, the large model extracts the semantic feature of "mammal" and the small model extracts the detail feature of "striped texture". The first feature and the second feature are spliced to obtain a comprehensive feature vector containing semantics and details. The comprehensive feature vector is input into the classifier for processing, and the category probability distribution is calculated through the fully connected layer and the Softmax function to finally determine the target recognition result (such as "tiger"). This strategy enhances the model's recognition ability for complex images by fusing features at different levels.
[0038] In some embodiments, when the preset fusion strategy corresponding to the application scenario is neither weighted fusion nor feature fusion, a decision fusion strategy is adopted. The image data to be identified is input into the target large model to obtain the fifth recognition result, such as being judged as a "car". The image data to be identified is input into the target small model to obtain the sixth recognition result, such as being judged as a "truck". Based on preset rules, such as "when the results of the large model and the small model conflict, the results of the large model shall prevail", the "car" output by the large model is directly selected as the target recognition result. Or a voting mechanism is adopted to give different voting weights to the judgments of the two models in combination with historical accuracy data, such as a large model weight of 60% and a small model weight of 40%, and the final category is determined according to the weighted voting results. This strategy uses flexible decision-making logic to balance model accuracy while ensuring efficiency.
[0039] The present invention's image recognition method, based on large and small model fusion, achieves breakthroughs in recognition accuracy, inference efficiency, and scenario adaptability by flexibly combining the advantages of large and small models. The collaborative design of multiple fusion strategies enables the system to dynamically switch strategies based on real-time needs, significantly improving the model's generalization capabilities and engineering value, providing efficient and intelligent solutions for edge computing, real-time monitoring, and other fields.
[0040] In some embodiments, the step of converting the original image data set into a feature vector includes: preprocessing each original image data in the original image data set based on preset preprocessing parameters to obtain standard image data; extracting preset types of image features based on the standard image data, and converting all image features to obtain a feature vector.
[0041] In some embodiments, the preset types include image feature types such as edges, textures, and colors. Preprocessing parameters include scaling, cropping, grayscale thresholds, or denoising intensity. Preprocessing includes scaling all raw image data in the original image dataset to a uniform size, cropping and focusing the effective area, grayscale simplification of color dimensions, and denoising to eliminate interfering pixels, to enhance data quality and consistency. For the preprocessed raw image data, image features such as edges, textures, and colors are extracted and converted into feature vectors suitable for model input.
[0042] In some implementations, a small target model is deployed on an end-device (such as a smart camera or mobile phone), and a large target model is deployed on a cloud platform. When the end-device operates solely on its own computing resources, the small target model on the end-device is directly called for target recognition and classification, leveraging the small target model's rapid response to achieve low-latency recognition. When the end-device and the cloud platform are interoperable, a fusion model is called for target recognition and classification, combining the high precision of the large model with the high efficiency of the small model to ensure recognition accuracy while also taking into account speed. For example, in a security monitoring scenario, daily monitoring is performed using a small end-device model to quickly screen images. When complex scenarios arise, the end-device uploads the data to the cloud for accurate recognition using the fusion model, effectively overcoming the shortcomings of the large model's high hardware requirements and the small model's insufficient accuracy.
[0043] The image recognition system based on large and small model fusion of the present invention converts the original image data set into feature vectors and trains a preset large model based on them. It can eliminate noise interference in the image and compress the data dimension through the feature extraction process, so that the large model focuses on learning key features, thereby obtaining a high-precision target large model. Target knowledge containing deep feature representations is extracted from the middle layer of the target large model to train a preset small model. With the help of knowledge distillation technology, the small model can inherit the core recognition capabilities of the large model while maintaining a lightweight architecture, achieving fast reasoning. Based on a preset fusion strategy corresponding to the application scenario, the target large model and the target small model are operated in coordination. The high precision of the large model can be used to compensate for the recognition defects of the small model in complex scenarios, and the high efficiency of the small model can be used to reduce the overall computing overhead. Ultimately, in the processing of the image data to be recognized, the combined effect of recognition accuracy close to that of the large model, high recognition efficiency and significantly reduced hardware resource usage is achieved. This effectively solves the industry pain points of the traditional large model's strong reliance on computing power and the insufficient accuracy of the small model, and provides an intelligent recognition solution that balances performance and efficiency for resource-constrained scenarios such as edge computing and real-time monitoring.
[0044] Based on the same inventive concept, according to another aspect of the present invention, Figure 3 As shown, an embodiment of the present invention further provides an image recognition system based on large and small model fusion, comprising: The large model training unit 110 is configured to convert the original image data set into feature vectors, and train a preset large model based on the feature vectors to obtain a target large model and deploy it on the cloud platform; The small model training unit 120 is configured to extract target knowledge from the middle layer of the target large model, train a preset small model based on the target knowledge, obtain the target small model, and deploy it on the end-side device; The recognition unit 130 is configured to call the target small model to identify the image data to be identified if the end-side device and the cloud platform are not interoperable; if the end-side device and the cloud platform are interoperable, the image data to be identified is identified based on the preset fusion strategy, target large model and target small model corresponding to the application scenario to obtain the corresponding target recognition result.
[0045] The image recognition system based on large and small model fusion of the present invention converts the original image data set into feature vectors and trains a preset large model based on them. It can eliminate noise interference in the image and compress the data dimension through the feature extraction process, so that the large model focuses on learning key features, thereby obtaining a high-precision target large model. Target knowledge containing deep feature representations is extracted from the middle layer of the target large model to train a preset small model. With the help of knowledge distillation technology, the small model can inherit the core recognition capabilities of the large model while maintaining a lightweight architecture, achieving fast reasoning. Based on a preset fusion strategy corresponding to the application scenario, the target large model and the target small model are operated in coordination. The high precision of the large model can be used to compensate for the recognition defects of the small model in complex scenarios, and the high efficiency of the small model can be used to reduce the overall computing overhead. Ultimately, in the processing of the image data to be recognized, the combined effect of recognition accuracy close to that of the large model, high recognition efficiency and significantly reduced hardware resource usage is achieved. This effectively solves the industry pain points of the traditional large model's strong reliance on computing power and the insufficient accuracy of the small model, and provides an intelligent recognition solution that balances performance and efficiency for resource-constrained scenarios such as edge computing and real-time monitoring.
[0046] Based on the same inventive concept, according to another aspect of the present invention, Figure 4 As shown, an embodiment of the present invention further provides a computer device 30, which includes a CPU processor and a GPU processor 310 and a memory 320. The memory 320 stores a computer program 321 that can be run on the CPU processor and the GPU processor. When the CPU processor and the GPU processor 310 execute the program, the steps of the above method are performed. The GPU processor is any artificial intelligence processor including an NPU (Neural Processing Unit).
[0047] Based on the same inventive concept, according to another aspect of the present invention, Figure 5 As shown, an embodiment of the present invention further provides a computer-readable storage medium 40 , which stores a computer program 410 for executing the above method when executed by a processor.
[0048] Finally, it should be noted that those skilled in the art will understand that all or part of the processes in the above-described method embodiments can be implemented using a computer program to instruct the relevant hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes in the above-described method embodiments. The program storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM). The above-described computer program embodiments can achieve the same or similar effects as any of the corresponding aforementioned method embodiments.
[0049] It will also be appreciated by those skilled in the art that the various exemplary logic blocks, modules, circuits and algorithmic steps described in conjunction with the disclosure herein can be implemented as electronic hardware, computer software or a combination of the two. In order to clearly illustrate this interchangeability of hardware and software, a general description has been given of the functions of various schematic components, blocks, modules, circuits and steps. Whether this function is implemented as software or hardware depends on specific applications and the design constraints imposed on the entire system. Those skilled in the art can implement the function in various ways for each specific application, but this implementation decision should not be interpreted as causing a departure from the disclosed scope of the embodiments of the present invention.
[0050] The above are exemplary embodiments disclosed in the present invention, but it should be noted that various changes and modifications can be made without departing from the scope of the disclosure of the embodiments of the present invention as defined in the claims. The functions, steps and / or actions of the method claims according to the disclosed embodiments described herein do not need to be performed in any particular order. The serial numbers of the embodiments disclosed in the above embodiments of the present invention are for description only and do not represent the advantages and disadvantages of the embodiments. In addition, although the elements disclosed in the embodiments of the present invention can be described or required in individual form, they can also be understood as multiple unless expressly limited to the singular.
[0051] It should be understood that, as used herein, the singular forms "a" and "an" are intended to include the plural forms as well, unless the context clearly supports an exception. It should also be understood that, as used herein, "and / or" is intended to include any and all possible combinations of one or more of the associated listed items.
[0052] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to limit the scope of the disclosure of the present invention (including the claims) to these examples. Within the spirit of the present invention, the technical features of the above embodiments or different embodiments may be combined, and many other variations exist in different aspects of the above embodiments, which are not provided in detail for the sake of clarity. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. An image recognition method based on large and small model fusion, characterized in that: include: Convert the original image dataset into feature vectors, train a preset large model based on the feature vectors, obtain the target large model and deploy it on the cloud platform; Extracting target knowledge from the middle layer of the target large model, training a preset small model based on the target knowledge, obtaining the target small model and deploying it on the end-side device; If the client device and the cloud platform are not interoperable, calling the target small model to identify the image data to be identified; If the terminal device is interoperable with the cloud platform, the image data to be identified is identified based on the preset fusion strategy corresponding to the application scenario, the target large model and the target small model to obtain a corresponding target recognition result.
2. The image recognition method based on large and small model fusion according to claim 1, characterized in that: The step of training a preset large model based on the feature vector to obtain a target large model includes: Inputting the feature vector into the preset large model, and mapping the feature vector to obtain an embedding vector based on the embedding dimension corresponding to the embedding layer in the preset large model; adding a position code to the embedding vector, inputting the embedding vector and its position code into an encoder of the preset large model, and training the preset large model based on the encoder and first preset training parameters; Obtaining a first recognition result corresponding to the trained preset large model, and determining whether an accuracy rate of the first recognition result meets a first preset value; If not satisfied, adjusting the first preset training parameters, and returning to the step of training the preset large model based on the encoder and the first preset training parameters; If satisfied, the trained preset large model is used as the target large model.
3. The image recognition method based on large and small model fusion according to claim 2, characterized in that: The step of extracting target knowledge from the middle layer of the target large model, training a preset small model based on the target knowledge, and obtaining the target small model includes: Extracting a feature map and network parameters based on the middle layer of the target large model, and using the feature map and the network parameters as the target knowledge; Training the preset small model based on knowledge distillation technology, second preset training parameters and the target knowledge; Obtaining a second recognition result corresponding to the trained preset small model, and determining whether the accuracy of the second recognition result meets a second preset value; If not, adjust the second preset training parameters, and return to the step of training the preset small model based on the knowledge distillation technology, the preset training parameters and the target knowledge; If satisfied, the trained preset small model is used as the target small model.
4. The image recognition method based on large and small model fusion according to claim 3, characterized in that: The step of identifying the image data to be identified based on the preset fusion strategy corresponding to the application scenario, the target large model, and the target small model to obtain the corresponding target recognition result includes: Determine whether the preset fusion strategy corresponding to the application scenario is weighted fusion; If the preset fusion strategy is weighted fusion, a first weight coefficient is assigned to the target large model and a second weight coefficient is assigned to the target small model according to the accuracy of the first recognition result corresponding to the target large model and the accuracy of the second recognition result corresponding to the target small model; Inputting the image data to be recognized into the target large model to obtain a third recognition result, and inputting the image data to be recognized into the target small model to obtain a fourth recognition result; The third recognition result and the fourth recognition result are weighted and summed based on the first weight coefficient and the second weight coefficient to obtain the target recognition result.
5. The image recognition method based on large and small model fusion according to claim 4, characterized in that: Also includes: If the preset fusion strategy is not weighted fusion, determining whether the preset fusion strategy is feature fusion; If the preset fusion strategy is feature fusion, the image data to be identified is input into the target large model to extract the first feature, and the image data to be identified is input into the target small model to extract the second feature; splicing the first feature and the second feature to obtain a comprehensive feature, and determining the target recognition result based on the comprehensive feature; If the preset fusion strategy is not feature fusion, the image data to be recognized is input into the target large model to obtain a fifth recognition result, and the image data to be recognized is input into the target small model to obtain a sixth recognition result; Based on a preset rule, one of the fifth recognition result and the sixth recognition result is selected as the target recognition result.
6. The image recognition method based on large and small model fusion according to claim 1, characterized in that: The step of converting the original image data set into a feature vector comprises: Preprocessing each original image data in the original image data set based on preset preprocessing parameters to obtain standard image data; Preset types of image features are extracted based on the standard image data, and the feature vector is obtained based on conversion of all the image features.
7. The image recognition method based on large and small model fusion according to claim 6, characterized in that: Also includes: Determining whether the target recognition result meets the preset user requirements; If not, adjust the preset preprocessing parameters, and return to the step of converting the original image data set into a feature vector, and training a preset large model based on the feature vector to obtain a target large model.
8. An image recognition system based on large and small model fusion, characterized in that: include: A large model training unit is configured to convert the original image data set into a feature vector, and train a preset large model based on the feature vector to obtain a target large model and deploy it on the cloud platform; A small model training unit is configured to extract target knowledge from the middle layer of the target large model, train a preset small model based on the target knowledge, obtain a target small model, and deploy it on a terminal device; The identification unit is configured to call the target small model to identify the image data to be identified if the terminal device and the cloud platform are not interoperable; if the terminal device and the cloud platform are interoperable, the image data to be identified is identified based on the preset fusion strategy corresponding to the application scenario, the target large model and the target small model to obtain the corresponding target recognition result.
9. A computer device comprising: At least one CPU processor and GPU processor; as well as A memory storing a computer program that can be run on the CPU processor and the GPU processor, wherein the CPU processor and the GPU processor execute the steps of the method according to any one of claims 1 to 7 when executing the program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are performed.
Citation Information
Patent Citations
Knowledge distillation and quantification technology for power scene edge calculation large model compression
CN115223049A
Unmanned aerial vehicle image cloud edge collaborative identification method based on knowledge distillation
CN116563731A
Training method for gesture recognition model, gesture recognition method, and apparatus
WO2021190046A1