Computer image processing equipment based on big data
Through knowledge distillation training of designing data processing modules and deep learning models in computer image processing equipment, the problem of unsatisfactory image processing effects due to different imaging formats is solved, and higher image processing accuracy and model adaptability are achieved.
Patent Information
- Application Number
- CN202510098316.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing computer image processing technology has different imaging formats, resulting in unsatisfactory image processing effects. Especially in the fields of medical and security, the diversity and complexity of images make it more difficult to train models.
A computer image processing device based on big data is designed to collect image data from multiple data sources through a data processing module, and perform normalization processing, sizing adjustment and data enhancement. At the same time, a deep learning teacher model with pre-trained weight initialization is built, and a lightweight student model is trained through knowledge distillation to improve the generalization ability and adaptability of the model.
Through the multi-source data collection and enhancement of large data sets, the model can better adapt to changes in different angles, lighting and local feature, and improve the accuracy and generalization ability of image processing. Lightweight student models can also quickly reason in resource-constrained environments to meet practical application needs.
Smart Images

Figure CN120047331A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and specifically to a computer image processing device based on big data. Background Art
[0002] With the rapid development of information technology, computer image processing plays a crucial role in many fields, especially in the medical and security fields. In the medical industry, doctors need to accurately analyze a large number of medical images, such as CT, MRI, ultrasound images, and pathological section images, etc., in order to diagnose diseases in a timely manner, evaluate the development of the condition, and formulate treatment plans. In the security field, the effective processing of surveillance images can achieve functions such as target detection and behavior recognition, which is of great significance for ensuring public safety and preventing and responding to various security incidents.
[0003] However, the current computer image processing technology faces many challenges. On the one hand, at the data level, the complexity and diversity of image data bring difficulties to model training. The sources of medical images and security images are extensive, and the images generated by different devices vary greatly in modalities, resolutions, imaging characteristics, etc. For example, the image data formats and contents generated by imaging devices in different departments of medical images are different, and the images captured by different locations and different types of surveillance cameras in the security field also vary widely in perspectives, lighting conditions, target scales, and behavior performances, resulting in unsatisfactory image processing effects. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a computer image processing device based on big data, which solves the problem of unsatisfactory image processing effects caused by different imaging formats in multiple image processing fields.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A computer image processing device based on big data, comprising:
[0006] A data processing module, which collects image data from multiple data sources, and performs normalization processing, size adjustment, and data augmentation operations on the collected image data;
[0007] A teacher model construction module, which is used to construct a deep learning teacher model with pre-trained weight initialization, and train the teacher model using the cross-entropy loss function;
[0008] A student model design module, which is used to design a lightweight student model architecture, and initialize it with randomly initialized or model weights trained using a small dataset of similar tasks;
[0009] The knowledge distillation training module inputs the images in the training dataset into the trained teacher model, calculates the distillation loss, computes the gradients through the backpropagation algorithm, and updates the parameters of the student model according to the gradients;
[0010] The deployment module deploys the trained student model to the device, processes the newly input images, and outputs the processing results. Meanwhile, this module can collect new image data to update the teacher model and the student model.
[0011] Preferably, the data processing module includes:
[0012] The data collection unit collects image data from multiple data sources, and performs normalization processing and size adjustment on the collected image data;
[0013] The data augmentation unit performs augmentation operations on the collected data, including random rotation, flipping, and cropping, and adjusts the label information accordingly during the data augmentation process to maintain the consistency between the image and the label.
[0014] Preferably, the teacher model construction module includes:
[0015] The feature capture unit selects the deep learning teacher model architecture and simultaneously captures the feature information of images at different scales;
[0016] The initialization unit initializes the weights of the similarity layer and the fully connected layer;
[0017] The training unit defines the loss function of the teacher model, and dynamically adjusts the division into the training set, validation set, and test set according to the size and distribution of the dataset;
[0018] The performance evaluation unit evaluates and adjusts the performance of the teacher model on the validation set.
[0019] Preferably, the student model design module includes:
[0020] The student model design unit adopts an inverted residual structure and a linear bottleneck layer to design a lightweight student model architecture;
[0021] The student model initialization unit uses one of the initialization methods, i.e., random initialization (randomly generating the initial parameters of the model) and initializing with the model weights of a model with similar task characteristics trained on a small dataset, to perform initialization processing on the model.
[0022] Preferably, the knowledge distillation training module includes:
[0023] The soft label generation unit, for each image in the training dataset, inputs it into the trained teacher model to obtain the output probability distribution of the teacher model, and this output probability distribution is the soft label;
[0024] A loss calculation unit inputs an image into the student model to obtain the output probability distribution of the student model, and uses the KL divergence to measure the difference between the output of the student model and the output of the teacher model.
[0025] Preferably, the deployment module includes:
[0026] A model inference unit deploys the trained student model to a computer image processing device based on big data;
[0027] An update and optimization unit regularly collects new image data and corresponding annotations to update and train the teacher model and the student model.
[0028] Preferably, the data augmentation unit further includes a noise addition unit, which simulates image changes under different lighting and environmental conditions by adding Gaussian noise.
[0029] Preferably, the architecture of the deep learning teacher model in the feature capture unit is selected from one of ResNet-101 and Inception-v4, and the weights of the fully connected layer are re-initialized according to the number of categories in the target dataset.
[0030] Preferably, the lightweight student model architecture in the student model design unit is MobileNet-v2, which uses an inverted residual structure and a linear bottleneck layer to reduce the amount of computation.
[0031] Preferably, after the deployment module outputs the processing result, for the image classification task, it directly outputs the category information, and for the object detection task, it outputs the category and bounding box coordinate information, and further analyzes and makes decisions according to the output result.
[0032] The present invention provides a computer image processing device based on big data. It has the following beneficial effects:
[0033] 1. The present invention collects image data from multiple data sources through a data processing module, covering multiple fields such as medical and security, providing accurate supervision information for model training. At the same time, the data augmentation unit uses a variety of data augmentation operations, enabling the model to learn richer image features, reducing the overfitting phenomenon, and also enabling the model to better adapt to different angles, lighting, and local feature changes, etc., improving the generalization ability and accuracy of the model in complex scenarios.
[0034] 2. The teacher model construction module of the present invention selects a suitable deep learning architecture according to different image processing tasks and data characteristics, and specifically selects images that can better process complex structures or multi-scale targets. During the initialization process, the pre-trained weights on a large-scale general image dataset are fully utilized, and the fully connected layers are reasonably randomly initialized and separately trained, which speeds up the convergence rate of the teacher model on the target dataset.
[0035] 3. The student model design module of the present invention adopts a lightweight convolutional neural network architecture. Its inverted residual structure and linear bottleneck layer design have significant advantages. By using 1×1 convolution to increase and decrease the dimension and depthwise separable convolution to reduce the computational amount, while maintaining a certain feature extraction ability, it meets the fast inference requirements in resource-constrained environments.
[0036] 4. In the present invention, the soft label generation unit reasonably sets the temperature parameter according to the dataset complexity and the similarity between categories, and converts the score vector output by the teacher model into a soft label, enabling the student model to learn the similarity between categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is the architecture diagram of the image processing device in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0039] Please refer to the attached Figure 1 , the embodiments of the present invention provide a computer image processing device based on big data, including:
[0040] 1. Data processing module
[0041] In this embodiment, the main function of the data processing module is to collect image data from multiple data sources, and perform normalization processing, size adjustment, and data augmentation operations on the collected image data, specifically as follows:
[0042] 1.1. Data collection unit
[0043] In this embodiment, the data collection unit obtains image data from a variety of different types of data sources, which cover a wide range of fields and scenarios. For example, in the medical field, medical image data is collected from different imaging devices used in various departments of a hospital. These medical images have different modalities, resolutions, and imaging characteristics. In the security field, surveillance video image data is collected from various surveillance cameras installed at different locations. In addition, relevant image data can also be obtained from publicly available image datasets on the Internet, image databases in specific industries, etc., to enrich the diversity of the dataset.
[0044] During the data collection process, the images are carefully annotated. For medical images, the annotation work is completed by professional doctors or medically trained imaging technicians. The annotation content includes information such as disease type, lesion location, and severity. For example, for lung CT images, it is marked whether there are lesions such as tumors, inflammation, nodules, etc., as well as the location, size, and morphological characteristics of the lesions.
[0045] For security images, the annotators make annotations according to the surveillance scenarios and target categories, including the categories of pedestrians, vehicles, various objects, as well as the behaviors of pedestrians such as walking, running, and staying, the driving direction and status information of vehicles. The annotation information corresponds to the image data, forming a complete annotated dataset to provide accurate supervision information for subsequent model training.
[0046] Images generated by different image acquisition devices have different sizes, while deep learning models usually require a fixed-size input. For larger images, a scaling algorithm is used for proportional scaling. For smaller images, the method of edge padding or replicating edge pixels is used to pad zeros or replicate edge pixels at their edges to facilitate the calculation and processing of the deep learning model;
[0047] 1.2. Data Enhancement Unit
[0048] In this embodiment, the data enhancement unit performs enhancement operations on the collected data. To increase the diversity of the dataset and reduce the overfitting phenomenon of the model, multi-source data enhancement operations are adopted. The random rotation operation rotates the image at a certain angle between -15° and 15° to simulate different shooting angles. For example, in security surveillance, the camera may have a certain tilt angle, and rotating the image can enable the model to better adapt to this change. The horizontal flipping and vertical flipping operations increase the symmetry information of the image, enabling the model to learn the features of the target in different directions. The cropping operation cuts out 80% of the area from a random position of the image according to a preset ratio, which can make the model focus on different local areas of the image and improve the recognition ability of the target position and local features.
[0049] Adding Gaussian noise can simulate noise interference in different lighting conditions and the image acquisition process, enhancing the robustness of the model. When performing these data augmentation operations, the label information needs to be adjusted accordingly. For classification tasks, if the image undergoes rotation or flipping operations, since the content of the image itself has not changed substantially, the label remains unchanged; if the cropping operation does not lose the key feature information of the target object, the label also remains unchanged, and if the key feature information is lost, the image data is excluded. For object detection tasks, after rotation, flipping, and cropping operations, the bounding box coordinates need to be adjusted accordingly according to the image transformation, while the label remains unchanged after adding noise operation.
[0050] 2. Teacher model construction module
[0051] In this embodiment, a deep learning teacher model with pre-trained weight initialization is constructed, and the cross-entropy loss function is used to train the teacher model, as follows:
[0052] 2.1 Training unit
[0053] In this embodiment, a suitable deep learning architecture is selected as the teacher model according to the complexity of the image processing task and the data characteristics. When processing images with rich textures and complex structures (such as tissue microstructure images in medical images, high-resolution security monitoring scene images, etc.), the ResNet-101 deep learning architecture is selected. ResNet-101 consists of multiple residual blocks, and each residual block contains multiple convolutional layers, batch normalization layers, and activation function layers inside. The convolution kernel size is usually 3*3 or 1*1. In the initial stage of the network, the stride of the convolutional layer is 2, which can effectively reduce the size of the image and extract higher-level features. As the network deepens, the stride becomes 1, which is used to further refine the features and can handle long-range dependencies and complex feature hierarchies in the image;
[0054] When processing images with multi-scale targets, such as pedestrians and vehicles at different distances in security scenarios, the Inception-v4 architecture is selected. Its module contains parallel computing branches and pooling branches with convolution kernels of different sizes. The convolution kernel sizes include 1*1, 3*3, and 5*5. Through the multi-branch structure, feature information at different scales can be captured simultaneously.
[0055] When initializing the teacher model, make full use of the pre-trained weights on large-scale general image datasets. For layers similar to the target categories in the ImageNet dataset, directly adopt the corresponding pre-trained weights. For example, in medical images, if dealing with human organ images, which have certain similarities in features with some object categories in ImageNet (such as the organ structures of animals, etc.), the weights of these layers can be directly reused;
[0056] For the fully connected layer, since it is related to the number of classes in the target dataset, according to the number of classes C in the target dataset, a weight matrix W is generated using a random initialization method fc , whose dimension is (n prev , C), where n prev is the feature dimension of the output of the previous layer. After re-initializing the weights of the fully connected layer, in order to enable the fully connected layer to better adapt to the target dataset, a small learning rate n init (with a value range of 0.001 to 0.01) is used to train the fully connected layer separately for several rounds (the number of rounds ranges from 1 to 5), and then it is trained together with the entire model. This initialization method can accelerate the convergence speed of the teacher model on the target dataset
[0057] Definition of loss function: For classification tasks, the cross-entropy loss function is adopted
[0058]
[0059] where N is the size of the dataset, C is the number of classes, is the one-hot encoded label of the i-th image, is the predicted output of the teacher model for the i-th image. The cross-entropy loss function measures the difference between the predicted output of the teacher model and the true label. By minimizing this loss function, the teacher model can learn the correct image classification pattern. For object detection tasks, in addition to the classification loss, a localization loss (such as using the intersection over union (IoU) loss, etc.) needs to be added to measure the difference between the predicted bounding box and the true bounding box
[0060] Partition of the training set, validation set, and test set: Dynamically adjust the partition ratio according to the size and distribution of the dataset. When the dataset size is large (N≥10000) and the data for each class is relatively balanced, the training set, validation set, and test set are partitioned in a ratio of 7:2:1. This partitioning method can ensure that the training set has enough data for model training, the validation set is used to adjust the hyperparameters of the model, and the test set is used to evaluate the final performance of the model. When the dataset size is small (N<10000) or there is a class imbalance problem, stratified sampling is used for partitioning. For example, in a medical image dataset, if the number of cases of a certain disease is small, stratified sampling can ensure that there is a certain proportion of images of this class of cases in the validation set and the test set, so as to more accurately evaluate the model's processing ability for images of each class
[0061] Parameter settings during training: The batch size for each extraction is determined according to the model complexity and computing resources. When the model complexity is high and computing resources are sufficient (such as using a high-end GPU), the batch size can be set to 64 or 128, which can make full use of computing resources and accelerate the training speed. When the model complexity is moderate and computing resources are limited (such as using an ordinary CPU), the batch size is set to 16 or 32 to avoid problems such as memory overflow. The learning rate η adopts a dynamic adjustment strategy during training. The initial learning rate η 0 is determined according to the model architecture and dataset characteristics, and its value range is from 0.001 to 0.1. After a certain number of training epochs (the epoch interval is determined according to the model convergence situation, generally 5 to 10 epochs), if the loss function value on the validation set does not decrease for several consecutive epochs (generally 3 to 5 epochs), the learning rate is multiplied by a decay factor (the decay factor value range is from 0.1 to 0.5). This method of dynamically adjusting the learning rate can make the model converge better during training and avoid falling into local optimal solutions.
[0062] 2.2 Performance Evaluation Unit
[0063] In this embodiment, after one epoch ends, the performance of the teacher model is evaluated on the validation set. For classification tasks, the accuracy is calculated
[0064]
[0065] where TP is the number of true positives, TN is the number of true negatives, FP is the number of false positives, and FN is the number of false negatives; Recall rate:
[0066]
[0067] F1 value:
[0068]
[0069] where,
[0070]
[0071] These metrics reflect the classification performance of the model from different perspectives. For object detection tasks, in addition to calculating the classification accuracy, the mean average precision (mAP) also needs to be calculated, which is obtained by calculating the average of the areas under the precision-recall curves of different category objects. According to these evaluation metrics, hyperparameters such as the learning rate, batch size, and number of training epochs are adjusted until the teacher model achieves satisfactory performance on the validation set (such as the accuracy reaching the above in classification tasks and the mAP reaching the above in object detection tasks). Finally, the final performance of the teacher model is evaluated on the test set to ensure that the model has good generalization ability.
[0072] 3. Student Model Design Module
[0073] In this embodiment, a lightweight convolutional neural network is used as the student model, and its number of layers and parameters is lower than that of the teacher model to meet the fast inference requirements in resource-constrained environments, as follows:
[0074] 3.1. Student Model Design Unit
[0075] In this embodiment, considering the requirements for processing speed and resource consumption in actual application scenarios, a lightweight student model architecture is designed. MobileNet-v2 with an inverted residual structure and a linear bottleneck layer is adopted. In the inverted residual structure, each bottleneck layer first performs dimensionality increase through a 1×1 convolution to increase the number of channels, so that richer feature information can be extracted. Then, depthwise separable convolution is used for feature extraction. Depthwise separable convolution includes depthwise convolution and pointwise convolution. Depthwise convolution performs convolution operations on each input channel separately, which can greatly reduce the computational amount and effectively capture the spatial features of the image. Pointwise convolution fuses the channels of the output after depthwise convolution through a 1×1 convolution to further integrate the feature information. Finally, dimensionality reduction is performed through a 1×1 convolution to restore to an appropriate number of channels. The linear bottleneck layer means that no activation function is used after the last 1×1 convolution, which can retain more feature information, avoid over-compressing the features, and is beneficial to improving the performance of the model. This lightweight structure enables the student model to run quickly on resource-constrained devices while maintaining a certain feature extraction ability.
[0076] 3.2. Student Model Initialization Unit
[0077] In this embodiment, random initialization or some pre-trained weight initialization methods are used, such as using the weights of a general image feature extraction model trained on a small data set to initialize the first few layers of the student model. Two initialization methods can be used. One is random initialization, that is, randomly generating the initial parameters of the model. The other is to use the weights of a model with similar task characteristics trained on a small data set for initialization. For example, when processing brain MRI images in medical images, if there is a model trained on other small brain image data sets, its weight can be used as an initialization reference. First, feature analysis is performed on the small data set and the target data set to determine the similarity between the two in terms of image content (such as brain tissue structure, lesion type, etc.), target category (such as normal tissue, lesion type, etc.), and feature distribution (such as grayscale distribution, texture features, etc.). For layers with similar features, the weights of the corresponding layers in the small data set training model are directly copied to the student model. For layers with large differences (such as fully connected layers or convolutional layers related to specific features of the target data set), a random initialization method is used, and the range of the initialization parameters is determined according to the input and output dimensions of the layer. For example, for a fully connected layer with an input dimension of m and an output dimension of n, the initial values of the elements of the weight matrix are to This initialization method enables the student model to have a certain perception of the target image features in the initial stage, which helps to speed up the training and improve the model performance.
[0078] 4. Knowledge distillation training module
[0079] In this embodiment, the main function of the knowledge distillation training module is to transfer knowledge from the trained teacher model to the student model. Specifically, it inputs the images in the training data set into the teacher model and quantifies the difference between the teacher model output and the student model output by calculating the distillation loss. Then, the gradients are calculated with the help of the back-propagation algorithm, and the parameters of the student model are updated according to these gradients, as follows:
[0080] 4.1. Soft label generation unit
[0081] In this embodiment, the images in the training data set are input into the trained teacher model to generate soft labels. The teacher model performs i Output a score vector Its dimension is C (number of categories). This score vector is calculated by the last convolutional layer and the fully connected layer of the teacher model. For classification tasks, each element of the score vector represents the score of the corresponding category. Then the score vector is converted into a soft label by the softmax function combined with the temperature parameter T:
[0082]
[0083] The temperature parameter T plays a crucial role in knowledge distillation. When T = 1, the softmax function outputs a normal probability distribution; when T > 1, the distribution of the soft labels is smoother, and the probability differences between different classes become smaller. For example, in medical images, for images of some similar diseases (such as different types of pneumonia, which may have similar manifestations in imaging), the soft labels at high temperatures can enable the student model to learn the similarities in features of these diseases, thus better performing classification. The value of the temperature parameter T is determined according to the complexity of the dataset and the similarity between classes. When the dataset contains a large number of classes (C ≥ 50) and the feature similarity between classes is high (by calculating the feature distance between classes, such as the Euclidean distance, if the average distance is less than a preset threshold), the value range of T is from 2 to 10; when the dataset has fewer classes (C < 50) and the differences between classes are obvious, the value range of T is from 1.5 to 5.
[0084] 4.2. Loss Calculation Unit
[0085] In this embodiment, the image x i is simultaneously input into the student model, and the probability distribution P student (y|x i ) output by the student model is obtained. The KL divergence is used to measure the difference between the output of the student model and the output of the teacher model as the distillation loss:
[0086]
[0087] When calculating the distillation loss, if the probability value of a certain class in the probability distribution P student (y|x i ) output by the student model is close to 0, to avoid numerical instability problems when calculating the KL divergence, a small smoothing value ∈ (the value range is from 10 -6 to 10 -8 ) is added to it, that is
[0088] P student (y = j|x i ) = max(P student (y = j|x i ), ∈)
[0089] The KL divergence is a measure method for measuring the difference between two probability distributions. It prompts the student model to learn the soft label distribution output by the teacher model, so that the student model can imitate the behavior of the teacher model and obtain knowledge from the teacher model.
[0090] 5. Deployment Module
[0091] In this embodiment, the deployment module is responsible for deploying the trained student model to the corresponding device, enabling the model to process newly input images in the actual environment and output the processing results, as follows:
[0092] 5.1. Model Inference Unit
[0093] In this embodiment, the trained student model is loaded from the storage medium, and corresponding space is allocated in the memory according to the above-mentioned student model structure parameters (such as the convolution kernel size and number of channels of each layer), and the model weight values obtained from training are loaded into the corresponding parameter positions. At the same time, some auxiliary data structures related to model inference are initialized, such as a buffer for storing intermediate calculation results. In a medical diagnostic device, ensure that the model loading process is compatible with the device's operating system and other running medical applications to avoid memory conflicts or other abnormal situations. On a security monitoring device, the model loading should meet the real-time requirements and minimize the startup time as much as possible so as to quickly engage in image analysis work;
[0094] The preprocessed image is sent to the student model for inference calculation: taking the student model as an example, the image data first enters the bottleneck layer of the inverted residual structure. In the bottleneck layer, a dimensionality increase operation is performed through a 1*1 convolution to increase the number of channels to extract richer feature information. Then, depthwise separable convolution is used for feature extraction. The depth convolution operates on each input channel separately, and the pointwise convolution fuses the outputs after the depth convolution. During the calculation process, make full use of the device's computing resources. For example, on a device with a GPU, the computing tasks are distributed to multiple computing cores of the GPU for parallel execution to accelerate the inference speed. On a device without a GPU, efficient CPU computing is achieved through optimized algorithms. The model sequentially extracts and processes the features of the image according to its structure, and after multiple layers of calculation, finally outputs the inference result. In medical image analysis, relevant information about disease diagnosis is output; in security image analysis, the results of object detection and behavior recognition are output.
[0095] 5.2. Update and Optimization Unit
[0096] In this embodiment, for the teacher model, the new data is integrated with the original large-scale dataset, and the training set, validation set, and test set are re-partitioned (the partitioning method can be adjusted according to the scale and class distribution of the new dataset). Then, the integrated dataset is used to retrain the teacher model. The training process is similar to that of the previous teacher model training, including selecting an appropriate loss function, adjusting training parameters, etc. After the teacher model training is completed and its performance is improved, the newly trained teacher model is used for knowledge distillation to train the student model. For the student model, knowledge distillation training is carried out using the updated teacher model. The new image data is input into the teacher model to calculate the distillation loss, and the gradient is calculated through the backpropagation algorithm to update the parameters of the student model, enabling the student model to learn new data features and the knowledge of the teacher model, thereby continuously improving the performance of the model in new data and new scenarios.
[0097] Working principle: The data processing module collects image data from multiple sources such as medical department imaging devices, security monitoring cameras, and publicly available network datasets. The data collection unit annotates the images in detail and makes them meet the model input requirements by using methods such as scaling, padding, or replicating edge pixels for images of different sizes. The data augmentation unit increases data diversity through operations such as rotation, flipping, cropping, and adding Gaussian noise, and at the same time adjusts the label information according to the task type and operation conditions;
[0098] The teacher model construction module selects an architecture according to the characteristics of the images. For example, ResNet-101 is used to process images with complex textures, and Inception-v4 is used for multi-scale target images. Some layers are initialized with pre-trained weights, and the fully connected layer is randomly initialized and then trained with a small learning rate separately. Training is carried out through cross-entropy loss and combined with localization loss. The dataset is partitioned according to the scale and distribution of the dataset, the training parameters are dynamically adjusted, and the hyperparameters are adjusted according to the performance evaluation on the validation set. Finally, the generalization ability is evaluated on the test set;
[0099] The student model design module adopts the lightweight architecture of MobileNet-v2, reduces the computational amount through specific convolutional operations, and is initialized with random initialization or the weights of a model for a similar task based on a small dataset;
[0100] The knowledge distillation training module inputs the images into the teacher model to generate soft labels, adjusts the soft label distribution in combination with the temperature parameter, and then inputs the images into the student model. The distillation loss is calculated using KL divergence (a smoothing value is added to avoid instability), and the parameters of the student model are updated through backpropagation;
[0101] The model inference unit of the deployment module loads the student model and processes the new input image, and outputs the result. The update and optimization unit integrates the new data with the original dataset, retrains the teacher model, and then performs knowledge distillation training on the student model with the new teacher model to improve the performance. First, a brief summary of each module is given, including the data processing module that collects and augments image data, the teacher model construction module that selects the architecture and initialization method and conducts training, the student model design module that adopts a lightweight architecture and initialization method, the knowledge distillation training module that transfers knowledge, and the deployment module that performs model inference and update and optimization. During the description, the key roles and features of each module are highlighted, and at the same time, attention is paid to controlling the number of words to avoid overly detailed technical elaboration, and the overall working principle is presented in a concise and clear manner.
[0102] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A computer image processing device based on big data, characterized in that: include: A data processing module collects image data from multiple data sources and performs normalization, size adjustment and data enhancement operations on the collected image data; The teacher model building module is used to build a deep learning teacher model with pre-trained weight initialization and train the teacher model using the cross entropy loss function; The student model design module is used to design a lightweight student model architecture and initialize it by random initialization or using the model weights trained with a small dataset of similar tasks; The knowledge distillation training module inputs the images in the training dataset into the trained teacher model, calculates the distillation loss, calculates the gradient through the back-propagation algorithm, and updates the parameters of the student model according to the gradient; The deployment module deploys the trained student model to the device, processes the newly input images, and outputs the processing results. At the same time, this module can collect new image data to update the teacher model and student model.
2. A computer image processing device based on big data according to claim 1, characterized in that: The data processing module comprises: A data collection unit collects image data from multiple data sources and performs normalization and size adjustment on the collected image data; The data augmentation unit performs augmentation operations including random rotation, flipping and cropping on the collected data, and adjusts the label information accordingly during the data augmentation process to maintain the consistency of the image and the label.
3. The computer image processing device based on big data according to claim 1, characterized in that: The teacher model building module includes: Feature capture unit, which selects the deep learning teacher model architecture and captures the feature information of images of different scales at the same time; Initialization unit, initializes the weights of the similarity layer and the fully connected layer; The training unit defines the loss function of the teacher model and dynamically adjusts the data set into training, validation, and test sets according to the size and distribution of the data set. The performance evaluation unit evaluates and adjusts the performance of the teacher model on the validation set.
4. The computer image processing device based on big data according to claim 1, characterized in that: The student model design module includes: The student model design unit uses an inverted residual structure and a linear bottleneck layer to design a lightweight student model architecture; The student model initialization unit uses random initialization, that is, randomly generating the initial parameters of the model and using the model weights with similar task characteristics trained on a small data set to initialize the model.
5. The computer image processing device based on big data according to claim 1, characterized in that: The knowledge distillation training module includes: The soft label generation unit inputs each image in the training data set into the trained teacher model to obtain the output probability distribution of the teacher model, which is the soft label; The loss calculation unit inputs the image into the student model, obtains the output probability distribution of the student model, and uses KL divergence to measure the difference between the student model output and the teacher model output.
6. The computer image processing device based on big data according to claim 1, characterized in that: The deployment module includes: The model inference unit deploys the trained student model to a computer image processing device based on big data; Update the optimization unit, regularly collect new image data and corresponding annotations, and update the training of the teacher model and the student model.
7. The computer image processing device based on big data according to claim 2, characterized in that: The data enhancement unit also includes a noise adding unit for simulating image changes under different lighting and environmental conditions by adding Gaussian noise.
8. The computer image processing device based on big data according to claim 3, characterized in that: The architecture of the deep learning teacher model in the feature capture unit is selected from one of ResNet-101 and Inception-v4, and the weights of the fully connected layer are reinitialized according to the number of categories of the target data set.
9. The computer image processing device based on big data according to claim 4, characterized in that: The lightweight student model architecture in the student model design unit is MobileNet-v2, which uses an inverted residual structure and a linear bottleneck layer to reduce the amount of calculation.
10. The computer image processing device based on big data according to claim 6, characterized in that: After outputting the processing results, the deployment module directly outputs category information for image classification tasks and outputs category and bounding box coordinate information for object detection tasks, and further performs analysis and decision making based on the output results.