Digital identification method and system for router, gateway, IPC and ONU

By building a lightweight digital recognition model and real-time image processing technology, the problem of excessive model parameters and low recognition accuracy on edge computing devices is solved, and efficient digital recognition and secure log transmission in dynamic lighting environments are achieved.

CN120356068APending Publication Date: 2025-07-22FUJIAN NEWLAND COMM SCI TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510269176.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When the prior art performs digital recognition on edge computing devices, the model parameters are too large and difficult to deploy, and the recognition accuracy rate is low and the error recognition rate is high, which cannot meet the needs of resource-constrained devices.

Method used

A digital recognition model based on the local feature extraction module and the global feature extraction module is constructed, deep separable convolution and expansion convolution are adopted, combined with knowledge distillation technology and dynamic pruning technology for compression, and light compensation and image enhancement are performed through CMOS sensors, and real-time recording and encryption identification log uploads are recorded and encrypted.

Benefits of technology

It realizes the lightweight deployment of digital recognition models on resource-constrained edge computing devices, improves recognition accuracy and image quality, reduces noise data, and enhances the security and traceability of identification logs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356068A_ABST
    Figure CN120356068A_ABST
Patent Text Reader

Abstract

The invention provides a digital recognition method and system for a router, a gateway, an IPC and an ONU in the technical field of edge computing. The method comprises the steps that S1, a large number of historical digital images are acquired to construct a data set; s2, creating a digital recognition model and setting a loss function; s3, training the digital recognition model through the data set and the loss function, and deploying the trained digital recognition model to an edge computing device of which the device type is a router, a gateway, an IPC or an ONU; s4, acquiring a real-time digital image by the edge computing equipment through a camera, and performing preprocessing at least including illumination compensation, image enhancement, dynamic noise reduction, ROI optimization cutting, graying and normalization on the real-time digital image to obtain an optimized digital image; and S5, the edge computing device inputs the optimized digital image into the digital recognition model to obtain a digital recognition result. The method has the advantage that the accuracy of digital recognition is improved in a lightweight manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of edge computing, and particularly to a digital recognition method and system for routers, gateways, IP cameras (IPCs), and optical network units (ONUs). Background Art

[0002] With the rapid development of technologies such as the Internet of Things and smart hardware, lightweight digital recognition technology has received increasing attention and applications, resulting in the need for digital recognition on resource-constrained edge computing devices (such as routers, gateways, IPCs (network cameras), and ONUs). For example: in the financial field, it is necessary to recognize digital information on bank bills to achieve automatic processing and information entry of bills, improving work efficiency and accuracy; in the transportation field, it is necessary to recognize digital information on license plates and digital numbers on road signs, such as speed limit signs; in the logistics and express delivery fields, it is necessary to recognize digital information on express delivery forms, such as waybills, to improve logistics processing efficiency; in the industrial and manufacturing fields, it is necessary to recognize digital identifiers on products for quality control and automated inspection; in the education field, it is necessary to recognize handwritten numbers on answer sheets for automatic marking of multiple-choice questions; in the medical field, it is necessary to recognize handwritten digital information by doctors, such as prescription dosages.

[0003] For digital recognition, traditionally, the CNN model is generally directly trained and then used for digital recognition without corresponding optimization of the CNN model, resulting in the parameter quantity of the CNN model generally exceeding 5MB, making it difficult to be deployed on edge computing devices with less than 32MB of memory; although there are some lightweight models traditionally, such as MobileNet, while reducing the parameter quantity, the feature extraction ability of MobileNet also decreases by about 15%-20%, unable to meet the requirements of digital recognition; and the recognition accuracy of traditional models fluctuates by more than 15 percentage points in a dynamic lighting environment, especially the misrecognition rate of the shallow network structure adopted is as high as 18.7% under complex background interference.

[0004] Therefore, how to provide a digital recognition method and system for routers, gateways, IPCs, and ONUs to achieve lightweight improvement in digital recognition accuracy has become an urgent technical problem to be solved. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a digital recognition method and system for routers, gateways, IPCs, and ONUs to achieve lightweight improvement in digital recognition accuracy.

[0006] In a first aspect, the present invention provides a digital recognition method for routers, gateways, IPCs, and ONUs, including the following steps:

[0007] Step S1: Obtain a large number of historical digital images under different scenarios and different lighting conditions, perform preprocessing on each of the historical digital images, including at least noise reduction, cropping, size unification, grayscale conversion, and normalization, and construct a dataset after annotating the numbers in the preprocessed historical digital images;

[0008] Step S2: Create a digital recognition model based on a local feature extraction module, a global feature extraction module, a feature fusion module, and a prediction output module, and set the loss function of the digital recognition model;

[0009] Step S3: Train the digital recognition model through the dataset and the loss function, and deploy the trained digital recognition model to an edge computing device of device type router, gateway, IPC, or ONU;

[0010] Step S4: The edge computing device collects real-time digital images through a camera, and performs preprocessing on the real-time digital images, including at least light compensation, image enhancement, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, to obtain optimized digital images;

[0011] Step S5: The edge computing device inputs the optimized digital image into the digital recognition model to obtain a digital recognition result, and displays the digital recognition result in real time through a display screen;

[0012] Step S6: Real-time record the recognition log including at least the optimized digital image, the digital recognition result, and the recognition time, encrypt the recognition log into an encrypted log, upload the encrypted log to the server for storage, and delete the encrypted log locally.

[0013] Further, in step S2, the local feature extraction module is used to extract local spatial features from digital images, and is constructed based on three depthwise separable convolutions. Each depthwise separable convolution is constructed based on a depth convolution and a pointwise convolution. The depth convolution is used to perform convolution operations independently on each channel of the input digital image to extract local spatial sub-features. The pointwise convolution is used to perform a linear combination of each local spatial sub-feature through a 1×1 convolution kernel to restore the interaction between channels and generate higher-level local spatial features;

[0014] The global feature extraction module is used to extract global spatial features from digital images and is constructed based on dilated convolutions, with the dilation rate taking a value of 2;

[0015] The feature fusion module is used to fuse local spatial features and global spatial features to obtain fused features and is constructed based on a channel attention mechanism;

[0016] The prediction output module is used to output the digital recognition result, which is constructed based on the global average pooling layer and the Softmax classifier, and the value of the temperature scaling factor of the Softmax classifier is 1.5;

[0017] The loss function uses the cross-entropy loss function.

[0018] Further, step S3 is specifically as follows:

[0019] Based on a preset segmentation ratio, the data set is divided into a training set, a validation set, and a test set. The digital recognition model is trained through the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, the digital recognition model is compressed through knowledge distillation technology and dynamic pruning technology; the trained digital recognition model is verified through the validation set to determine whether the recognition accuracy is greater than a preset accuracy threshold. If not, the verification fails, and the training set is expanded and training continues. If so, the verification is successful; the successfully verified digital recognition model is tested through the test set to determine whether the confidence level is greater than a preset confidence level threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test is successful, and the training ends;

[0020] The trained digital recognition model is deployed to edge computing devices of device types such as routers, gateways, IP cameras, or ONUs.

[0021] Further, step S4 is specifically as follows:

[0022] The edge computing device collects real-time digital images through a CMOS sensor with automatic white balance and exposure compensation, and performs preprocessing on the real-time digital images, including at least light compensation, image enhancement, dynamic noise reduction, ROI optimization and cropping, grayscale conversion, and normalization to obtain optimized digital images;

[0023] The light compensation adopts a hybrid strategy of HSV color space histogram equalization and local contrast enhancement; the image enhancement is based on the CLAHE algorithm to enhance the contrast of the image and suppress noise; the dynamic noise reduction is based on a hybrid morphological filtering algorithm.

[0024] Further, step S6 is specifically as follows:

[0025] The recognition log including at least the optimized digital images, digital recognition results, and recognition time is recorded in real time. The device serial number of the edge computing device itself is obtained, and image data and text data are extracted from the recognition log through a media type detection algorithm;

[0026] The image data is compressed by the ZSTD algorithm to obtain image compressed data, and the text data is compressed by the Brotli algorithm to obtain text compressed data;

[0027] The first MAC value of the image compressed data and the device serial number is calculated by the HMAC algorithm, and the first encrypted data is obtained by encrypting the first MAC value and the image compressed data through the AES-256 algorithm;

[0028] The second MAC value of the text compressed data and the device serial number is calculated by the HMAC algorithm, and the second encrypted data is obtained by encrypting the second MAC value and the text compressed data through the SM9 algorithm;

[0029] The signature value of the recognition log is calculated by the EdDSA algorithm;

[0030] The first encrypted data, the second encrypted data, the signature value, and the device serial number are encapsulated into an encrypted log through the AES algorithm, and the encrypted log is uploaded to the server for storage in real time through the TLS protocol, and the encrypted log on the local side is cleared.

[0031] In a second aspect, the present invention provides a digital recognition system for routers, gateways, IP cameras, and ONUs, including the following modules:

[0032] A dataset construction module, configured to obtain a large number of historical digital images under different scenarios and different lighting conditions, perform preprocessing on each of the historical digital images at least including noise reduction, cropping, size unification, grayscale conversion, and normalization, and construct a dataset after labeling the numbers in the preprocessed historical digital images;

[0033] A digital recognition model creation module, configured to create a digital recognition model based on a local feature extraction module, a global feature extraction module, a feature fusion module, and a prediction output module, and set a loss function of the digital recognition model;

[0034] A digital recognition model training module, configured to train the digital recognition model through the dataset and the loss function, and deploy the trained digital recognition model to an edge computing device of a device type of a router, a gateway, an IP camera, or an ONU;

[0035] A real-time digital image acquisition module, configured to enable an edge computing device to collect real-time digital images through a camera, and perform preprocessing on the real-time digital images at least including light compensation, image enhancement, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization to obtain optimized digital images;

[0036] The digital recognition module is used for the edge computing device to input the optimized digital image into the digital recognition model, obtain the digital recognition result, and display the digital recognition result in real time through the display screen;

[0037] The recognition log management module is used to record in real time the recognition log including at least the optimized digital image, the digital recognition result, and the recognition time, encrypt the recognition log into an encrypted log, upload the encrypted log to the server for storage, and delete the encrypted log locally.

[0038] Further, in the digital recognition model creation module, the local feature extraction module is used to extract local spatial features from the digital image, and is constructed based on three depthwise separable convolutions. Each depthwise separable convolution is constructed based on a depth convolution and a pointwise convolution; the depth convolution is used to perform convolution operations independently on each channel of the input digital image to extract local spatial sub-features; the pointwise convolution is used to perform a linear combination of each local spatial sub-feature through a 1×1 convolution kernel to restore the interaction between channels and generate higher-level local spatial features;

[0039] The global feature extraction module is used to extract global spatial features from the digital image and is constructed based on dilated convolutions, with the dilation rate taking a value of 2;

[0040] The feature fusion module is used to fuse the local spatial features and the global spatial features to obtain the fused features and is constructed based on the channel attention mechanism;

[0041] The prediction output module is used to output the digital recognition result, and is constructed based on the global average pooling layer and the Softmax classifier, and the temperature scaling factor of the Softmax classifier takes a value of 1.5;

[0042] The loss function uses the cross-entropy loss function.

[0043] Further, the digital recognition model training module is specifically used for:

[0044] Divide the dataset into a training set, a validation set, and a test set based on a preset splitting ratio. Train the digital recognition model using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, compress the digital recognition model using knowledge distillation technology and dynamic pruning technology. Verify the trained digital recognition model using the validation set to determine whether the recognition accuracy is greater than a preset accuracy threshold. If not, the verification fails, and the training set is expanded and training continues. If so, the verification succeeds. Test the digitally recognized model that has passed the verification using the test set to determine whether the confidence level is greater than a preset confidence threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test succeeds and training ends;

[0045] Deploy the trained digital recognition model to edge computing devices of device types including routers, gateways, IP cameras, or ONUs.

[0046] Further, the real-time digital image acquisition module is specifically used for:

[0047] The edge computing device acquires real-time digital images through a CMOS sensor with automatic white balance and exposure compensation, and performs preprocessing on the real-time digital images, including at least light compensation, image enhancement, dynamic noise reduction, ROI optimization and cropping, grayscale conversion, and normalization, to obtain optimized digital images;

[0048] The light compensation adopts a hybrid strategy of HSV color space histogram equalization and local contrast enhancement; the image enhancement is based on the CLAHE algorithm to enhance the contrast of the image and suppress noise; the dynamic noise reduction is based on a hybrid morphological filtering algorithm.

[0049] Further, the recognition log management module is specifically used for:

[0050] Real-time record recognition logs including at least the optimized digital images, digital recognition results, and recognition times, obtain the device serial number of the edge computing device itself, and extract image data and text data from the recognition logs through a media type detection algorithm;

[0051] Compress the image data using the ZSTD algorithm to obtain image compressed data, and compress the text data using the Brotli algorithm to obtain text compressed data;

[0052] Calculate the first MAC value of the image compressed data and the device serial number through the HMAC algorithm, and encrypt the first MAC value and the image compressed data using the AES-256 algorithm to obtain the first encrypted data;

[0053] Calculate the second MAC value of the text compression data and the device serial number through the HMAC algorithm, and encrypt the second MAC value and the text compression data through the SM9 algorithm to obtain the second encrypted data;

[0054] Calculate the signature value of the recognition log through the EdDSA algorithm;

[0055] Package the first encrypted data, the second encrypted data, the signature value, and the device serial number into an encrypted log through the AES algorithm, upload the encrypted log to the server for storage in real time through the TLS protocol, and clear the encrypted log locally.

[0056] The advantages of the present invention are as follows:

[0057] 1. By obtaining a large number of historical digital images under different scenarios and lighting conditions, preprocessing and annotating each historical digital image with at least noise reduction, cropping, size unification, grayscale conversion, and normalization to construct a dataset; then creating a digital recognition model based on a local feature extraction module, a global feature extraction module, a feature fusion module, and a prediction output module, setting the loss function of the digital recognition model, training the digital recognition model through the dataset and the loss function, and deploying the trained digital recognition model to an edge computing device with a device type of router, gateway, IPC, or ONU; the edge computing device captures real-time digital images through a camera, preprocesses the real-time digital images with at least lighting compensation, image enhancement, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization to obtain optimized digital images, inputs the optimized digital images into the digital recognition model to obtain digital recognition results, records recognition logs including at least the optimized digital images, digital recognition results, and recognition times in real time, encrypts the recognition logs into encrypted logs and uploads them to the server for storage, and deletes the local encrypted logs; since the local feature extraction module is constructed based on depthwise separable convolution and the global feature extraction module is constructed based on dilated convolution, both depthwise separable convolution and dilated convolution are lightweight network structures, and the digital recognition model is compressed through knowledge distillation technology and dynamic pruning technology during the training process, effectively reducing the number of parameters (complexity) and volume of the digital recognition model, and locally deleting the uploaded encrypted logs in real time to reduce storage overhead, enabling the digital recognition model to be easily deployed on edge computing devices with limited computing power and storage resources; by fusing the local spatial features and global spatial features extracted by the local feature extraction module and the global feature extraction module, that is, introducing a channel attention mechanism through the feature fusion module to achieve an optimized combination of multi-scale features, effectively improving the feature extraction ability of the digital recognition model, combined with a CMOS sensor with automatic white balance and exposure compensation to capture real-time digital images, preprocessing the real-time digital images with at least lighting compensation, image enhancement, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, effectively improving the image quality, facilitating the recognition of the digital recognition model, and finally achieving lightweight improvement of the accuracy of digital recognition.

[0058] 2. By preprocessing each historical digital image with at least noise reduction, cropping, size unification, grayscale conversion, and normalization, the quality of the dataset is effectively improved, thereby improving the training effect of the digital recognition model.

[0059] 3. By recording recognition logs including at least optimized digital images, digital recognition results, and recognition times in real time, encrypting the recognition logs into encrypted logs and uploading them to the server for storage, it is convenient for later traceability.

[0060] 4. Train the digital recognition model using the training set until the loss value of the loss function is less than the preset loss threshold. During the training process, compress the digital recognition model using knowledge distillation technology and dynamic pruning technology. Then, calculate the accuracy through the validation set to verify the trained digital recognition model, and calculate the confidence through the test set to test the digital recognition model that has passed the verification. That is, during the training process of the digital recognition model, continuous compression, verification, and testing are performed to effectively balance the model volume and recognition accuracy of the digital recognition model.

[0061] 5. By performing preprocessing on the real-time digital image, including at least illumination compensation, image enhancement, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, effectively reduce the noise data carried in the real-time digital image, thereby greatly improving the accuracy of digital recognition.

[0062] 6. Obtain the device serial number of the edge computing device itself and extract image data and text data from the recognition log; compress the image data using the ZSTD algorithm to obtain compressed image data, and compress the text data using the Brotli algorithm to obtain compressed text data; calculate the first MAC value of the compressed image data and the device serial number using the HMAC algorithm, and encrypt the first MAC value and the compressed image data using the AES-256 algorithm to obtain the first encrypted data; calculate the second MAC value of the compressed text data and the device serial number using the HMAC algorithm, and encrypt the second MAC value and the compressed text data using the SM9 algorithm to obtain the second encrypted data; calculate the signature value of the recognition log using the EdDSA algorithm; encapsulate the first encrypted data, the second encrypted data, the signature value, and the device serial number into an encrypted log using the AES algorithm, and upload the encrypted log to the server for storage in real time through the TLS protocol; that is, perform different encryptions on the recognition log based on the data type (image data and text data), and finally merge the encryptions. Moreover, the image data and text data respectively combine different encryption algorithms, and the TLS protocol is a secure transmission protocol. At least 8 security measures are taken before and after (device serial number, first MAC value, AES-256 algorithm, second MAC value, SM9 algorithm, EdDSA algorithm, AES algorithm, TLS protocol) to prevent the recognition log from being stolen and tampered with in plain text, and combined with compression algorithms, thereby greatly improving the security of the transmission and storage of the recognition log, effectively improving the reliability of traceability, and greatly improving the transmission efficiency of the recognition log. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The following further describes the present invention with reference to the accompanying drawings and embodiments.

[0064] Figure 1 It is a flowchart of a digital recognition method for routers, gateways, IP cameras, and ONUs according to the present invention.

[0065] Figure 2 This is a schematic structural diagram of a digital recognition system for routers, gateways, IP cameras, and ONUs according to the present invention. Specific implementation manners

[0066] The overall idea of the technical solution in the embodiment of this application is as follows: Digital recognition is performed through a digital recognition model constructed by a local feature extraction module, a global feature extraction module, a feature fusion module, and a prediction output module. Since the local feature extraction module is constructed based on depthwise separable convolution and the global feature extraction module is constructed based on dilated convolution, both depthwise separable convolution and dilated convolution are lightweight network structures. And during the training process of the digital recognition model, compression is performed through knowledge distillation technology and dynamic pruning technology. Encrypted logs that have been uploaded are deleted locally in real time to reduce storage overhead, so that the digital recognition model can be easily deployed on edge computing devices with limited computing power resources and storage resources; the feature fusion module introduces a channel attention mechanism to realize the optimal combination of multi-scale features (local spatial features and global spatial features), effectively improving the feature extraction ability of the digital recognition model. Combined with a CMOS sensor with automatic white balance and exposure compensation to collect real-time digital images, preprocessing of the real-time digital images is performed, including at least light compensation, image enhancement, dynamic noise reduction, ROI optimization and cropping, grayscale conversion, and normalization, effectively improving the image quality to achieve lightweight improvement of the accuracy of digital recognition.

[0067] Please refer to Figures 1 to 2 As shown, a preferred embodiment of a digital recognition method for routers, gateways, IP cameras, and ONUs according to the present invention includes the following steps:

[0068] Step S1: Obtain a large number of historical digital images under different scenarios and different lighting conditions, perform preprocessing on each of the historical digital images, including at least noise reduction, cropping, size unification, grayscale conversion, and normalization, and label the numbers in the preprocessed historical digital images to construct a data set;

[0069] By performing preprocessing on each historical digital image, including at least noise reduction, cropping, size unification, grayscale conversion, and normalization, the quality of the data set is effectively improved, and thus the training effect of the digital recognition model is improved.

[0070] Step S2: Create a digital recognition model based on a local feature extraction module, a global feature extraction module, a feature fusion module, and a prediction output module, and set the loss function of the digital recognition model;

[0071] Step S3: Train the digital recognition model through the data set and the loss function, and deploy the trained digital recognition model to an edge computing device of a device type of router, gateway, IP camera, or ONU;

[0072] Step S4: The edge computing device acquires real-time digital images through a camera, and performs preprocessing on the real-time digital images, including at least illumination compensation, image enhancement, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, to obtain optimized digital images;

[0073] By performing preprocessing on the real-time digital images, including at least illumination compensation, image enhancement, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, the noise data carried in the real-time digital images is effectively reduced, thereby greatly improving the accuracy of digital recognition.

[0074] Step S5: The edge computing device inputs the optimized digital images into a digital recognition model to obtain a digital recognition result, and the digital recognition result is displayed in real time through a display screen;

[0075] Step S6: The recognition log including at least the optimized digital images, digital recognition results, and recognition time is recorded in real time, the recognition log is encrypted into an encrypted log, the encrypted log is uploaded to a server for storage, and the encrypted log on the local side is cleared.

[0076] By recording in real time the recognition log including at least the optimized digital images, digital recognition results, and recognition time, encrypting the recognition log into an encrypted log and uploading it to the server for storage, it is convenient for later traceability.

[0077] In the step S2, the local feature extraction module is used to extract local spatial features from digital images. It is constructed based on three depthwise separable convolutions. Each depthwise separable convolution is constructed based on a depthwise convolution and a pointwise convolution. By decomposing the standard convolution into two independent steps, the amount of calculation and the number of parameters are significantly reduced, while maintaining good feature extraction capabilities. The depthwise convolution is used to perform convolution operations independently on each channel of the input digital image to extract local spatial sub-features. The pointwise convolution is used to perform a linear combination of each local spatial sub-feature through a 1×1 convolution kernel to restore the interaction between channels and generate higher-level local spatial features;

[0078] The global feature extraction module is used to extract global spatial features from digital images. It is constructed based on dilated convolutions, and the dilation rate takes a value of 2. When the dilation rate takes a value of 2, the convolution kernel samples once every 1 pixel, which is used to enable the convolution kernel to cover a larger receptive field without increasing additional parameters;

[0079] The depthwise separable convolution is mainly used to reduce the computational complexity and the number of parameters, and is suitable for scenarios that require lightweight models, such as mobile devices and edge computing; the dilated convolution is mainly used to expand the receptive field, capture context information in a larger range (global spatial features), and at the same time maintain the spatial resolution of the feature map (digital image), and is suitable for tasks that require capturing long-range dependencies, such as semantic segmentation and time series analysis.

[0080] By constructing a parallel structure of depthwise separable convolution and dilated convolution (dual-branch feature extraction network), the recognition accuracy is maintained while the number of parameters is reduced to 2.3MB.

[0081] The feature fusion module is used to fuse local spatial features and global spatial features to obtain fused features, and is constructed based on the channel attention mechanism (Squeeze-and-Excitation Module, SE module); the channel attention mechanism dynamically adjusts the importance of each channel in the feature map by explicitly modeling the dependencies between channels, thereby improving the feature expression ability and performance of the model. The formula is:

[0082] F = α × F1 + (1 - α) × F2;

[0083] where F represents the fused feature; F1 represents the local spatial feature; F2 represents the global spatial feature; α represents the fusion weight, α ∈ [0, 1], and is automatically learned by the SE module;

[0084] The prediction output module is used to output the digital recognition result, and is constructed based on the global average pooling layer (Global Average Pooling, GAP) and the Softmax classifier, and the value of the temperature scaling factor of the Softmax classifier is 1.5; the temperature scaling factor is a technique used to adjust the output probability distribution of a neural network model, usually applied in the Softmax classifier to calibrate the prediction confidence of the model; by replacing the traditional fully connected layer with the global average pooling layer, the number of parameters can be effectively reduced, overfitting can be prevented, the robustness of the model can be enhanced, and the network structure can be simplified;

[0085] The loss function uses the cross-entropy loss function.

[0086] The cross-entropy loss function measures the loss by calculating the difference between the probability distribution predicted by the model and the probability distribution of the true labels; for image classification tasks, especially digit recognition, the output of the model is usually a probability distribution representing the confidence of the input image belonging to each category; the cross-entropy loss function provides more direct gradient information for the model, which helps to converge faster during training and gives higher penalties to misclassified predictions, enabling the model to better learn the distinctions between categories; that is, the cross-entropy loss function can effectively measure the difference between the predicted probability distribution and the true distribution and accelerate the convergence of the model.

[0087] The specific steps of step S3 are as follows:

[0088] Based on a preset segmentation ratio, the dataset is divided into a training set, a validation set, and a test set. The digit recognition model is trained using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, the digit recognition model is compressed using knowledge distillation technology and dynamic pruning technology; the trained digit recognition model is validated using the validation set to determine whether the recognition accuracy is greater than a preset accuracy threshold. If not, the validation fails, and the training set is expanded and training continues. If so, the validation is successful; the successfully validated digit recognition model is tested using the test set to determine whether the confidence is greater than a preset confidence threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test is successful, and training ends;

[0089] The trained digit recognition model is deployed to edge computing devices with device types of routers, gateways, IP cameras, or ONUs.

[0090] The digit recognition model is trained using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, the digit recognition model is compressed using knowledge distillation technology and dynamic pruning technology. Then, the accuracy is calculated using the validation set to validate the trained digit recognition model, and the confidence is calculated using the test set to test the successfully validated digit recognition model. That is, during the training process of the digit recognition model, compression, validation, and testing are continuously performed to effectively balance the model size and recognition accuracy of the digit recognition model.

[0091] Knowledge distillation technology is a machine learning model compression method aimed at transferring the knowledge of a large model to a small model to improve the model performance and generalization ability; the core idea of knowledge distillation is to transform the knowledge of a complex model into a more concise and effective representation, enabling it to maintain high performance while reducing computational complexity and resource requirements. Dynamic pruning technology aims to remove parts of the neural network that have little impact on the model performance (such as accuracy), such as neurons, connections (weights), etc., thereby reducing the complexity and computational resource requirements of the model.

[0092] The specific steps of step S4 are as follows:

[0093] The edge computing device collects real-time digital images through a CMOS sensor with automatic white balance and exposure compensation, and performs preprocessing on the real-time digital images, including at least light compensation, image enhancement, dynamic noise reduction, ROI optimization and cropping, grayscale conversion, and normalization, to obtain optimized digital images;

[0094] The light compensation adopts a hybrid strategy of HSV color space histogram equalization and local contrast enhancement; the image enhancement is based on the CLAHE algorithm to enhance the contrast of the image and suppress noise; the dynamic noise reduction is based on a hybrid morphological filtering algorithm. The ROI optimization and cropping are achieved through a mask image, an image segmentation algorithm (such as GrabCut), segmentation based on Graph Cut, or ROI selection based on Bayesian optimization.

[0095] Adopting a hybrid strategy of HSV color space histogram equalization and local contrast enhancement is an effective image enhancement method, especially suitable for scenarios where it is necessary to simultaneously improve the overall contrast and local details of the image. The specific steps are as follows: 1. Convert the RGB image to an HSV image; 2. Perform histogram equalization on the brightness channel (V channel) of the HSV image; 3. Convert the processed HSV image back to an RGB image.

[0096] The hybrid morphological filtering algorithm is a non-linear filtering technology based on mathematical morphology, which performs operations such as erosion, dilation, opening (erosion followed by dilation), and closing (dilation followed by erosion) on the image by constructing specific structural elements; these operations can be used alone or in combination to achieve different image processing effects.

[0097] The specific steps of step S6 are as follows:

[0098] Real-time record the recognition log including at least the optimized digital image, digital recognition result, and recognition time, obtain the device serial number of the edge computing device itself, and extract image data and text data from the recognition log through a media type detection algorithm;

[0099] Compress the image data through the ZSTD algorithm to obtain image compressed data, and compress the text data through the Brotli algorithm to obtain text compressed data;

[0100] Calculate the first MAC value of the image compressed data and the device serial number through the HMAC algorithm, and encrypt the first MAC value and the image compressed data through the AES-256 algorithm to obtain the first encrypted data;

[0101] Calculate the second MAC value of the text compression data and the device serial number through the HMAC algorithm, and encrypt the second MAC value and the text compression data through the SM9 algorithm to obtain the second encrypted data;

[0102] Calculate the signature value of the recognition log through the EdDSA algorithm;

[0103] Package the first encrypted data, the second encrypted data, the signature value, and the device serial number into an encrypted log through the AES algorithm, upload the encrypted log to the server for storage in real time through the TLS protocol, and clear the encrypted log locally.

[0104] Extract the image data and text data from the recognition log by obtaining the device serial number of the edge computing device itself; compress the image data through the ZSTD algorithm to obtain image compression data, and compress the text data through the Brotli algorithm to obtain text compression data; calculate the first MAC value of the image compression data and the device serial number through the HMAC algorithm, and encrypt the first MAC value and the image compression data through the AES-256 algorithm to obtain the first encrypted data; calculate the second MAC value of the text compression data and the device serial number through the HMAC algorithm, and encrypt the second MAC value and the text compression data through the SM9 algorithm to obtain the second encrypted data; calculate the signature value of the recognition log through the EdDSA algorithm; package the first encrypted data, the second encrypted data, the signature value, and the device serial number into an encrypted log through the AES algorithm, and upload the encrypted log to the server for storage in real time through the TLS protocol; that is, different encryptions are performed on the recognition log based on the data types (image data and text data), and then combined for encryption. Moreover, the image data and text data respectively combine different encryption algorithms, and the TLS protocol is a secure transmission protocol. At least eight security measures are taken before and after (device serial number, first MAC value, AES-256 algorithm, second MAC value, SM9 algorithm, EdDSA algorithm, AES algorithm, TLS protocol) to prevent the recognition log from being stolen and tampered with in plain text, and combined with compression algorithms, thereby greatly improving the security of the transmission and storage of the recognition log, effectively enhancing the reliability of traceability, and greatly improving the transmission efficiency of the recognition log.

[0105] A preferred embodiment of a digital recognition system for routers, gateways, IP cameras, and ONUs according to the present invention includes the following modules:

[0106] A dataset construction module for obtaining a large number of historical digital images under different scenarios and different lighting conditions, performing preprocessing on each of the historical digital images including at least noise reduction, cropping, size unification, grayscale conversion, and normalization, and constructing a dataset after labeling the numbers on the preprocessed historical digital images;

[0107] By performing preprocessing on each historical digital image, including at least noise reduction, cropping, size unification, grayscale conversion, and normalization, the quality of the dataset is effectively improved, thereby enhancing the training effect of the digital recognition model.

[0108] Digital recognition model creation module, used to create a digital recognition model based on the local feature extraction module, global feature extraction module, feature fusion module, and prediction output module, and set the loss function of the digital recognition model;

[0109] Digital recognition model training module, used to train the digital recognition model through the dataset and the loss function, and deploy the trained digital recognition model to edge computing devices of device types such as routers, gateways, IP cameras, or ONUs;

[0110] Real-time digital image acquisition module, used for edge computing devices to collect real-time digital images through cameras, and perform preprocessing on the real-time digital images, including at least light compensation, image enhancement, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, to obtain optimized digital images;

[0111] By performing preprocessing on the real-time digital images, including at least light compensation, image enhancement, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, the noise data carried in the real-time digital images is effectively reduced, thereby greatly improving the accuracy of digital recognition.

[0112] Digital recognition module, used for edge computing devices to input the optimized digital images into the digital recognition model to obtain digital recognition results, and display the digital recognition results in real time through a display screen;

[0113] Recognition log management module, used to record in real time recognition logs including at least the optimized digital images, digital recognition results, and recognition times, encrypt the recognition logs into encrypted logs, upload the encrypted logs to the server for storage, and delete the encrypted logs locally.

[0114] By recording in real time recognition logs including at least optimized digital images, digital recognition results, and recognition times, encrypting the recognition logs into encrypted logs and uploading them to the server for storage, it is convenient for later traceability.

[0115] In the digital recognition model creation module, the local feature extraction module is used to extract local spatial features from digital images. It is constructed based on three depthwise separable convolutions. Each depthwise separable convolution is constructed based on a depthwise convolution and a pointwise convolution. By decomposing the standard convolution into two independent steps, it significantly reduces the computational amount and the number of parameters, while maintaining good feature extraction capabilities. The depthwise convolution is used to independently perform convolution operations on each channel of the input digital image to extract local spatial sub-features. The pointwise convolution is used to linearly combine each local spatial sub-feature through a 1×1 convolution kernel to restore the interaction between channels and generate higher-level local spatial features.

[0116] The global feature extraction module is used to extract global spatial features from digital images. It is constructed based on dilated convolutions, and the dilation rate takes a value of 2. When the dilation rate takes a value of 2, the convolution kernel samples once every 1 pixel, which is used to enable the convolution kernel to cover a larger receptive field without increasing additional parameters.

[0117] The depthwise separable convolution is mainly used to reduce the computational amount and the number of parameters, and is suitable for scenarios that require lightweight models, such as mobile devices and edge computing. The dilated convolution is mainly used to expand the receptive field, capture context information (global spatial features) in a larger range, and at the same time maintain the spatial resolution of the feature map (digital image). It is suitable for tasks that require capturing long-range dependencies, such as semantic segmentation and time series analysis.

[0118] By constructing a parallel structure of depthwise separable convolution and dilated convolution (dual-branch feature extraction network), while reducing the number of parameters to 2.3MB, a high recognition accuracy is maintained.

[0119] The feature fusion module is used to fuse local spatial features and global spatial features to obtain fused features. It is constructed based on the channel attention mechanism (Squeeze-and-Excitation Module, SE module). The channel attention mechanism dynamically adjusts the importance of each channel in the feature map by explicitly modeling the dependencies between channels, thereby enhancing the feature expression ability and performance of the model. The formula is:

[0120] F = α × F1 + (1 - α) × F2;

[0121] Among them, F represents the fused feature; F1 represents the local spatial feature; F2 represents the global spatial feature; α represents the fusion weight, α ∈ [0, 1], and is automatically learned by the SE module.

[0122] The prediction output module is used to output the digital recognition result, which is constructed based on the Global Average Pooling (GAP) layer and the Softmax classifier, and the value of the temperature scaling factor of the Softmax classifier is 1.5. The temperature scaling factor is a technique used to adjust the output probability distribution of a neural network model, usually applied to the Softmax classifier to calibrate the prediction confidence of the model. By replacing the traditional fully connected layer with the global average pooling layer, the number of parameters can be effectively reduced, overfitting can be prevented, the robustness of the model can be enhanced, and the network structure can be simplified.

[0123] The loss function uses the cross-entropy loss function.

[0124] The cross-entropy loss function measures the loss by calculating the difference between the probability distribution predicted by the model and the probability distribution of the true label. For image classification tasks, especially digital recognition, the output of the model is usually a probability distribution representing the confidence of the input image belonging to each category. The cross-entropy loss function provides more direct gradient information for the model, which helps to converge faster during the training process and gives higher penalties to misclassified predictions, enabling the model to better learn the discrimination between categories. That is, the cross-entropy loss function can effectively measure the difference between the predicted probability distribution and the true distribution and accelerate the convergence of the model.

[0125] The digital recognition model training module is specifically used for:

[0126] Dividing the dataset into a training set, a validation set, and a test set based on a preset splitting ratio, training the digital recognition model with the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, compress the digital recognition model through knowledge distillation technology and dynamic pruning technology. Validate the trained digital recognition model with the validation set to determine whether the recognition accuracy is greater than a preset accuracy threshold. If not, the validation fails, and the training set is expanded and training continues. If so, the validation succeeds. Test the digitally recognized model that has passed the validation with the test set to determine whether the confidence is greater than a preset confidence threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test succeeds and training ends.

[0127] Deploy the trained digital recognition model to edge computing devices of device types such as routers, gateways, IP cameras, or ONUs.

[0128] The digital recognition model is trained using a training set until the loss value of the loss function is less than a preset loss threshold. During the training process, the digital recognition model is compressed using knowledge distillation technology and dynamic pruning technology. Then, the accuracy is calculated using a validation set to verify the trained digital recognition model. The confidence level is calculated using a test set to test the digital recognition model that has passed the verification. That is, during the training process of the digital recognition model, compression, verification, and testing are continuously performed to effectively balance the model volume and recognition accuracy of the digital recognition model.

[0129] Knowledge distillation technology is a machine learning model compression method aimed at transferring the knowledge of a large model to a small model to improve the model performance and generalization ability. The core idea of knowledge distillation is to transform the knowledge of a complex model into a more concise and effective representation, reducing the computational complexity and resource requirements while maintaining high performance. Dynamic pruning technology aims to remove parts of the neural network that have little impact on the model performance (such as accuracy), such as neurons, connections (weights), etc., thereby reducing the model complexity and computational resource requirements.

[0130] The real-time digital image acquisition module is specifically used for:

[0131] The edge computing device acquires real-time digital images through a CMOS sensor with automatic white balance and exposure compensation, and performs preprocessing on the real-time digital images, including at least light compensation, image enhancement, dynamic noise reduction, ROI optimization and cropping, grayscale conversion, and normalization, to obtain optimized digital images;

[0132] The light compensation adopts a hybrid strategy of HSV color space histogram equalization and local contrast enhancement; the image enhancement is based on the CLAHE algorithm to enhance the image contrast and suppress noise; the dynamic noise reduction is based on the hybrid morphological filtering algorithm. The ROI optimization and cropping are achieved through a mask image, an image segmentation algorithm (such as GrabCut), segmentation based on Graph Cut, or ROI selection based on Bayesian optimization.

[0133] Adopting a hybrid strategy of HSV color space histogram equalization and local contrast enhancement is an effective image enhancement method, especially suitable for scenarios that require simultaneously enhancing the overall contrast and local details of an image. The specific steps are as follows: 1. Convert the RGB image to an HSV image; 2. Perform histogram equalization on the brightness channel (V channel) of the HSV image; 3. Convert the processed HSV image back to an RGB image.

[0134] The hybrid morphological filtering algorithm is a non-linear filtering technique based on mathematical morphology. It performs operations such as erosion, dilation, opening (erosion followed by dilation), and closing (dilation followed by erosion) on the image by constructing specific structuring elements. These operations can be used alone or in combination to achieve different image processing effects.

[0135] The recognition log management module is specifically used for:

[0136] Real-time recording of the recognition log including at least the optimized digital image, digital recognition result, and recognition time, obtaining the device serial number of the edge computing device itself, and extracting image data and text data from the recognition log through the media type detection algorithm;

[0137] Compressing the image data through the ZSTD algorithm to obtain image compressed data, and compressing the text data through the Brotli algorithm to obtain text compressed data;

[0138] Calculating the first MAC value of the image compressed data and the device serial number through the HMAC algorithm, and encrypting the first MAC value and the image compressed data through the AES-256 algorithm to obtain the first encrypted data;

[0139] Calculating the second MAC value of the text compressed data and the device serial number through the HMAC algorithm, and encrypting the second MAC value and the text compressed data through the SM9 algorithm to obtain the second encrypted data;

[0140] Calculating the signature value of the recognition log through the EdDSA algorithm;

[0141] Encapsulating the first encrypted data, the second encrypted data, the signature value, and the device serial number into an encrypted log through the AES algorithm, uploading the encrypted log to the server for storage in real time through the TLS protocol, and clearing the encrypted log locally.

[0142] By obtaining the device serial number of the edge computing device itself, extract image data and text data from the recognition log; compress the image data through the ZSTD algorithm to obtain image compressed data, and compress the text data through the Brotli algorithm to obtain text compressed data; calculate the first MAC value of the image compressed data and the device serial number through the HMAC algorithm, and encrypt the first MAC value and the image compressed data through the AES-256 algorithm to obtain the first encrypted data; calculate the second MAC value of the text compressed data and the device serial number through the HMAC algorithm, and encrypt the second MAC value and the text compressed data through the SM9 algorithm to obtain the second encrypted data; calculate the signature value of the recognition log through the EdDSA algorithm; encapsulate the first encrypted data, the second encrypted data, the signature value, and the device serial number into an encrypted log through the AES algorithm, and upload the encrypted log to the server for storage in real time through the TLS protocol; that is, different encryptions are performed on the recognition log based on the data type (image data and text data), and then combined for encryption at the end. Moreover, different encryption algorithms are combined for image data and text data respectively, and the TLS protocol is a secure transmission protocol. At least 8 security measures are taken before and after (device serial number, first MAC value, AES-256 algorithm, second MAC value, SM9 algorithm, EdDSA algorithm, AES algorithm, TLS protocol) to prevent the recognition log from being stolen and tampered with in plain text, and compression algorithms are combined, thereby greatly improving the security of the transmission and storage of the recognition log, effectively improving the reliability of traceability, and greatly improving the transmission efficiency of the recognition log.

[0143] In summary, the advantages of the present invention are as follows:

[0144] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.

Claims

1. A digital recognition method for routers, gateways, IP cameras, and ONUs, characterized in that: The steps include the following: Step S1: Obtain a large number of historical digital images under different scenarios and different lighting conditions. Perform preprocessing on each of the historical digital images, including at least noise reduction, cropping, size unification, grayscale conversion, and normalization. After preprocessing, label each of the historical digital images and construct a dataset. Step S2: Create a digital recognition model based on a local feature extraction module, a global feature extraction module, a feature fusion module, and a prediction output module. Set the loss function of the digital recognition model. Step S3: Train the digital recognition model using the dataset and the loss function. Deploy the trained digital recognition model to an edge computing device with a device type of router, gateway, IPC, or ONU. Step S4: The edge computing device captures real-time digital images through a camera. Perform preprocessing on the real-time digital images, including at least light compensation, image enhancement, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, to obtain optimized digital images. Step S5: The edge computing device inputs the optimized digital images into the digital recognition model to obtain digital recognition results, and displays the digital recognition results in real time through a display screen. Step S6: Record the recognition log in real time, including at least the optimized digital images, digital recognition results, and recognition time. Encrypt the recognition log into an encrypted log, upload the encrypted log to a server for storage, and delete the encrypted log locally.

2. The digital recognition method for a router, gateway, IPC, and ONU according to claim 1, characterized in that: In step S2, the local feature extraction module is used to extract local spatial features from digital images. It is constructed based on three depthwise separable convolutions, and each depthwise separable convolution is constructed based on a depth convolution and a pointwise convolution. The depth convolution is used to perform convolution operations independently on each channel of the input digital image to extract local spatial sub-features. The pointwise convolution is used to linearly combine each local spatial sub-feature through a 1×1 convolutional kernel to restore the interaction between channels and generate higher-level local spatial features. The global feature extraction module is used to extract global spatial features from digital images. It is constructed based on dilated convolutions, and the dilation rate is set to 2. The feature fusion module is used to fuse local spatial features and global spatial features to obtain fused features. It is constructed based on a channel attention mechanism. The prediction output module is used to output digital recognition results. It is constructed based on a global average pooling layer and a Softmax classifier, and the temperature scaling factor of the Softmax classifier is set to 1.

5. The loss function uses a cross-entropy loss function.

3. A digital recognition method for routers, gateways, IP cameras, and ONUs according to claim 1, characterized in that: The specific content of step S3 is as follows: Divide the dataset into a training set, a validation set, and a test set based on a preset splitting ratio. Train the digit recognition model using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, compress the digit recognition model using knowledge distillation technology and dynamic pruning technology. Validate the trained digit recognition model using the validation set to determine whether the recognition accuracy is greater than a preset accuracy threshold. If not, the validation fails, and the training set is expanded and training continues. If so, the validation succeeds. Test the digit recognition model that has passed the validation using the test set to determine whether the confidence level is greater than a preset confidence threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test succeeds, and training ends; Deploy the trained digit recognition model to an edge computing device with a device type of router, gateway, IPC, or ONU.

4. A digital recognition method for a router, gateway, IPC, and ONU according to claim 1, characterized in that: The specific steps of step S4 are as follows: The edge computing device collects real-time digit images through a CMOS sensor with automatic white balance and exposure compensation, and performs preprocessing on the real-time digit images, including at least light compensation, image enhancement, dynamic noise reduction, ROI optimization and cropping, grayscale conversion, and normalization to obtain optimized digit images; The light compensation adopts a hybrid strategy of HSV color space histogram equalization and local contrast enhancement; the image enhancement is based on the CLAHE algorithm to enhance the contrast of the image and suppress noise; the dynamic noise reduction is based on a hybrid morphological filtering algorithm.

5. A digital recognition method for routers, gateways, IP cameras, and ONUs according to claim 1, characterized in that: The specific steps of step S6 are as follows: Record in real time an identification log including at least the optimized digit images, digit recognition results, and recognition times, obtain the device serial number of the edge computing device itself, and extract image data and text data from the identification log through a media type detection algorithm; Compress the image data through the ZSTD algorithm to obtain compressed image data, and compress the text data through the Brotl i algorithm to obtain compressed text data; Calculate the first MAC value of the compressed image data and the device serial number through the HMAC algorithm, and encrypt the first MAC value and the compressed image data through the AES-256 algorithm to obtain first encrypted data; Calculate the second MAC value of the compressed text data and the device serial number through the HMAC algorithm, and encrypt the second MAC value and the compressed text data through the SM9 algorithm to obtain second encrypted data; Calculate the signature value of the identification log through the EdDSA algorithm; Package the first encrypted data, second encrypted data, signature value, and device serial number into an encrypted log through the AES algorithm, upload the encrypted log to the server for storage in real time through the TLS protocol, and delete the encrypted log locally.

6. A digital recognition system for routers, gateways, IP cameras, and ONUs, characterized in that: It includes the following modules: A dataset construction module, which is used to obtain a large number of historical digit images in different scenarios and under different lighting conditions, perform preprocessing on each of the historical digit images, including at least noise reduction, cropping, size unification, grayscale conversion, and normalization, and construct a dataset after annotating the digits of the preprocessed historical digit images; A digital recognition model creation module, which is used to create a digital recognition model based on a local feature extraction module, a global feature extraction module, a feature fusion module, and a prediction output module, and set the loss function of the digital recognition model; A digital recognition model training module, which is used to train the digital recognition model through the data set and the loss function, and deploy the trained digital recognition model to edge computing devices with device types of routers, gateways, IP cameras, or ONUs; A real-time digital image acquisition module, which is used for edge computing devices to collect real-time digital images through cameras, and perform preprocessing on the real-time digital images, including at least illumination compensation, image enhancement, dynamic noise reduction, ROI optimization and cropping, grayscale conversion, and normalization to obtain optimized digital images; A digital recognition module, which is used for edge computing devices to input the optimized digital images into the digital recognition model to obtain digital recognition results, and display the digital recognition results in real time through a display screen; An identification log management module, which is used to record in real time the identification logs including at least the optimized digital images, digital recognition results, and recognition times, encrypt the identification logs into encrypted logs, upload the encrypted logs to a server for storage, and delete the encrypted logs locally.

7. The digital recognition system for routers, gateways, IP cameras, and ONUs according to claim 6, wherein: In the digital recognition model creation module, the local feature extraction module is used to extract local spatial features from digital images. It is constructed based on three depthwise separable convolutions, and each depthwise separable convolution is constructed based on a depth convolution and a pointwise convolution; the depth convolution is used to perform convolution operations independently on each channel of the input digital image to extract local spatial sub-features; the pointwise convolution is used to linearly combine each local spatial sub-feature through a 1×1 convolution kernel to restore the interaction between channels and generate higher-level local spatial features; The global feature extraction module is used to extract global spatial features from digital images. It is constructed based on dilated convolutions, and the dilation rate is set to 2; The feature fusion module is used to fuse local spatial features and global spatial features to obtain fused features. It is constructed based on a channel attention mechanism; The prediction output module is used to output digital recognition results. It is constructed based on a global average pooling layer and a Softmax classifier, and the temperature scaling factor of the Softmax classifier is set to 1.5; The loss function uses a cross-entropy loss function.

8. The digital recognition system for a router, gateway, IPC, and ONU according to claim 6, characterized in that: The digital recognition model training module is specifically used for: Divide the dataset into a training set, a validation set, and a test set based on a preset splitting ratio. Train a digital recognition model using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, compress the digital recognition model using knowledge distillation technology and dynamic pruning technology. Validate the trained digital recognition model using the validation set to determine whether the recognition accuracy is greater than a preset accuracy threshold. If not, the validation fails, and the training set is expanded and training continues. If so, the validation succeeds. Test the digital recognition model that has passed the validation using the test set to determine whether the confidence level is greater than a preset confidence threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test succeeds, and the training ends. Deploy the trained digital recognition model to edge computing devices of device types such as routers, gateways, IP cameras, or ONUs.

9. A digital recognition system for routers, gateways, IP cameras, and ONUs according to claim 6, characterized in that: The real-time digital image acquisition module is specifically used for: The edge computing device acquires real-time digital images through a CMOS sensor with automatic white balance and exposure compensation, and performs preprocessing on the real-time digital images, including at least light compensation, image enhancement, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, to obtain optimized digital images. The light compensation adopts a hybrid strategy of HSV color space histogram equalization and local contrast enhancement. The image enhancement is based on the CLAHE algorithm to enhance the contrast of the image and suppress noise. The dynamic noise reduction is based on a hybrid morphological filtering algorithm.

10. A digital recognition system for routers, gateways, IP cameras, and ONUs as claimed in claim 6, characterized in that: The recognition log management module is specifically used for: Real-time record recognition logs including at least the optimized digital images, digital recognition results, and recognition times, obtain the device serial number of the edge computing device itself, and extract image data and text data from the recognition logs through a media type detection algorithm. Compress the image data using the ZSTD algorithm to obtain image compressed data, and compress the text data using the Brotl i algorithm to obtain text compressed data. Calculate the first MAC value of the image compressed data and the device serial number through the HMAC algorithm, and encrypt the first MAC value and the image compressed data using the AES-256 algorithm to obtain the first encrypted data. Calculate the second MAC value of the text compressed data and the device serial number through the HMAC algorithm, and encrypt the second MAC value and the text compressed data using the SM9 algorithm to obtain the second encrypted data. Calculate the signature value of the recognition log through the EdDSA algorithm. Package the first encrypted data, the second encrypted data, the signature value, and the device serial number into an encrypted log using the AES algorithm, upload the encrypted log to the server for storage in real time through the TLS protocol, and delete the encrypted log locally.