Posture detection method and system for router, gateway and camera
By building a lightweight attitude detection model on edge computing devices and combining light compensation and image preprocessing, the problem of high resource requirements and low accuracy of attitude detection on resource-constrained devices is solved, and efficient and accurate attitude detection is achieved.
Patent Information
- Application Number
- CN202510269048.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-22
Smart Images

Figure CN120356067A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of edge computing, and particularly to a method and system for detecting the posture of routers, gateways, and cameras. Background Art
[0002] The rise of edge computing devices such as routers, gateways, and cameras (indoor cameras, outdoor cameras, action cameras) has brought many innovative application scenarios, among which posture detection is applied to many application scenarios; for example: in the field of security monitoring, abnormal behavior and intrusion behavior are judged through posture detection; in the field of medical and health, rehabilitation training is carried out through posture detection; in the field of sports training, the movements of athletes are captured through posture detection for standardized training; in the field of smart home, smart home devices are controlled through posture detection; in the field of human-computer interaction, intention recognition is carried out through posture detection for interaction.
[0003] For the detection of posture, traditionally, a cloud processing architecture is generally adopted, that is, the edge computing device transmits the collected video (image frames) to the cloud for detection, and then returns the detection result to the terminal. The traditional method has data transmission bottlenecks and privacy leakage risks; the transmission of a single-frame 1080P image to the cloud generates an average delay of 300 - 500 ms, and the bandwidth occupancy reaches 20 - 50 Mbps in the video stream processing scenario. The original image data has the risk of being intercepted when transmitted over the public network, and it is not applicable to scenarios such as medical and security. Therefore, there is a need for local posture detection on edge computing devices.
[0004] However, the number of parameters of traditional backbone networks such as VGG16 reaches 138M, and it is difficult to deploy the posture detection model based on the traditional backbone network on resource-constrained edge computing devices. Although there are also some lightweight posture detection models, the detection accuracy of these models is poor, especially the detection accuracy under dynamic lighting conditions drops by 40%, and the key point positioning drifts significantly.
[0005] Therefore, how to provide a method and system for detecting the posture of routers, gateways, and cameras to reduce the resource requirements for posture detection and improve the accuracy of posture detection has become an urgent technical problem to be solved. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method and system for detecting the posture of routers, gateways, and cameras to reduce the resource requirements for posture detection and improve the accuracy of posture detection.
[0007] In a first aspect, the present invention provides a method for detecting the posture of routers, gateways, and cameras, including the following steps:
[0008] Step S1: Obtain a large number of historical pose images, perform preprocessing on each of the historical pose images including at least noise reduction, cropping, and size unification, annotate the poses of the preprocessed historical pose images, and then perform sample augmentation on each of the annotated historical pose data including at least rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, saturation adjustment, hue adjustment, and adding random noise to construct a dataset;
[0009] Step S2: Create a pose detection model based on MobileNetV3, EfficientNet-Lite, a feature fusion module, and an output module, and set the loss function of the pose detection model;
[0010] The input end of the feature fusion module is connected to MobileNetV3 and EfficientNet-Lite, and the output end is connected to the output module; both MobileNetV3 and EfficientNet-Lite are used to extract pose features from pose images; the feature fusion module is used to fuse the pose features extracted by MobileNetV3 and EfficientNet-Lite to obtain 128-dimensional fused features; the output module is used to output pose detection results based on the fused features;
[0011] Step S3: Train the pose detection model through the dataset and the loss function, compress the pose detection model during the training process, and deploy the trained pose detection model to an edge computing device of device type router, gateway, or camera;
[0012] Step S4: The edge computing device collects surveillance videos in real time through a camera, detects image frames with humans in the surveillance videos through a pre-trained object detection model, and then extracts real-time pose images from the surveillance videos;
[0013] Step S5: The edge computing device performs preprocessing on the real-time pose image including at least light compensation, dynamic noise reduction, ROI optimized cropping, grayscaling, and normalization to obtain an optimized pose image;
[0014] Step S6: The edge computing device inputs the optimized pose image into the pose detection model to obtain a pose detection result, and superimposes the pose detection result on the video stream;
[0015] Step S7: Real-time record a detection log including at least the optimized pose image, the pose detection result, and the detection time, encrypt the detection log into an encrypted log, upload the encrypted log to the server for storage, and delete the encrypted log locally.
[0016] Further, in the step S2, the MobileNetV3 is constructed based on depthwise separable convolution and SENet units, and the number of network layers is 12; the compound scaling factor of the EfficientNet-Lite is 0.75, and the compound scaling factor is used to adjust the network width, network depth, and resolution; the loss function is constructed based on the OKS function and the CIoU function.
[0017] Further, the step S3 is specifically as follows:
[0018] Based on a preset segmentation ratio, the data set is divided into a training set, a validation set, and a test set. The pose detection model is trained using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, the learning rate is adaptively adjusted by the LAMB optimizer, and the pose detection model is compressed using knowledge distillation technology, dynamic pruning technology, and model quantization technology; the trained pose detection model is verified using the validation set to determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the verification fails, and the training set is expanded and training continues. If so, the verification is successful; the successfully verified pose detection model is tested using the test set to determine whether the confidence level is greater than a preset confidence level threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test is successful, and training ends;
[0019] The trained pose detection model is deployed to an edge computing device with a device type of router, gateway, or camera.
[0020] Further, in the step S4, the object detection model is EdgeYOLO-Tiny;
[0021] In the step S5, the light compensation is specifically as follows: The contrast of the low-light region in the real-time pose image is enhanced by the CLAHE algorithm for light compensation;
[0022] The dynamic noise reduction is specifically as follows: The real-time pose image is dynamically denoised by the cooperation of the Kalman filter algorithm and the bilateral filter algorithm.
[0023] Further, the step S7 is specifically as follows:
[0024] The real-time recording includes at least the detection log of the optimized pose image, the pose detection result, and the detection time, obtains the device serial number of the edge computing device itself, extracts the image data and text data from the detection log, separates the EXIF information and pixel data in the image data, and performs a hash calculation on the EXIF information and the device serial number through the SHA3-256 algorithm to obtain the first data fingerprint, and performs a hash calculation on the text data and the device serial number through the SHA3-512 algorithm to obtain the second data fingerprint;
[0025] Set a row-column permutation rule, perform a row-column permutation operation on the pixel data based on the row-column permutation rule to obtain encrypted pixel data, and encrypt the EXIF information, the encrypted pixel data, and the first data fingerprint through the IEDA algorithm to obtain the first encrypted data;
[0026] Set a cyclic shift rule, perform a shift on the characters of the text data based on the cyclic shift rule to obtain encrypted text data, and encrypt the encrypted text data and the second data fingerprint through the RC6 algorithm to obtain the second encrypted data;
[0027] Encrypt the first encrypted data and the second encrypted data into an encrypted log through the AES-KW algorithm, upload the encrypted log to the server for storage in real time through the SSL protocol, and clear the encrypted log locally.
[0028] In a second aspect, the present invention provides a pose detection system for routers, gateways, and cameras, including the following modules:
[0029] The dataset construction module is used to obtain a large number of historical pose images, perform preprocessing on each of the historical pose images including at least noise reduction, cropping, and size unification, perform pose annotation on each of the preprocessed historical pose images, and then perform sample augmentation on each of the annotated historical pose data including at least rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, saturation adjustment, hue adjustment, and adding random noise to construct a dataset;
[0030] The pose detection model creation module is used to create a pose detection model based on MobileNetV3, EfficientNet-Lite, a feature fusion module, and an output module, and set the loss function of the pose detection model;
[0031] The input end of the feature fusion module is connected to MobileNetV3 and EfficientNet-Lite, and the output end is connected to the output module; both MobileNetV3 and EfficientNet-Lite are used to extract pose features from the pose images; the feature fusion module is used to fuse the pose features extracted by MobileNetV3 and EfficientNet-Lite to obtain 128-dimensional fused features; the output module is used to output the pose detection result based on the fused features;
[0032] The pose detection model deployment module is used to train the pose detection model through the data set and the loss function, compress the pose detection model during the training process, and deploy the trained pose detection model to edge computing devices of device types such as routers, gateways, or cameras;
[0033] The real-time pose image extraction module is used for the edge computing device to collect the monitoring video in real time through the camera, detect the image frames with human bodies in the monitoring video through a pre-trained object detection model, and then extract the real-time pose image from the monitoring video;
[0034] The real-time pose image processing module is used for the edge computing device to perform preprocessing on the real-time pose image, including at least light compensation, dynamic noise reduction, ROI optimization and cropping, grayscale conversion, and normalization, to obtain an optimized pose image;
[0035] The pose detection module is used for the edge computing device to input the optimized pose image into the pose detection model to obtain the pose detection result, and superimpose the pose detection result on the video stream;
[0036] The detection log management module is used to record in real time the detection logs including at least the optimized pose image, the pose detection result, and the detection time, encrypt the detection logs into encrypted logs, upload the encrypted logs to the server for storage, and delete the encrypted logs locally.
[0037] Furthermore, in the pose detection model creation module, MobileNetV3 is constructed based on depthwise separable convolution and SENet units, and the number of network layers is 12; the compound scaling factor of EfficientNet-Lite takes a value of 0.75, and the compound scaling factor is used to adjust the network width, network depth, and resolution; the loss function is constructed based on the OKS function and the CIoU function.
[0038] Furthermore, the pose detection model deployment module is specifically used for:
[0039] Divide the dataset into a training set, a validation set, and a test set based on a preset splitting ratio. Train the pose detection model using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, adaptively adjust the learning rate using the LAMB optimizer, and compress the pose detection model using knowledge distillation technology, dynamic pruning technology, and model quantization technology. Validate the trained pose detection model using the validation set to determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the validation fails, and the training set is expanded and training continues. If so, the validation succeeds. Test the pose detection model that has passed the validation using the test set to determine whether the confidence level is greater than a preset confidence threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test succeeds, and training ends;
[0040] Deploy the trained pose detection model to an edge computing device with a device type of router, gateway, or camera.
[0041] Further, in the real-time pose image extraction module, the object detection model is EdgeYOLO-Tiny;
[0042] In the real-time pose image processing module, the light compensation is specifically: enhance the contrast of the low-light area in the real-time pose image through the CLAHE algorithm to perform light compensation;
[0043] The dynamic noise reduction is specifically: perform dynamic noise reduction on the real-time pose image through the cooperation of the Kalman filtering algorithm and the bilateral filtering algorithm.
[0044] Further, the detection log management module is specifically used for:
[0045] Real-time record the detection log including at least the optimized pose image, pose detection result, and detection time, obtain the device serial number of the edge computing device itself, extract the image data and text data from the detection log, separate the EXIF information and pixel data in the image data, calculate the first data fingerprint by hashing the EXIF information and the device serial number through the SHA3-256 algorithm, and calculate the second data fingerprint by hashing the text data and the device serial number through the SHA3-512 algorithm;
[0046] Set a row-column permutation rule, perform a row-column permutation operation on the pixel data based on the row-column permutation rule to obtain encrypted pixel data, and encrypt the EXIF information, encrypted pixel data, and the first data fingerprint through the IEDA algorithm to obtain the first encrypted data;
[0047] Set a cyclic displacement rule, displace the characters of the text data based on the cyclic displacement rule to obtain encrypted text data, and encrypt the encrypted text data and the second data fingerprint through the RC6 algorithm to obtain second encrypted data;
[0048] Encrypt the first encrypted data and the second encrypted data into an encrypted log through the AES-KW algorithm, upload the encrypted log to the server for storage in real time through the SSL protocol, and clear the encrypted log locally.
[0049] The advantages of the present invention are as follows:
[0050] 1. By obtaining a large number of historical pose images, preprocessing, annotating, and augmenting the samples of each historical pose image to construct a dataset; then creating a pose detection model based on MobileNetV3, EfficientNet-Lite, a feature fusion module, and an output module, and setting the loss function of the pose detection model; then training the pose detection model through the dataset and the loss function, compressing the pose detection model during the training process, and deploying the trained pose detection model to an edge computing device with a device type of router, gateway, or camera; then the edge computing device collects monitoring videos in real time through a camera, detects the image frames with humans in the monitoring videos through a pre-trained object detection model, extracts real-time pose images from the monitoring videos, performs preprocessing including at least illumination compensation, dynamic noise reduction, ROI optimization and cropping, grayscale conversion, and normalization on the real-time pose images to obtain optimized pose images, inputs the optimized pose images into the pose detection model to obtain pose detection results, and superimposes and displays the pose detection results on the video stream; at the same time, real-time records of detection logs including at least the optimized pose images, pose detection results, and detection time are recorded, the detection logs are encrypted into encrypted logs and uploaded to the server for storage, and the local encrypted logs are cleared; since the pose detection model is constructed based on lightweight MobileNetV3 and EfficientNet-Lite, and the number of network layers of MobileNetV3 is reduced to 12 layers, and the compound scaling factor of EfficientNet-Lite is set to 0.75 to reduce the network width, network depth, and resolution, and the pose detection model is compressed through knowledge distillation technology, dynamic pruning technology, and model quantization technology during the training process, effectively reducing the number of parameters (complexity) and volume of the pose detection model, and locally deleting the uploaded encrypted logs in real time to reduce storage overhead, effectively reducing the computing overhead of the edge computing device; through cross-detection (fusion detection), image preprocessing, and an improved loss function of the dual channels of MobileNetV3 and EfficientNet-Lite, the generalization ability and detection performance of pose detection are effectively improved, ultimately greatly reducing the resource requirements for pose detection, greatly improving the accuracy and real-time performance of pose detection, and greatly reducing the power consumption of the edge computing device for pose detection.
[0051] 2. By performing sample augmentation on each annotated historical pose data including at least rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, saturation adjustment, hue adjustment, and adding random noise, the sample size of the dataset is effectively augmented, and thus the generalization performance of the pose detection model is greatly improved.
[0052] 3. By performing preprocessing on the real-time pose image before pose detection, including at least illumination compensation, dynamic noise reduction, ROI optimization and cropping, grayscale conversion, and normalization, the image noise is effectively reduced and the image quality is improved, thereby greatly enhancing the accuracy of pose detection.
[0053] 4. By recording in real time the detection log including at least the optimized pose image, pose detection results, and detection time, encrypting the detection log into an encrypted log and uploading it to the server for storage, it is convenient for later traceability.
[0054] 5. The OKS function is a scale-normalized loss function that can assign different weights according to the importance of key points. For example, head key points (such as eyes, nose) are more easily penalized for pixel-level errors than body key points (such as shoulders, knees); the CIoU function can significantly improve the accuracy and efficiency of pose detection by comprehensively considering the overlapping area, center point distance, and aspect ratio. It not only reduces the cases of false detection and missed detection, but also speeds up the convergence speed of the model, making it perform well in complex scenes and small object detection; by setting the loss function to be constructed based on the OKS function and the CIoU function, that is, weighting the OKS function and the CIoU function, the advantages of the OKS function and the CIoU function are combined, greatly enhancing the training effect of the pose detection model, and thus greatly enhancing the accuracy of pose detection.
[0055] 6. The pose detection model is trained with a training set until the loss value of the loss function is less than a preset loss threshold. During the training process, the learning rate is adaptively adjusted by the LAMB optimizer, and the pose detection model is compressed by knowledge distillation technology, dynamic pruning technology, and model quantization technology; the detection accuracy is calculated through a validation set to verify the trained pose detection model, and the confidence level is calculated through a test set to test the pose detection model that has passed the verification. That is, during the training process of the pose detection model, compression, verification, and testing are continuously performed to effectively balance the model volume and detection accuracy of the pose detection model, and the LAMB optimizer is also combined to greatly accelerate the model convergence speed.
[0056] 7. By setting the object detection model to EdgeYOLO-Tiny, EdgeYOLO-Tiny is a lightweight object detection model designed specifically for edge computing devices, with high inference speed and small model volume, facilitating deployment on resource-constrained edge computing devices.
[0057] 8. By obtaining the device serial number of the edge computing device itself, extracting image data and text data from the detection log, separating the EXIF information and pixel data in the image data, performing hash calculation on the EXIF information and the device serial number through the SHA3-256 algorithm to obtain the first data fingerprint, and performing hash calculation on the text data and the device serial number through the SHA3-512 algorithm to obtain the second data fingerprint; performing row-column permutation operation on the pixel data based on the row-column permutation rule to obtain encrypted pixel data, and encrypting the EXIF information, the encrypted pixel data, and the first data fingerprint through the IEDA algorithm to obtain the first encrypted data; performing character displacement on the text data based on the cyclic shift rule to obtain encrypted text data, and encrypting the encrypted text data and the second data fingerprint through the RC6 algorithm to obtain the second encrypted data; encrypting the first encrypted data and the second encrypted data into an encrypted log through the AES-KW algorithm, and uploading the encrypted log to the server for storage in real time through the SSL protocol; that is, different encryptions are performed on the detection log based on the data types (image data and text data), and then the encrypted data is merged. Moreover, the image data and the text data respectively combine different encryption algorithms and data transformation rules, and the SSL protocol is a secure transmission protocol. At least nine security measures are taken before and after (device serial number, first data fingerprint, second data fingerprint, row-column permutation rule, IEDA algorithm, cyclic shift rule, RC6 algorithm, AES-KW algorithm, SSL protocol) to prevent the detection log from being stolen and tampered with in plaintext, thereby greatly improving the security of the detection log transmission and storage, and effectively enhancing the reliability of traceability.
[0058] 9. By performing sample augmentation operations including brightness adjustment, contrast adjustment, saturation adjustment, and hue adjustment, and then training the pose detection model using the augmented dataset, and performing light compensation on the real-time pose image before pose recognition, effectively overcoming the problem of the decrease in detection accuracy under traditional dynamic lighting conditions, and greatly improving the accuracy of pose detection. Brief Description of the Drawings
[0059] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.
[0060] Figure 1 It is a flowchart of a pose detection method for routers, gateways, and cameras according to the present invention.
[0061] Figure 2 It is a schematic structural diagram of a pose detection system for routers, gateways, and cameras according to the present invention. Detailed Embodiments
[0062] The general idea of the technical solution in the embodiments of this application is as follows: Perform pose detection using a pose detection model created based on MobileNetV3, EfficientNet-Lite, a feature fusion module, and an output module. The number of network layers of MobileNetV3 is reduced to 12 layers, and the compound scaling factor of EfficientNet-Lite is set to 0.75 to reduce the network width, network depth, and resolution. During the training process of the pose detection model, compression is performed through knowledge distillation technology, dynamic pruning technology, and model quantization technology, effectively reducing the volume of the pose detection model. Encrypted logs that have been uploaded are locally and real-time deleted to reduce storage overhead; cross-detection, image preprocessing, and an improved loss function are performed through the dual channels of MobileNetV3 and EfficientNet-Lite, effectively improving the generalization ability and detection performance of pose detection, thereby reducing the resource requirements for pose detection and improving the accuracy of pose detection.
[0063] Please refer to Figures 1 to 2 As shown, a preferred embodiment of a method for pose detection of a router, gateway, and camera according to the present invention includes the following steps:
[0064] Step S1: Obtain a large number of historical pose images, perform preprocessing on each of the historical pose images, including at least noise reduction, cropping, and size unification, perform pose annotation on each of the preprocessed historical pose images, and then perform sample augmentation on each of the annotated historical pose data, including at least rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, saturation adjustment, hue adjustment, and adding random noise, to construct a dataset;
[0065] By performing sample augmentation on each of the annotated historical pose data, including at least rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, saturation adjustment, hue adjustment, and adding random noise, the sample size of the dataset is effectively increased, and thus the generalization performance of the pose detection model is greatly improved.
[0066] Step S2: Create a pose detection model based on MobileNetV3, EfficientNet-Lite, a feature fusion module, and an output module, and set the loss function of the pose detection model;
[0067] The input end of the feature fusion module is connected to MobileNetV3 and EfficientNet-Lite, and the output end is connected to the output module; both MobileNetV3 and EfficientNet-Lite are used to extract pose features from the pose image; the feature fusion module is used to fuse the pose features extracted by MobileNetV3 and EfficientNet-Lite to obtain 128-dimensional fused features; the output module is used to output the pose detection result based on the fused features;
[0068] MobileNetV3 is a lightweight neural network architecture designed specifically for mobile devices and embedded systems, aiming to improve efficiency and performance; EfficientNet-Lite is a lightweight version of the EfficientNet series, optimized for mobile and edge devices, with advantages such as efficient computing performance, low power consumption, fast inference, and easy deployment.
[0069] Step S3: Train the pose detection model through the dataset and the loss function. During the training process, compress the pose detection model, and deploy the trained pose detection model to edge computing devices of device types such as routers, gateways, or cameras;
[0070] Step S4: The edge computing device collects the monitoring video in real time through the camera, detects the image frames with people in the monitoring video through the pre-trained object detection model, and then extracts the real-time pose image from the monitoring video;
[0071] Step S5: The edge computing device preprocesses the real-time pose image, including at least light compensation, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, to obtain an optimized pose image;
[0072] By preprocessing the real-time pose image, including at least light compensation, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, before pose detection, the image noise is effectively reduced and the image quality is improved, thereby greatly improving the accuracy of pose detection.
[0073] Step S6: The edge computing device inputs the optimized pose image into the pose detection model to obtain the pose detection result, and superimposes the pose detection result on the video stream;
[0074] Step S7: Record the detection log including at least the optimized pose image, the pose detection result, and the detection time in real time, encrypt the detection log into an encrypted log, upload the encrypted log to the server for storage, and delete the encrypted log locally.
[0075] By recording in real time the detection log including at least the optimized pose image, pose detection result, and detection time, encrypting the detection log into an encrypted log and uploading it to the server for storage, which is convenient for later traceability.
[0076] In the step S2, the MobileNetV3 is constructed based on depthwise separable convolution and SENet units, and the number of network layers is 12; the composite scaling factor of the EfficientNet-Lite takes the value of 0.75, and the composite scaling factor is used to adjust the network width, network depth, and resolution; the loss function is constructed based on the OKS function and the CIoU function.
[0077] The OKS function is a scale-normalized loss function that can assign different weights according to the importance of key points. For example, head key points (such as eyes, nose) are more vulnerable to pixel-level errors than body key points (such as shoulders, knees); the CIoU function can significantly improve the accuracy and efficiency of pose detection by comprehensively considering the overlapping area, the distance between the center points, and the aspect ratio. It not only reduces the cases of false detection and missed detection, but also speeds up the convergence rate of the model, making it perform well in complex scenarios and small target detection; by setting the loss function to be constructed based on the OKS function and the CIoU function, that is, weighting the OKS function and the CIoU function, combining the advantages of the OKS function and the CIoU function, greatly improves the training effect of the pose detection model, and thus greatly improves the accuracy of pose detection.
[0078] The step S3 is specifically as follows:
[0079] Based on a preset segmentation ratio, the data set is divided into a training set, a validation set, and a test set. The pose detection model is trained through the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, the learning rate is adaptively adjusted by the LAMB optimizer, and the pose detection model is compressed through knowledge distillation technology, dynamic pruning technology, and model quantization technology; the trained pose detection model is verified through the validation set to judge whether the detection accuracy is greater than a preset accuracy threshold. If not, the verification fails, and the training set is expanded and training continues. If so, the verification is successful; the pose detection model that has passed the verification is tested through the test set to judge whether the confidence level is greater than a preset confidence level threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test is successful, and the training ends;
[0080] Deploy the trained pose detection model to an edge computing device with a device type of router, gateway, or camera.
[0081] The pose detection model is trained with a training set until the loss value of the loss function is less than a preset loss threshold. During the training process, the learning rate is adaptively adjusted by the LAMB optimizer, and the pose detection model is compressed by knowledge distillation technology, dynamic pruning technology, and model quantization technology; the detection accuracy is calculated through a validation set to verify the trained pose detection model, and the confidence level is calculated through a test set to test the successfully verified pose detection model. That is, during the training process of the pose detection model, compression, verification, and testing are continuously performed to effectively balance the model volume and detection accuracy of the pose detection model. Combining with the LAMB optimizer also greatly accelerates the model convergence speed.
[0082] Knowledge distillation technology is a machine learning model compression method aimed at transferring the knowledge of a large model to a small model to improve model performance and generalization ability; the core idea of knowledge distillation is to transform the knowledge of a complex model into a more concise and effective representation, which can reduce the computational complexity and resource requirements while maintaining high performance. Dynamic pruning technology aims to remove parts of the neural network that have little impact on model performance (such as accuracy), such as neurons, connections (weights), etc., thereby reducing the complexity of the model and computational resource requirements. Model quantization technology reduces the storage space and computational complexity of the model by converting model parameters from high precision (such as 32-bit floating-point numbers) to low precision (such as 8-bit integers or 16-bit floating-point numbers).
[0083] In step S4, the object detection model is EdgeYOLO-Tiny;
[0084] By setting the object detection model to EdgeYOLO-Tiny, EdgeYOLO-Tiny is a lightweight object detection model designed specifically for edge computing devices, with high inference speed and small model volume, facilitating deployment on resource-constrained edge computing devices.
[0085] In step S5, the light compensation is specifically: enhancing the contrast of low-light regions in the real-time pose image through the CLAHE algorithm for light compensation;
[0086] The dynamic noise reduction is specifically: dynamically reducing the noise of the real-time pose image through the cooperation of the Kalman filtering algorithm and the bilateral filtering algorithm.
[0087] The ROI optimization and cropping are achieved through a mask image, an image segmentation algorithm (such as GrabCut), GraphCut-based segmentation, or ROI selection based on Bayesian optimization.
[0088] Step S7 is specifically:
[0089] The real-time recording includes at least the detection log of the optimized pose image, the pose detection result, and the detection time, obtains the device serial number of the edge computing device itself, extracts image data and text data from the detection log, separates the EXIF information and pixel data in the image data, performs a hashing calculation on the EXIF information and the device serial number through the SHA3-256 algorithm to obtain a first data fingerprint, and performs a hashing calculation on the text data and the device serial number through the SHA3-512 algorithm to obtain a second data fingerprint;
[0090] Set a row-column permutation rule, perform a row-column permutation operation on the pixel data based on the row-column permutation rule to obtain encrypted pixel data, and encrypt the EXIF information, the encrypted pixel data, and the first data fingerprint through the IEDA algorithm to obtain first encrypted data;
[0091] Set a cyclic shift rule, shift the characters of the text data based on the cyclic shift rule to obtain encrypted text data, and encrypt the encrypted text data and the second data fingerprint through the RC6 algorithm to obtain second encrypted data;
[0092] Encrypt the first encrypted data and the second encrypted data into an encrypted log through the AES-KW algorithm, upload the encrypted log to the server for storage in real time through the SSL protocol, and delete the encrypted log locally.
[0093] By obtaining the device serial number of the edge computing device itself, extracting image data and text data from the detection log, separating the EXIF information and pixel data in the image data, performing hash calculation on the EXIF information and the device serial number through the SHA3-256 algorithm to obtain the first data fingerprint, and performing hash calculation on the text data and the device serial number through the SHA3-512 algorithm to obtain the second data fingerprint; performing row-column permutation operation on the pixel data based on the row-column permutation rule to obtain encrypted pixel data, and encrypting the EXIF information, the encrypted pixel data, and the first data fingerprint through the IEDA algorithm to obtain the first encrypted data; performing displacement on the characters of the text data based on the cyclic shift rule to obtain encrypted text data, and encrypting the encrypted text data and the second data fingerprint through the RC6 algorithm to obtain the second encrypted data; encrypting the first encrypted data and the second encrypted data into an encrypted log through the AES-KW algorithm, and uploading the encrypted log to the server for storage in real time through the SSL protocol; that is, encrypting the detection log differently based on the data types (image data and text data), and finally merging the encryption, and the image data and the text data respectively combine different encryption algorithms and data transformation rules, and the SSL protocol is a secure transmission protocol, and at least nine security measures are taken before and after (device serial number, first data fingerprint, second data fingerprint, row-column permutation rule, IEDA algorithm, cyclic shift rule, RC6 algorithm, AES-KW algorithm, SSL protocol), avoiding the detection log from being stolen and tampered with in plain text, thereby greatly improving the security of the detection log transmission and storage, and effectively improving the reliability of traceability.
[0094] A preferred embodiment of an attitude detection system for routers, gateways, and cameras according to the present invention includes the following modules:
[0095] A dataset construction module, configured to obtain a large number of historical attitude images, perform preprocessing on each of the historical attitude images including at least noise reduction, cropping, and size unification, perform attitude annotation on each of the preprocessed historical attitude images, and then perform sample augmentation including at least rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, saturation adjustment, hue adjustment, and adding random noise on each of the annotated historical attitude data to construct a dataset;
[0096] By performing sample augmentation including at least rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, saturation adjustment, hue adjustment, and adding random noise on each of the annotated historical attitude data, the sample size of the dataset is effectively augmented, thereby greatly improving the generalization performance of the attitude detection model.
[0097] A pose detection model creation module, which is used to create a pose detection model based on MobileNetV3, EfficientNet-Lite, a feature fusion module, and an output module, and set the loss function of the pose detection model;
[0098] The input end of the feature fusion module is connected to MobileNetV3 and EfficientNet-Lite, and the output end is connected to the output module; both MobileNetV3 and EfficientNet-Lite are used to extract pose features from pose images; the feature fusion module is used to fuse the pose features extracted by MobileNetV3 and EfficientNet-Lite to obtain 128-dimensional fused features; the output module is used to output pose detection results based on the fused features;
[0099] MobileNetV3 is a lightweight neural network architecture designed specifically for mobile devices and embedded systems, aiming to improve efficiency and performance; EfficientNet-Lite is a lightweight version of the EfficientNet series, optimized for mobile and edge devices, with advantages such as efficient computing performance, low power consumption, fast inference, and easy deployment.
[0100] A pose detection model deployment module, which is used to train the pose detection model through the data set and the loss function, compress the pose detection model during the training process, and deploy the trained pose detection model to edge computing devices of device types such as routers, gateways, or cameras;
[0101] A real-time pose image extraction module, which is used for the edge computing device to collect monitoring videos in real time through a camera, detect image frames with people in the monitoring videos through a pre-trained object detection model, and then extract real-time pose images from the monitoring videos;
[0102] A real-time pose image processing module, which is used for the edge computing device to perform preprocessing on the real-time pose image, including at least light compensation, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, to obtain an optimized pose image;
[0103] By performing preprocessing on the real-time pose image, including at least light compensation, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, before pose detection, the image noise is effectively reduced and the image quality is improved, thereby greatly improving the accuracy of pose detection.
[0104] A pose detection module, which is used for the edge computing device to input the optimized pose image into the pose detection model to obtain a pose detection result, and superimpose the pose detection result on the video stream;
[0105] The detection log management module is used to record in real time the detection logs including at least the optimized pose image, the pose detection result, and the detection time, encrypt the detection logs into encrypted logs, upload the encrypted logs to the server for storage, and delete the encrypted logs locally.
[0106] By recording in real time the detection logs including at least the optimized pose image, the pose detection result, and the detection time, encrypting the detection logs into encrypted logs and uploading them to the server for storage, it is convenient for later traceability.
[0107] In the pose detection model creation module, the MobileNetV3 is constructed based on depthwise separable convolution and SENet units, and the number of network layers is 12; the compound scaling factor of the EfficientNet-Lite takes the value of 0.75, and the compound scaling factor is used to adjust the network width, network depth, and resolution; the loss function is constructed based on the OKS function and the CIoU function.
[0108] The OKS function is a scale-normalized loss function that can assign different weights according to the importance of key points. For example, head key points (such as eyes, nose) are more susceptible to pixel-level errors than body key points (such as shoulders, knees); the CIoU function can significantly improve the accuracy and efficiency of pose detection by comprehensively considering the overlapping area, the distance between the center points, and the aspect ratio. It not only reduces the cases of false detection and missed detection, but also speeds up the convergence rate of the model, making it perform well in complex scenarios and small target detection; by setting the loss function to be constructed based on the OKS function and the CIoU function, that is, weighting the OKS function and the CIoU function, combining the advantages of the OKS function and the CIoU function, the training effect of the pose detection model is greatly improved, and thus the accuracy of pose detection is greatly improved.
[0109] The pose detection model deployment module is specifically used for:
[0110] Divide the dataset into a training set, a validation set, and a test set based on a preset segmentation ratio. Train the pose detection model using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, adaptively adjust the learning rate through the LAMB optimizer, and compress the pose detection model using knowledge distillation technology, dynamic pruning technology, and model quantization technology. Verify the trained pose detection model using the validation set to determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the verification fails, and the training set is expanded and training continues. If so, the verification is successful. Test the pose detection model that has passed verification using the test set to determine whether the confidence level is greater than a preset confidence threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test is successful, and training ends;
[0111] Deploy the trained pose detection model to an edge computing device with a device type of router, gateway, or camera.
[0112] Train the pose detection model using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, adaptively adjust the learning rate through the LAMB optimizer, and compress the pose detection model using knowledge distillation technology, dynamic pruning technology, and model quantization technology. Calculate the detection accuracy using the validation set to verify the trained pose detection model, and calculate the confidence level using the test set to test the pose detection model that has passed verification. That is, during the training process of the pose detection model, compression, verification, and testing are continuously performed to effectively balance the model volume and detection accuracy of the pose detection model. Combining the LAMB optimizer also greatly accelerates the model convergence speed.
[0113] Knowledge distillation technology is a machine learning model compression method aimed at transferring the knowledge of a large model to a small model to improve the model performance and generalization ability. The core idea of knowledge distillation is to transform the knowledge of a complex model into a more concise and effective representation, enabling it to maintain high performance while reducing computational complexity and resource requirements. Dynamic pruning technology aims to remove parts of the neural network that have little impact on the model performance (such as accuracy), such as neurons, connections (weights), etc., thereby reducing the model complexity and computational resource requirements. Model quantization technology reduces the storage space and computational complexity of the model by converting the model parameters from high precision (such as 32-bit floating-point numbers) to low precision (such as 8-bit integers or 16-bit floating-point numbers).
[0114] In the real-time pose image extraction module, the object detection model is EdgeYOLO-Tiny;
[0115] By setting the target detection model as EdgeYOLO-Tiny, which is a lightweight target detection model designed specifically for edge computing devices, with high inference speed and small model size, facilitating deployment on resource-constrained edge computing devices.
[0116] In the real-time pose image processing module, the specific illumination compensation is as follows: enhancing the contrast of low-illumination regions in the real-time pose image through the CLAHE algorithm for illumination compensation.
[0117] The specific dynamic noise reduction is as follows: dynamically reducing the noise of the real-time pose image through the cooperation of the Kalman filtering algorithm and the bilateral filtering algorithm.
[0118] The ROI optimization and cropping is achieved through a mask image, an image segmentation algorithm (such as GrabCut), segmentation based on GraphCut, or ROI selection based on Bayesian optimization.
[0119] The detection log management module is specifically used for:
[0120] Real-time recording of detection logs including at least the optimized pose image, pose detection results, and detection time, obtaining the device serial number of the edge computing device itself, extracting image data and text data from the detection logs, separating the EXIF information and pixel data in the image data, calculating the first data fingerprint by hashing the EXIF information and the device serial number through the SHA3-256 algorithm, and calculating the second data fingerprint by hashing the text data and the device serial number through the SHA3-512 algorithm;
[0121] Setting a row-column permutation rule, performing a row-column permutation operation on the pixel data based on the row-column permutation rule to obtain encrypted pixel data, and encrypting the EXIF information, encrypted pixel data, and the first data fingerprint through the IEDA algorithm to obtain the first encrypted data;
[0122] Setting a cyclic shift rule, shifting the characters of the text data based on the cyclic shift rule to obtain encrypted text data, and encrypting the encrypted text data and the second data fingerprint through the RC6 algorithm to obtain the second encrypted data;
[0123] Encrypting the first encrypted data and the second encrypted data into an encrypted log through the AES-KW algorithm, uploading the encrypted log to the server for storage in real time through the SSL protocol, and clearing the encrypted log locally.
[0124] By obtaining the device serial number of the edge computing device itself, extracting image data and text data from the detection log, separating the EXIF information and pixel data in the image data, performing hash calculation on the EXIF information and the device serial number through the SHA3-256 algorithm to obtain the first data fingerprint, and performing hash calculation on the text data and the device serial number through the SHA3-512 algorithm to obtain the second data fingerprint; performing row-column permutation operation on the pixel data based on the row-column permutation rule to obtain encrypted pixel data, and encrypting the EXIF information, the encrypted pixel data, and the first data fingerprint through the IEDA algorithm to obtain the first encrypted data; performing character displacement on the text data based on the cyclic shift rule to obtain encrypted text data, and encrypting the encrypted text data and the second data fingerprint through the RC6 algorithm to obtain the second encrypted data; encrypting the first encrypted data and the second encrypted data into an encrypted log through the AES-KW algorithm, and uploading the encrypted log to the server for storage in real time through the SSL protocol; that is, encrypting the detection log differently based on the data types (image data and text data), and finally merging and encrypting. Moreover, the image data and the text data respectively combine different encryption algorithms and data transformation rules, and the SSL protocol is a secure transmission protocol. At least nine security measures are taken before and after (device serial number, first data fingerprint, second data fingerprint, row-column permutation rule, IEDA algorithm, cyclic shift rule, RC6 algorithm, AES-KW algorithm, SSL protocol) to prevent the detection log from being stolen and tampered with in plain text, thereby greatly improving the security of the detection log transmission and storage, and effectively improving the reliability of traceability.
[0125] In summary, the advantages of the present invention are as follows:
[0126] 1. By obtaining a large number of historical pose images, preprocessing, annotating, and augmenting the samples of each historical pose image, a dataset is constructed. Then, a pose detection model is created based on MobileNetV3, EfficientNet-Lite, a feature fusion module, and an output module, and the loss function of the pose detection model is set. Next, the pose detection model is trained using the dataset and the loss function. During the training process, the pose detection model is compressed, and the trained pose detection model is deployed to an edge computing device of the device type of router, gateway, or camera. Then, the edge computing device collects monitoring videos in real time through a camera, detects the image frames with humans in the monitoring videos through a pre-trained object detection model, extracts real-time pose images from the monitoring videos, performs preprocessing including at least illumination compensation, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization on the real-time pose images to obtain optimized pose images, inputs the optimized pose images into the pose detection model to obtain pose detection results, and superimposes and displays the pose detection results on the video stream. At the same time, the detection logs including at least the optimized pose images, pose detection results, and detection time are recorded in real time, the detection logs are encrypted into encrypted logs and uploaded to the server for storage, and the local encrypted logs are cleared. Since the pose detection model is constructed based on lightweight MobileNetV3 and EfficientNet-Lite, and the number of network layers of MobileNetV3 is reduced to 12 layers, and the compound scaling factor of EfficientNet-Lite is set to 0.75 to reduce the network width, network depth, and resolution, and the pose detection model is compressed through knowledge distillation technology, dynamic pruning technology, and model quantization technology during the training process, effectively reducing the number of parameters (complexity) and volume of the pose detection model, and locally deleting the uploaded encrypted logs in real time to reduce the storage overhead, effectively reducing the computing overhead of the edge computing device. Through cross-detection (fusion detection), image preprocessing, and an improved loss function using the dual channels of MobileNetV3 and EfficientNet-Lite, the generalization ability and detection performance of pose detection are effectively improved, ultimately greatly reducing the resource requirements for pose detection, greatly improving the accuracy and real-time performance of pose detection, and greatly reducing the power consumption of the edge computing device for pose detection.
[0127] 2. By performing sample augmentation on each annotated historical pose data including at least rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, saturation adjustment, hue adjustment, and adding random noise, the sample size of the dataset is effectively augmented, and thus the generalization performance of the pose detection model is greatly improved.
[0128] 3. Before attitude detection, preprocess the real-time attitude image by at least including light compensation, dynamic noise reduction, ROI optimization and cropping, grayscale conversion, and normalization, effectively reducing image noise and improving image quality, thereby greatly improving the accuracy of attitude detection.
[0129] 4. By real-time recording the detection log including at least the optimized attitude image, attitude detection result, and detection time, encrypt the detection log into an encrypted log and upload it to the server for storage, which is convenient for later traceability.
[0130] 5. The OKS function is a scale-normalized loss function that can assign different weights according to the importance of key points. For example, head key points (such as eyes, nose) are more easily punished for pixel-level errors than body key points (such as shoulders, knees); the CIoU function can significantly improve the accuracy and efficiency of attitude detection by comprehensively considering the overlapping area, center point distance, and aspect ratio. It not only reduces the cases of false detection and missed detection, but also speeds up the convergence rate of the model, making it perform well in complex scenarios and small target detection; by setting the loss function to be constructed based on the OKS function and the CIoU function, that is, weighting the OKS function and the CIoU function, combining the advantages of the OKS function and the CIoU function, greatly improves the training effect of the attitude detection model, and thus greatly improves the accuracy of attitude detection.
[0131] 6. Train the attitude detection model with the training set until the loss value of the loss function is less than the preset loss threshold. During the training process, adaptively adjust the learning rate through the LAMB optimizer, and compress the attitude detection model through knowledge distillation technology, dynamic pruning technology, and model quantization technology; calculate the detection accuracy through the validation set to verify the trained attitude detection model, and calculate the confidence through the test set to test the attitude detection model that has passed the verification. That is, during the training process of the attitude detection model, continuously compress, verify, and test to effectively balance the model volume and detection accuracy of the attitude detection model. Combining with the LAMB optimizer also greatly accelerates the model convergence rate.
[0132] 7. By setting the object detection model to EdgeYOLO-Tiny, EdgeYOLO-Tiny is a lightweight object detection model designed for edge computing devices, with high inference speed and small model volume, which is convenient for deployment on resource-constrained edge computing devices.
[0133] 8. By obtaining the device serial number of the edge computing device itself, extracting image data and text data from the detection log, separating the EXIF information and pixel data in the image data, performing hash calculation on the EXIF information and the device serial number through the SHA3-256 algorithm to obtain the first data fingerprint, and performing hash calculation on the text data and the device serial number through the SHA3-512 algorithm to obtain the second data fingerprint; performing row-column permutation operation on the pixel data based on the row-column permutation rule to obtain encrypted pixel data, and encrypting the EXIF information, the encrypted pixel data, and the first data fingerprint through the IEDA algorithm to obtain the first encrypted data; performing character displacement on the text data based on the cyclic shift rule to obtain encrypted text data, and encrypting the encrypted text data and the second data fingerprint through the RC6 algorithm to obtain the second encrypted data; encrypting the first encrypted data and the second encrypted data into an encrypted log through the AES-KW algorithm, and uploading the encrypted log to the server for storage in real time through the SSL protocol; that is, encrypting the detection log differently based on the data types (image data and text data), and finally merging and encrypting. Moreover, the image data and the text data respectively combine different encryption algorithms and data transformation rules, and the SSL protocol is a secure transmission protocol. At least nine security measures are taken before and after (device serial number, first data fingerprint, second data fingerprint, row-column permutation rule, IEDA algorithm, cyclic shift rule, RC6 algorithm, AES-KW algorithm, SSL protocol) to prevent the detection log from being stolen and tampered with in plain text, thereby greatly improving the security of the detection log transmission and storage, and effectively enhancing the reliability of traceability.
[0134] 9. By performing sample augmentation operations including brightness adjustment, contrast adjustment, saturation adjustment, and hue adjustment, and then training the pose detection model using the dataset after sample augmentation, and performing light compensation on the real-time pose image before pose recognition, effectively overcoming the problem of the decrease in detection accuracy under traditional dynamic lighting conditions, and greatly improving the accuracy of pose detection.
[0135] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.
Claims
1. A method for attitude detection of routers, gateways, and cameras, characterized in that: The steps are as follows: Step S1: Obtain a large number of historical pose images, perform preprocessing on each of the historical pose images, including at least noise reduction, cropping, and size unification. Perform pose annotation on each of the preprocessed historical pose images, and then perform sample augmentation on each of the annotated historical pose data, including at least rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, saturation adjustment, hue adjustment, and adding random noise, to construct a dataset; Step S2: Create a pose detection model based on MobileNetV3, EfficientNet-Lite, a feature fusion module, and an output module, and set the loss function of the pose detection model; The input end of the feature fusion module is connected to MobileNetV3 and EfficientNet-Lite, and the output end is connected to the output module; both MobileNetV3 and EfficientNet-Lite are used to extract pose features from pose images; the feature fusion module is used to fuse the pose features extracted by MobileNetV3 and EfficientNet-Lite to obtain 128-dimensional fused features; the output module is used to output pose detection results based on the fused features; Step S3: Train the pose detection model through the dataset and the loss function. During the training process, compress the pose detection model, and deploy the trained pose detection model to an edge computing device with a device type of router, gateway, or camera; Step S4: The edge computing device uses a camera to collect surveillance videos in real time, detects image frames with people in the surveillance videos through a pre-trained object detection model, and then extracts real-time pose images from the surveillance videos; Step S5: The edge computing device performs preprocessing on the real-time pose images, including at least light compensation, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, to obtain optimized pose images; Step S6: The edge computing device inputs the optimized pose images into the pose detection model to obtain pose detection results, and superimposes the pose detection results on the video stream; Step S7: Record detection logs in real time, including at least the optimized pose images, pose detection results, and detection times. Encrypt the detection logs into encrypted logs, upload the encrypted logs to a server for storage, and delete the encrypted logs locally.
2. The attitude detection method for a router, gateway, and camera according to claim 1, characterized in that: In step S2, MobileNetV3 is constructed based on depthwise separable convolution and SENet units, and the number of network layers is 12; the compound scaling factor of EfficientNet-Lite takes a value of 0.75, and the compound scaling factor is used to adjust network width, network depth, and resolution; the loss function is constructed based on the OKS function and the CIoU function.
3. The attitude detection method for a router, gateway, and camera according to claim 1, characterized in that: Specifically, step S3 is as follows: The data set is divided into a training set, a validation set, and a test set based on a preset segmentation ratio. The pose detection model is trained using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, the learning rate is adaptively adjusted by the LAMB optimizer, and the pose detection model is compressed using knowledge distillation technology, dynamic pruning technology, and model quantization technology; the trained pose detection model is verified using the validation set to determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the verification fails, and the training set is expanded and training continues. If so, the verification is successful; the pose detection model that has passed the verification is tested using the test set to determine whether the confidence level is greater than a preset confidence level threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test is successful, and training ends; The trained pose detection model is deployed to an edge computing device with a device type of router, gateway, or camera.
4. The attitude detection method for a router, gateway, and camera according to claim 1, characterized in that: In step S4, the object detection model is EdgeYOLO-Tiny; In step S5, the light compensation is specifically: enhancing the contrast of the low-light region in the real-time pose image through the CLAHE algorithm for light compensation; The dynamic noise reduction is specifically: dynamically reducing the noise of the real-time pose image through the cooperation of the Kalman filtering algorithm and the bilateral filtering algorithm.
5. The attitude detection method for a router, gateway, and camera according to claim 1, wherein: Step S7 is specifically: The detection log including at least the optimized pose image, pose detection result, and detection time is recorded in real time. The device serial number of the edge computing device itself is obtained. The image data and text data are extracted from the detection log. The EXIF information and pixel data in the image data are separated. The first data fingerprint is obtained by performing a hash calculation on the EXIF information and the device serial number through the SHA3-256 algorithm. The second data fingerprint is obtained by performing a hash calculation on the text data and the device serial number through the SHA3-512 algorithm; A row-column permutation rule is set, and the row-column permutation operation is performed on the pixel data based on the row-column permutation rule to obtain encrypted pixel data. The EXIF information, encrypted pixel data, and the first data fingerprint are encrypted through the IEDA algorithm to obtain the first encrypted data; A cyclic shift rule is set, and the characters of the text data are shifted based on the cyclic shift rule to obtain encrypted text data. The encrypted text data and the second data fingerprint are encrypted through the RC6 algorithm to obtain the second encrypted data; The first encrypted data and the second encrypted data are encrypted into an encrypted log through the AES-KW algorithm. The encrypted log is uploaded to the server for storage in real time through the SSL protocol, and the local encrypted log is cleared.
6. An attitude detection system for routers, gateways, and cameras, characterized in that: Includes the following modules: A dataset construction module, configured to obtain a large number of historical pose images, perform preprocessing on each of the historical pose images, including at least noise reduction, cropping, and size unification, perform pose annotation on each of the preprocessed historical pose images, and then perform sample augmentation on each of the annotated historical pose data, including at least rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, saturation adjustment, hue adjustment, and adding random noise, to construct a dataset; A pose detection model creation module, configured to create a pose detection model based on MobileNetV3, EfficientNet-Lite, a feature fusion module, and an output module, and set a loss function of the pose detection model; The input end of the feature fusion module is connected to MobileNetV3 and EfficientNet-Lite, and the output end is connected to the output module; both MobileNetV3 and EfficientNet-Lite are used to extract pose features from pose images; the feature fusion module is used to fuse the pose features extracted by MobileNetV3 and EfficientNet-Lite to obtain 128-dimensional fused features; the output module is used to output a pose detection result based on the fused features; A pose detection model deployment module, configured to train the pose detection model through the dataset and the loss function, compress the pose detection model during the training process, and deploy the trained pose detection model to an edge computing device of which the device type is a router, a gateway, or a camera; A real-time pose image extraction module, configured to enable the edge computing device to collect a monitoring video in real time through a camera, detect an image frame with a human body in the monitoring video through a pre-trained object detection model, and then extract a real-time pose image from the monitoring video; A real-time pose image processing module, configured to enable the edge computing device to perform preprocessing on the real-time pose image, including at least light compensation, dynamic noise reduction, ROI optimized cropping, grayscale conversion, and normalization, to obtain an optimized pose image; A pose detection module, configured to enable the edge computing device to input the optimized pose image into the pose detection model to obtain a pose detection result, and superimpose the pose detection result on the video stream; A detection log management module, configured to record in real time a detection log including at least the optimized pose image, the pose detection result, and the detection time, encrypt the detection log into an encrypted log, upload the encrypted log to a server for storage, and delete the encrypted log locally.
7. The attitude detection system for routers, gateways, and cameras according to claim 6, characterized in that: In the pose detection model creation module, MobileNetV3 is constructed based on depthwise separable convolution and SENet units, and the number of network layers is 12; the compound scaling factor of EfficientNet-Lite takes a value of 0.75, and the compound scaling factor is used to adjust the network width, network depth, and resolution; the loss function is constructed based on the OKS function and the CIoU function.
8. The attitude detection system for a router, gateway, and camera according to claim 6, characterized in that: The pose detection model deployment module is specifically configured to: Divide the dataset into a training set, a validation set, and a test set based on a preset segmentation ratio. Train the pose detection model using the training set until the loss value of the loss function is less than a preset loss threshold. During the training process, adaptively adjust the learning rate using the LAMB optimizer, and compress the pose detection model using knowledge distillation technology, dynamic pruning technology, and model quantization technology. Validate the trained pose detection model using the validation set to determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the validation fails, and the training set is expanded and training continues. If so, the validation is successful. Test the pose detection model with successful validation using the test set to determine whether the confidence level is greater than a preset confidence threshold. If not, the test fails, and the training set is expanded and training continues. If so, the test is successful, and the training ends. Deploy the trained pose detection model to an edge computing device with a device type of router, gateway, or camera.
9. The attitude detection system for a router, gateway, and camera according to claim 6, wherein: In the real-time pose image extraction module, the object detection model is EdgeYOLO-Tiny. In the real-time pose image processing module, the light compensation is specifically: enhance the contrast of the low-light area in the real-time pose image through the CLAHE algorithm for light compensation. The dynamic noise reduction is specifically: perform dynamic noise reduction on the real-time pose image through the cooperation of the Kalman filter algorithm and the bilateral filter algorithm.
10. A posture detection system for routers, gateways, and cameras according to claim 6, characterized in that: The detection log management module is specifically used for: Real-time record the detection log including at least the optimized pose image, pose detection result, and detection time, obtain the device serial number of the edge computing device itself, extract image data and text data from the detection log, separate the EXIF information and pixel data in the image data, calculate the first data fingerprint by hashing the EXIF information and the device serial number through the SHA3-256 algorithm, and calculate the second data fingerprint by hashing the text data and the device serial number through the SHA3-512 algorithm; Set a row-column permutation rule, perform a row-column permutation operation on the pixel data based on the row-column permutation rule to obtain encrypted pixel data, and encrypt the EXIF information, encrypted pixel data, and the first data fingerprint through the IEDA algorithm to obtain the first encrypted data; Set a cyclic shift rule, shift the characters of the text data based on the cyclic shift rule to obtain encrypted text data, and encrypt the encrypted text data and the second data fingerprint through the RC6 algorithm to obtain the second encrypted data; Encrypt the first encrypted data and the second encrypted data into an encrypted log through the AES-KW algorithm, upload the encrypted log to the server for storage in real time through the SSL protocol, and clear the local encrypted log.