Low-resolution real-time key information detection and protection method

By pre-training and using the LRRT-Det object detection network on a low-resolution network, combining 2D-TEM chaotic sequence and ring diffusion algorithm, real-time key information detection and encryption on low-computing equipment is achieved, solving the problems of low-resolution image detection accuracy and speed, and improving information security.

CN120032099AActive Publication Date: 2025-05-23SOUTH CHINA UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510037279.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-23
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The prior art is difficult to realize real-time critical information detection and encryption on low-computing equipment, especially under low-resolution image input, with low detection accuracy and slow speed. At the same time, traditional encryption algorithms are not suitable for low-resolution images.

Method used

The lightweight low-resolution network is pre-trained with a higher resolution, and the LRRT-Det target detection network is proposed. Combined with 2D-TEM chaotic sequence and improved ring diffusion algorithm, the RGB three channels are complexly exchanged to realize the encryption of the detection area and the auxiliary data encryption of the embedded ROI area.

Benefits of technology

The accuracy and speed of object detection are significantly improved under low-resolution image input, and the security of image information is improved through chaotic encryption methods, which is suitable for low-computing equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032099A_ABST
    Figure CN120032099A_ABST
Patent Text Reader

Abstract

The invention discloses a low-resolution real-time key information detection and protection method. The method comprises the following steps: acquiring an image by using a network camera; shuffleNetV2 is used as a backbone network for extraction and sampling, a Ge-Fusion FPN module is used for multi-scale feature fusion, an FCOS target detection head is used, and meanwhile, a data distillation and high-resolution pre-training method is used, so that the detection accuracy of the model under low resolution is improved; meanwhile, a chaotic encryption scheme is used to encrypt the detected key information area, and the key information area is fused to the original image. By applying the technical scheme of the invention, the target detection speed and precision under low resolution can be improved, the method can be applied to the field of intelligent security and protection, and the privacy performance of the network camera is improved. According to the invention, by introducing a chaotic encryption method, real-time key information protection is carried out on the detected image, and the security is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of intelligent security, and in particular relates to a low-resolution real-time key information detection and protection method. Background Art

[0002] Network cameras provide users with a more convenient user experience through cloud server storage, but the interaction between a large amount of data and cloud servers is prone to data leakage. Therefore, how to better protect user data security has become one of the key issues; however, traditional encryption systems, such as data encryption standard (DES), advanced encryption standard (AES), Shamir algorithm (RSA), etc., are not suitable for real-time image encryption due to their large data capacity, similar grayscale values, and large pixel correlation. In addition, these algorithms use complex structures to increase randomness, thereby improving the security of the algorithm. Therefore, traditional encryption algorithms require larger memory and computing power, while network cameras are mainly low-computing devices, and their computing power is not up to the task. At the same time, most existing algorithms are used to encrypt the entire image. Sometimes multimedia data contains a lot of redundant information. Therefore, in order to improve encryption efficiency, avoid waste of resources, and protect privacy in a hierarchical manner, so that people with different permissions can access key data of different levels, and balance the data availability and user privacy of network camera image information, it is necessary to encrypt only the key areas of the image. Key information detection and encryption is divided into two steps: detection and encryption. In the part of target detection, the current mainstream Yolov7, Yolox, Yolov11 (Khanam R, Hussain M.Yolov11: An overview of the key architectural enhancements [J]. arXiv preprint arXiv: 2410.17725, 2024.) and others are all detected under the condition of 640x640 input. In the low-resolution scene of the network camera, for real-time performance, a lower resolution image input is required to maintain a good detection effect. The Nanodet series has good real-time performance, but under the input of 320x320, the accuracy still has room for improvement; in the part of key information encryption, Wei Song (Song, W., Fu, C., Zheng, Y. et al. Protection of image ROI using chaos-based encryption and DCNN-based object detection. Neural Comput&Applic 34, 5743–5756 (2022). https: / / doi.org / 10.1007 / s00521-021-06725-w) proposed an encryption algorithm for key area ROI and embedded the auxiliary data of the key area into the image after encryption. However, this method still needs to be performed on a 1080Ti desktop device, which consumes a lot of computing power and is difficult to implement the entire detection and encryption algorithm on a low-computing embedded device.However, network security cameras require smoother frames. In the security field where network surveillance cameras are used, Aribilola (Aribilola I, Asghar MN, Kanwal N, et al. SecureCam: Selective Detection and Encryption enabled Application for Dynamic Camera Surveillance Videos [J]. IEEE Transactions on Consumer Electronics, 2022.) et al. proposed using H264 encoding to encrypt images in video streams based on pixel differences between consecutive frames, but this encryption method still makes it difficult to detect static key information areas. Therefore, developing a method for real-time key information detection and protection on low-computing power devices is of great significance in the field of intelligent security. Summary of the invention

[0003] To solve the above problems, the present invention proposes a low-resolution real-time key information detection and protection method. Aiming at the problems of low accuracy and slow speed of key information detection under low resolution, the lightweight low-resolution network is pre-trained with higher resolution, which achieves a good accuracy improvement. The LRRT-Det target detection network is proposed to obtain the key information in the network camera. At the same time, a 2D-TEM chaotic sequence is proposed, and the characteristics of the 2D-TEM sequence are combined to improve the two-stage annular diffusion algorithm, perform more complex exchange of the RGB three channels, encrypt the detection area, and encrypt and transmit the auxiliary data of the embedded ROI area, and transmit the key information by embedding the encrypted key information.

[0004] The present invention is achieved by at least one of the following technical solutions.

[0005] A low-resolution real-time key information detection and protection method comprises the following steps:

[0006] (1) Images acquired using a webcam;

[0007] (2) Input the image obtained by the network camera into the trained target detection model, which extracts and fuses features and outputs categories and detection box areas of different scales;

[0008] (3) Use chaotic encryption to encrypt the content in the detection frame in real time, and embed the encrypted area coordinate information into the unencrypted area of ​​the image.

[0009] Furthermore, the target detection model is a low-resolution real-time target detection model LRRT-Det, including a feature extraction backbone network ShuffleNetV2, a generalized feature pyramid GeFusion-FPN network based on attention fusion, and a fully convolutional single-stage FCOS target detection head;

[0010] The feature extraction backbone network is used to extract multi-scale features of the image. GeFusion-FPN uses a feature selector to fuse features of different scales. The FCOS target detection head is used to finally output the regional coordinates and detection category of the detection box.

[0011] Furthermore, the feature extraction backbone network uses ShuffleNetV2 to extract multi-scale features of the image, and inputs three scale feature maps of 1 / 8, 1 / 16, and 1 / 32 into GeFusion-FPN for feature fusion.

[0012] Furthermore, Ge-Fusion FPN uses a feature selector to fuse features of different scales. The fusion formula is:

[0013] C=A×score [:,0] +B×score [:,1] ;

[0014] A and B are the two features of the input Ge-Fusion FPN, score [:,0] is the weight score corresponding to feature A, score [:,1] is the weight score corresponding to feature B, and C is the fusion feature output by Ge-Fusion FPN.

[0015] Furthermore, the training of the target detection model includes the following steps:

[0016] S1, use the high-resolution images of the original dataset to train the LRRT-Det model step by step;

[0017] S2, use high-resolution images and large-scale object detection model Yolov11 for data distillation;

[0018] Furthermore, in step S1, the LRRT-Det model is first trained under the scale of 416x416 image input to obtain training parameters, and the obtained training parameters are used to initialize the target detection LRRT-Det model to guide the training under the 320x320 image input LRRT-Det;

[0019] Data distillation in step S2: Use the Coco2017 and Widerface datasets to train the Yolov11 model with a 640x640 image input scale, and use the output of the Yolov11 model as a soft label. The soft label includes not only the probability distribution of the target category, but also the relative relationship between each category, thereby providing more learning information for the student model;

[0020] The teacher model outputs the category probability distribution and detection box coordinates of each target through the Softmax operation. The output z of the teacher model teacher Soft labels are obtained by smoothing with temperature T:

[0021]

[0022] Among them, z teacher is the output of the teacher model, T is the temperature coefficient, which is used to control the smoothness of the output; the student model minimizes the loss function L total , to achieve the best detection effect, L total Using the hard labels and soft labels of target detection, the losses of the two are weighted and combined, and the total loss function is expressed as:

[0023]

[0024] Among them, λ 1 , 2 , 3 are weighted coefficients, which control the contribution ratio of hard labels, soft labels and detection boxes respectively. represents the predicted category of the student model, and y represents the true category. is the soft output of the teacher model, is the detection box coordinate output by the student model, b is the actual detection box coordinate, is the hard label classification loss, which is used to measure the difference between the class predicted by the student model and the true label, is the soft label classification loss, which measures the difference between the categories predicted by the student model and the soft labels generated by the teacher model, is the detection box regression loss, which is used to measure the difference between the detection box predicted by the student model and the true detection box.

[0025] Furthermore, in step (3), the chaotic sequence iteration equation used is:

[0026] x k+1 =mod(10·tanh(p 1 *y k )·exp(x k +y k ),1);

[0027] y k+1 =mod(10·tanh(x k )·exp(p 2 ·(x k +y k )),1);

[0028] Among them, exp() represents the natural exponential function, tanh() represents the hyperbolic tangent function, and p 1 、p 2 is the control parameter value (0,1), set by the user, y k 、x k is the chaotic value of the current k-th step, and mod(.) represents the modulus operation.

[0029] Furthermore, in step (3), the encryption step is as follows: for each detection frame of size w*h, w and h are the number of rows and columns of the detection frame area respectively, firstly, different keys are associated according to the detection category, and then the chaotic sequence is iterated to divide the sequence into a scrambled part and a diffusion part, the image is scrambled using the elementary matrix row and column transformation method, the pixels are diffused using the XOR method, and finally, the encrypted image replaces the position of the detection frame in the original image.

[0030] Furthermore, in step (3), information embedding is performed by using a digital watermarking method.

[0031] A computer device of the present invention comprises: a memory and a processor and a computer program stored in the memory. When the computer program is executed on the processor, a low-resolution real-time critical information detection and protection method as described in any one of claims 1 to 9 is implemented.

[0032] The present invention has the following beneficial effects compared with the prior art:

[0033] 1. The present invention introduces a weight allocation strategy through the feature fusion module (FPN module) in the conventional target detection algorithm to become a generalized feature pyramid based on attention fusion (GeFusion-FPN), so as to better fuse feature information of different scales and improve the accuracy of lightweight target detection under low-resolution image input in an embedded environment.

[0034] 2. The present invention introduces a chaotic encryption method to protect key information of the detection image in real time, thereby improving security. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A flowchart of a low-resolution real-time key information detection and protection method according to an embodiment of the present invention;

[0036] Figure 2This is a schematic diagram of the training of the target detection network LRRT-Det according to an embodiment of the present invention;

[0037] Figure 3 This is a schematic diagram of the structure of the target detection network LRRT-Det according to an embodiment of the present invention;

[0038] Figure 4 It is a structural schematic diagram of a data fusion module in a generalized feature pyramid GeFusion-FPN based on attention fusion according to an embodiment of the present invention;

[0039] Figure 5 This is a chaos algorithm encryption flow chart of an embodiment of the present invention. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the technical scheme in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. It should be understood that the specific practicality described is to explain the present application and is not used to limit the present application. The specific implementation methods of the present invention are further described below in conjunction with the drawings:

[0041] like Figure 1 As shown, this example provides a key information detection and protection method for low-resolution scenes, including the following steps:

[0042] (1) Images acquired using a webcam.

[0043] (2) The image obtained by the network camera is input into the trained target detection model LRRT-Det. The target detection model LRRT-Det uses the detection frame to extract image features, fuses the extracted image features, and outputs categories and detection frame areas of different scales respectively.

[0044] The target detection model is a low-resolution real-time target detection model (LRRT-Det model) including a feature extraction backbone network ShfflenetV2, a generalized feature pyramid GeFusion-FPN (generalized aggregate feature fusion module) network based on attention fusion, and a FCOS (full convolution single stage) target detection head; wherein GeFusion-FPN uses a feature selector to perform attention fusion on features of different scales to improve target detection accuracy.

[0045] As an embodiment, the feature extraction backbone network uses ShuffleNetV2 to extract multi-scale features of the image, and inputs the multi-scale feature maps into the generalized feature pyramid network based on attention fusion for feature fusion.

[0046] Specifically, first, the feature extraction backbone network samples the input 320x320 image by 8, 16, and 32 times to generate features of three scales: 40x40, 20x20, and 0x10, which are then input into the Ge-Fusion FPN network.

[0047] Ge-Fusion FPN introduces Figure 3 The weight generator of Ge-Fusion FPN performs attention fusion on features of different scales to improve the accuracy of target detection. The key idea of ​​Ge-Fusion FPN is to further optimize the representation of features by utilizing multi-scale features, combining adaptive weight selection and attention mechanism, thereby improving the performance of target detection models in different scales and complex scenarios.

[0048] The Ge-Fusion network uses the attention feature fusion method to process the input of ShufflenetV2 and the down-sampling part of FPN on the original feature fusion module (FPN) network, such as Figure 4 As shown in the figure, assuming that the two input features are A and B, and the fusion feature is C, the weight generator module first concatenates the two data to be fused A and B in the channel dimension, then extracts the features through three convolutional layers, and finally obtains the weighted weight Score of each channel through softmax. The features A and B are weightedly fused according to the weight of the Score. The fusion formula is:

[0049] C=A×score [:,0] +B×score [:,1] ;

[0050] A and B are the two features of the input Ge-Fusion FPN, score [:,0] is the weight score corresponding to input A, score [:,1] is the weight score corresponding to input B, and C is the fusion feature output by Ge-Fusion FPN. Finally, the model's learning of details is further improved through bottom-up feature fusion.

[0051] Finally, the fused data of the three scales are sent to the detection head. The detection head is divided into two branches: classification and regression. Each branch first has two convolutional layers, and then output convolutional layers for classification and regression, which output categories and detection box areas of different scales respectively.

[0052] As an example, Figure 2As shown, the training method of the target detection model LRRT-Det provided in this example includes the following steps:

[0053] 1) Use a webcam to obtain key data in various scenarios as a data set, such as facial information, computer screen data, etc., and mark the data location and category information, and divide the training set and validation set into a ratio of 9:1.

[0054] 2) Train the Yolov11[3] model based on the training set as the teacher network to guide the training.

[0055] 3) Use the teacher network for data distillation to guide the training of the LRRT-Det model under high resolution, while ensuring that the LRRT-Det model has low parameters and low computational complexity, and improves the high accuracy of the model under low resolution;

[0056] 4) Use the high-resolution LRRT-Det model as the initial parameters and use the teacher network to train the low-resolution LRRT-Det model. Figure 3 As shown, the structure of the final low-resolution detection model LRRT-Det is shown in the figure.

[0057] In distillation, the teacher model obtains the output of the category probability distribution and detection box coordinates of each target through the normalized exponential function (Softmax) operation. To obtain the soft label, the output z of the teacher model is teacher Smoothing by temperature T makes the information transfer between categories richer:

[0058]

[0059] Among them, z teacher is the soft label output by the teacher model, T is the temperature coefficient, which controls the smoothness of the output, It is the soft output after the temperature coefficient weighting and Softmax operation. The student model LRRT-Det is achieved by minimizing L total Loss function, to achieve the best detection effect, L total Using the hard label (true label) and soft label (output of the teacher model) of target detection, the losses of the two are weighted and combined, and the total loss function can be expressed as:

[0060]

[0061] Among them, λ 1 , 2 , 3 are weighted coefficients, which control the contribution ratio of hard labels, soft labels and detection boxes respectively. represents the predicted category of the student model, and y represents the true category. is the soft output of the teacher model, is the detection box coordinate output by the student model, b is the actual detection box coordinate, is the hard label classification loss, which measures the difference between the class predicted by the student model and the true label. is the soft label classification loss, which measures the difference between the categories predicted by the student model and the soft labels generated by the teacher model. is the detection box regression loss, which is used to measure the difference between the detection box predicted by the student model and the true detection box.

[0062] (3) Use the chaotic encryption method to encrypt the content in the detection frame in real time, and use the steganography technology to embed the coordinate information of the encrypted area into the unencrypted area of ​​the image.

[0063] After obtaining the key information image, such as Figure 5 The chaotic encryption method is used to protect the information security of the image, which includes the sequence generation part, the scrambling part and the diffusion part:

[0064] 1) For each ROI area, firstly, according to the detected key information area of ​​size w*h, where w and h are the number of rows and columns of the detection box area, respectively, then use the chaotic sequence iteration times, we get w*h+1000+w+h valid chaotic sequences, and then remove the first 1000 chaotic sequences to eliminate the influence of the initial value and increase the disorder.

[0065] 2) In the scrambling part, first take the first w sequences x = {x k}, where k = 1, 2, 3, ..., w, where x k (x) k is the value of the first w chaotic sequences, and a new descending matrix x′={x′ k}where k=1,2,3,......,w,x′ k is x k The new sequence after the elements are arranged in descending order, and at the same time determine {x k} to {x k}, forming a set of mapping addresses T = {t k}, t k For {x k The kth element x in k In {x′ k}, using {t k}Generate a w*w elementary matrix P:

[0066]

[0067] Where P (m,n) For the elements of the matrix P with m rows and n columns, we then use h sequences from (w+1) to (w+h) and use the mapping {t k} method to generate an h*h elementary matrix Q, and scramble the RGB color image, that is, the pixel matrix A of size w*h, by means of matrix multiplication P*A*Q to perform elementary row and column transformations.

[0068] 3) In the diffusion part, firstly, unused w*h sequences are taken to form a set X = {x k1}where k1=1,2,3,......,w*h,where x k1 is the value of w*h chaotic sequences required for the diffusion operation, X is x k1 The set of components, and for each x k1 There is X k1 =x k1 *2 24 , where X k1 is x k1 The amplified chaotic sequence value is used to extract the last three sequences from the first 1000 sequences using the user key as s 1 、s 2 、s 3 Facilitates subsequent operations.

[0069] Pixel(1,1)=SR(1,1)+SG(1,1)*2 8 +SB(1,1)*2 16 ;

[0070]

[0071] SR(1,1)=mod(Pixel(1,1),2 8 );

[0072] SG(1,1)=mod(div(Pixel(1,1),2 8 ),2 8 );

[0073] SB(1,1)=mod(div(Pixel(1,1),2 16 ),2 8 );

[0074] Where SR(1,1) is the R channel component of the image in the RGB color space with the coordinates (1,1), SG(1,1) is the G channel component of the image in the RGB color space with the coordinates (1,1), SB(1,1) is the B channel component of the image in the RGB color space with the coordinates (1,1), and X[1] is the current chaotic sequence value. Pixel(1,1) is the pixel value after SR(1,1), SG(1,1), and SB(1,1) are concatenated. mod(div(.),2 8 ) represents taking 2 for div(.) 8 The modulus, that is, div(.) divided by 2 8 The remainder of div(Pixel(1,1),2 16 ) represents Pixel(1,1) divided by 2 16 The integer part of the quotient of .

[0075] In addition, the next RGB space scrambling is determined by the previous pixel value Pixel value, and an array D = {d i}, where i = (1, 2, 3), and its values ​​are as follows:

[0076]

[0077] Where Lx and Ly represent the horizontal and vertical coordinates of the pixels in the subsequent iteration process as follows:

[0078]

[0079] Wherein i is the current iteration number, i=2,3,......,w*h; mod((i-2),h) represents the remainder of i-2 divided by h, and div((i-2),h) represents the integer part of the quotient of i-2 divided by h.

[0080] According to the above formula, D={d i}, Lx, Ly, and the remaining pixels are iterated by the following formula:

[0081]

[0082] SR(x,y)=mod(Pixel(x,y),2 8 );

[0083] SG(x,y)=mod(div(Pixel(x,y),2 8 ),2 8 );

[0084] SB(x,y)=mod(div(Pixel(x,y),2 16 ),28 );

[0085] Where X[n 1 ]=2,3,...,w*h,x=w,w-1...,1,y=h,h-1,...1,X[n 1 ] represents n 1 The chaotic sequence X after the expansion of the position, SR (x, y) is the R channel component of the image in the RGB color space of the (1, 1) coordinate, SG (x, y) is the G channel component of the image in the RGB color space of the (1, 1) coordinate, SB (x, y) is the B channel component of the image in the RGB color space of the (1, 1) coordinate, Pixel (x, y) is the pixel value after SR (x, y), SG (x, y), SB (x, y) are spliced, represents the XOR operation, d 0 , d 1 , d 2 represents D = {d i} three components.

[0086] Finally, the encryption of the image is completed, and then the key information area image is fused back to the original image, and the detection box coordinate information is added to the image using digital watermark technology to complete the encryption of the key information.

[0087] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well.

Claims

1. A low-resolution real-time critical information detection and protection method, characterized in that: The following steps are involved: (1) Images acquired using a webcam; (2) Input the image obtained by the network camera into the trained target detection model, which extracts and fuses features and outputs categories and detection box areas of different scales; (3) Use chaotic encryption to encrypt the content in the detection frame in real time, and embed the encrypted area coordinate information into the unencrypted area of ​​the image.

2. A low-resolution real-time critical information detection and protection method according to claim 1, characterized in that: The target detection model is a low-resolution real-time target detection model LRRT-Det, which includes a feature extraction backbone network ShuffleNetV2, a generalized feature pyramid GeFusion-FPN network based on attention fusion, and a fully convolutional single-stage FCOS target detection head; The feature extraction backbone network is used to extract multi-scale features of the image. GeFusion-FPN uses a feature selector to fuse features of different scales. The FCOS target detection head is used to finally output the regional coordinates and detection category of the detection box.

3. A low-resolution real-time critical information detection and protection method according to claim 2, characterized in that: The feature extraction backbone network uses ShuffleNetV2 to extract multi-scale features of the image, and inputs three scale feature maps of 1 / 8, 1 / 16, and 1 / 32 into GeFusion-FPN for feature fusion.

4. A low-resolution real-time critical information detection and protection method according to claim 2, characterized in that: Ge-Fusion FPN uses a feature selector to fuse features of different scales. The fusion formula is: C=A×score [:,0] +B×score [:,1] ; A and B are the two features of the input Ge-Fusion FPN, score [:,0] is the weight score corresponding to feature A, score [:,1] is the weight score corresponding to feature B, and C is the fusion feature output by Ge-Fusion FPN.

5. A low-resolution real-time critical information detection and protection method according to claim 1, characterized in that: The training of the object detection model consists of the following steps: S1, use the high-resolution images of the original dataset to train the LRRT-Det model step by step; S2. Data distillation using high-resolution images and the large object detection model Yolov11.

6. A low-resolution real-time critical information detection and protection method according to claim 5, characterized in that: In step S1, the LRRT-Det model is first trained under the scale of 416x416 image input to obtain training parameters, and the obtained training parameters are used to initialize the target detection LRRT-Det model to guide the training of LRRT-Det under the scale of 320x320 image input; Data distillation in step S2: Use the Coco2017 and Widerface datasets to train the Yolov11 model with a 640x640 image input scale, and use the output of the Yolov11 model as a soft label. The soft label includes not only the probability distribution of the target category, but also the relative relationship between each category, thereby providing more learning information for the student model; The teacher model outputs the category probability distribution and detection box coordinates of each target through the Softmax operation. The output z of the teacher model teacher Soft labels are obtained by smoothing with temperature T: Among them, z teacher is the output of the teacher model, T is the temperature coefficient, which is used to control the smoothness of the output; the student model minimizes the loss function L total , to achieve the best detection effect, L total Using the hard labels and soft labels of target detection, the losses of the two are weighted and combined, and the total loss function is expressed as: Among them, λ1, λ2, and λ3 are weighted coefficients, which control the contribution ratio of hard labels, soft labels, and detection boxes respectively. represents the predicted category of the student model, and y represents the true category. is the soft output of the teacher model, is the detection box coordinate output by the student model, b is the actual detection box coordinate, is the hard label classification loss, which is used to measure the difference between the class predicted by the student model and the true label, is the soft label classification loss, which measures the difference between the categories predicted by the student model and the soft labels generated by the teacher model, is the detection box regression loss, which is used to measure the difference between the detection box predicted by the student model and the true detection box.

7. A low-resolution real-time critical information detection and protection method according to claim 1, characterized in that: In step (3), the chaotic sequence iteration equation used is: x k+1 =mod(10·tanh(p1*y k )·exp(x k +y k ),1); y k+1 =mod(10·tanh(x k )·exp(p2·(x k +y k )),1); Among them, exp() represents the natural exponential function, tanh() represents the hyperbolic tangent function, p1 and p2 are control parameters (0,1) set by the user, and y k 、x k is the chaotic value of the current k-th step, and mod(.) represents the modulus operation.

8. A low-resolution real-time critical information detection and protection method according to claim 7, characterized in that: In step (3), the encryption step is as follows: for each detection frame of size w*h, w and h are the number of rows and columns of the detection frame area respectively, firstly associate different keys according to the detection category, then use the chaotic sequence iteration to divide the sequence into a scrambled part and a diffusion part, use the elementary matrix row and column transformation method to scramble the image, use the XOR method to diffuse the pixels, and finally replace the position of the detection frame in the original image with the encrypted image.

9. A low-resolution real-time critical information detection and protection method according to claim 1, characterized in that: In step (3), information embedding is to embed information using a digital watermarking method.

10. A computer device, characterized in that: It comprises: a memory and a processor and a computer program stored in the memory. When the computer program is executed on the processor, a low-resolution real-time critical information detection and protection method as described in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Low-resolution real-time gesture recognition method

    CN115797976A

  • Multi-scale distillation for low-resolution detection

    US20230153943A1