Method, system and equipment for detecting illegal use of mobile phone

By applying the object detection method of deep learning technology in surveillance videos, the situation of illegal use of mobile phones is automatically identified and detected, and the missed discovery problem caused by relying on manual observation in the existing technology is solved, and the reliability of confidential document management is improved.

CN120014557APending Publication Date: 2025-05-16GUANGDONG ELECTRIC POWER SCI RES INST ENERGY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145203.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art determines that illegal use of mobile phones is too dependent on manual observation through surveillance video, which leads to monotonous surveillance images that easily cause fatigue and easy to miss discovery, reducing the reliability of confidential document management.

Method used

The object detection method in deep learning technology is adopted to acquire and preprocess the monitoring image, generate the monitoring feature set, and input the preset initial mobile phone detection model for training to generate the target mobile phone detection model. The model includes a backbone network and a detection network, which is used to extract feature and detect targets of the surveillance image to be identified and monitored, and automatically identify illegal use of mobile phones.

Benefits of technology

It avoids missed discovery caused by artificial fatigue, improves the reliability of confidential document management, and realizes automatic detection and monitoring of illegal use of mobile phones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014557A_ABST
    Figure CN120014557A_ABST
Patent Text Reader

Abstract

The invention discloses a method, a system and equipment for detecting illegal use of a mobile phone, and relates to the technical field of detection of illegal use of the mobile phone. A plurality of training monitoring images are acquired, data preprocessing is performed on all the training monitoring images, a monitoring feature set is generated, and the monitoring feature set is input into a preset initial mobile phone detection model for training; a target mobile phone detection model is generated, the target mobile phone detection model comprises a backbone network and a detection network, when a to-be-recognized monitoring image is received, feature extraction is performed on the to-be-recognized monitoring image through the backbone network, a plurality of monitoring feature images are output in sequence, and target detection is performed on all the monitoring feature images through the detection network; and obtaining a detection result corresponding to the to-be-identified monitoring image. The technical problems that when monitoring personnel observe the monitoring video of a security room and judge whether a person uses a mobile phone illegally or not, manual observation is excessively relied on, the situation of missing discovery is prone to occurring, and the reliability of confidential document management is reduced are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of detecting illegal use of mobile phones, and in particular to a method, system and device for detecting illegal use of mobile phones. Background Art

[0002] For a long time, confidential documents within the power system have been kept in paper form in specific confidential rooms. However, as mobile phones become more and more powerful, they transmit information through various means such as phone calls, voice, video, and photos. Therefore, preventing information leakage caused by illegal use of mobile phones has become a key concern for many units.

[0003] At present, the existing technology mainly relies on monitoring personnel to observe the surveillance video of the confidential room to determine whether someone is using the mobile phone in violation of the regulations. However, this method relies too much on manual observation, and the monotonous monitoring screen can easily make people tired, which can easily lead to missed discoveries, reducing the reliability of confidential document management. Summary of the invention

[0004] The present invention provides a method, system and device for detecting illegal use of mobile phones, which solves the technical problem that the prior art mainly relies on monitoring personnel to observe the monitoring video of the confidential room to determine whether someone is using the mobile phone illegally, but this method is too dependent on manual observation, and the monotonous monitoring screen is easy to make people tired, and it is easy to miss the discovery, which reduces the reliability of confidential document management.

[0005] A first aspect of the present invention provides a method for detecting illegal use of a mobile phone, comprising:

[0006] Acquire multiple training monitoring images, perform data preprocessing on all of the training monitoring images, and generate a monitoring feature set;

[0007] The monitoring feature set is used to input a preset initial mobile phone detection model for training to generate a target mobile phone detection model, wherein the target mobile phone detection model includes a backbone network and a detection network;

[0008] When a surveillance image to be identified is received, feature extraction is performed on the surveillance image to be identified through the backbone network, and a plurality of surveillance feature images are output in sequence;

[0009] The detection network is used to perform target detection on all the monitoring feature images to obtain detection results corresponding to the monitoring images to be identified.

[0010] Optionally, the backbone network includes a first convolution group, a second convolution group and a third convolution group, and when a surveillance image to be identified is received, the step of extracting features of the surveillance image to be identified through the backbone network and sequentially outputting a plurality of surveillance feature images includes:

[0011] When receiving a surveillance image to be identified, extracting features of the surveillance image to be identified by using the first convolution group to generate a first surveillance feature map, wherein the first convolution group includes a first detection module and a second detection module connected in sequence;

[0012] Performing feature extraction on the first monitoring feature map through the second convolution sub-group to generate a second monitoring feature map, wherein the second convolution group includes a second detection module, an attention mechanism module, a second detection module and a third detection module connected in sequence;

[0013] The third convolution group is used to perform feature extraction on the second monitoring feature map to generate a third monitoring feature map, wherein the third convolution group includes a second detection module and a deconvolution layer connected in sequence.

[0014] Optionally, the detection network includes a first fused convolution group, a second fused convolution group and a target detection module, and the step of using the detection network to perform target detection on all the monitoring feature maps to obtain a detection result corresponding to the monitoring image to be identified includes:

[0015] Using the first fused convolution group to perform feature fusion on the first monitoring feature map and the third monitoring feature map to obtain a first fusion map and a fourth monitoring feature map, wherein the first fused convolution group includes a feature fusion layer, a second detection module, an attention mechanism module, a third detection module and a deconvolution layer connected in sequence;

[0016] Performing feature fusion on the second monitoring feature map, the third monitoring feature map, the first fusion map, and the fourth monitoring feature map through the second fusion convolution group to obtain a target feature map;

[0017] The target detection module performs target detection on the target feature map to obtain a detection result corresponding to the surveillance image to be identified.

[0018] Optionally, the step of using the first fused convolution group to perform feature fusion on the first monitoring feature map and the third monitoring feature map to obtain a first fused map and a fourth monitoring feature map includes:

[0019] Performing feature fusion on the first monitoring feature map and the third monitoring feature map through a feature fusion layer to obtain a first fusion map;

[0020] Using a second detection module to extract features from the first fusion image to obtain a first fusion feature image;

[0021] Using an attention mechanism module to perform global adaptive pooling and linear transformation operations on the first fused feature map to obtain a second fused feature map;

[0022] Performing feature extraction on the second fused feature map by a third detection module to obtain a third fused feature map;

[0023] An upsampling operation is performed on the third fused feature map through a deconvolution layer to obtain a fourth monitoring feature map.

[0024] Optionally, the first detection module includes a second detection module and a feature fusion layer connected in sequence; the specific processing process of the first detection module is:

[0025] Using a second detection module to extract features from the input first feature map to generate a second feature map;

[0026] The first feature map and the second feature map are subjected to feature fusion through a feature fusion layer to obtain a third feature map.

[0027] Optionally, the second detection module includes a 3×3 standard convolution layer, a batch normalization layer and an activation layer connected in sequence; the specific processing process of the second detection module is:

[0028] Extract features from the fourth feature map of the input through a 3×3 standard convolutional layer to generate a fifth feature map;

[0029] Performing normalization processing on the fifth feature map through a batch normalization layer to obtain a sixth feature map;

[0030] The sixth feature map is nonlinearly transformed by using an activation layer to obtain a seventh feature map.

[0031] Optionally, the second fused convolution group includes two extraction branches, a fusion branch, a feature fusion layer and a first detection module, and the step of performing feature fusion on the second monitoring feature map, the third monitoring feature map, the first fusion map and the fourth monitoring feature map through the second fused convolution group to obtain a target feature map includes:

[0032] Performing feature extraction on the fourth monitoring feature map through an extraction branch to obtain a fifth monitoring feature map, wherein the extraction branch includes a feature fusion layer and a second detection module connected in sequence;

[0033] Performing feature fusion on the fourth monitoring feature map, the second monitoring feature map and the third monitoring feature map through a fusion branch to obtain a fourth fused feature map, wherein the fusion branch includes two feature fusion layers and a second detection module;

[0034] Using an extraction branch to perform feature extraction on the third monitoring feature graph to obtain a sixth monitoring feature graph;

[0035] Using a feature fusion layer to perform feature fusion on the fifth monitoring feature map, the fourth fused feature map and the sixth monitoring feature map to obtain a fifth fused feature map;

[0036] The first detection module performs feature extraction on the fifth fused feature map to obtain a target feature map.

[0037] Optionally, the step of using the monitoring feature set to input a preset initial mobile phone detection model for training to generate a target mobile phone detection model includes:

[0038] Using the monitoring feature set to input a preset initial mobile phone detection model for training, and outputting training detection data;

[0039] Calculate the training loss function value of the monitoring feature set according to the training detection data;

[0040] When the training loss function value is greater than or equal to the preset standard loss value, the network parameters of the preset initial mobile phone detection model are adjusted using the gradient descent method until the training loss function value is less than the preset standard loss value;

[0041] When the training loss function value is less than the preset standard loss value, a target mobile phone detection model is generated.

[0042] A second aspect of the present invention provides a detection system for illegal use of a mobile phone, comprising:

[0043] A preprocessing module, used to obtain a plurality of training monitoring images, perform data preprocessing on all the training monitoring images, and generate a monitoring feature set;

[0044] A training module, used to use the monitoring feature set to input a preset initial mobile phone detection model for training to generate a target mobile phone detection model, wherein the target mobile phone detection model includes a backbone network and a detection network;

[0045] An extraction module, configured to extract features of a surveillance image to be identified through the backbone network when receiving the surveillance image to be identified, and output a plurality of surveillance feature images in sequence;

[0046] The detection module is used to use the detection network to perform target detection on all the monitoring feature images to obtain the detection results corresponding to the monitoring images to be identified.

[0047] A third aspect of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for detecting illegal use of a mobile phone as described in any one of the above items.

[0048] It can be seen from the above technical solutions that the present invention has the following advantages:

[0049] Acquire multiple training monitoring images, perform data preprocessing on all training monitoring images, generate a monitoring feature set, use the monitoring feature set to input a preset initial mobile phone detection model for training, and generate a target mobile phone detection model, wherein the target mobile phone detection model includes a backbone network and a detection network. When receiving a monitoring image to be identified, extract features of the monitoring image to be identified through the backbone network, output multiple monitoring feature maps in sequence, use the detection network to perform target detection on all monitoring feature maps, and obtain the detection result corresponding to the monitoring image to be identified. This solves the technical problem that the monitoring personnel rely too much on manual observation to judge whether someone is using a mobile phone illegally when observing the monitoring video of the confidential room, and the monotonous monitoring screen is easy to make people tired, which is easy to miss the discovery and reduces the reliability of confidential document management. This application integrates the target detection in deep learning technology into the traditional monitoring video, avoids the missed discovery caused by manual fatigue, and improves the reliability of confidential document management. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0051] Figure 1 A flowchart of a method for detecting illegal use of a mobile phone provided in Embodiment 1 of the present invention;

[0052] Figure 2 A flowchart of a method for detecting illegal use of a mobile phone provided in Embodiment 2 of the present invention;

[0053] Figure 3 A schematic diagram of the structure of a target mobile phone detection model provided in Embodiment 2 of the present invention;

[0054] Figure 4 A schematic diagram of the structure of the attention mechanism module provided in the second embodiment of the present invention;

[0055] Figure 5 A structural block diagram of a detection system for illegal use of mobile phones provided in Embodiment 3 of the present invention;

[0056] Figure 6 This is a structural block diagram of a computer device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION

[0057] The embodiments of the present invention provide a method, system and device for detecting illegal use of mobile phones, which are used to solve the technical problem that the existing technology mainly relies on monitoring personnel to observe the monitoring video of the confidential room to determine whether someone is using the mobile phone illegally, but this method is too dependent on manual observation, and the monotonous monitoring screen is easy to make people tired, and it is easy to miss the discovery, which reduces the reliability of confidential document management.

[0058] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0059] See also Figure 1 , Figure 1 This is a flowchart of a method for detecting illegal use of a mobile phone provided in Embodiment 1 of the present invention.

[0060] The present invention provides a method for detecting illegal use of a mobile phone, comprising:

[0061] Step 101: Acquire multiple training monitoring images, perform data preprocessing on all the training monitoring images, and generate a monitoring feature set;

[0062] In the embodiment of the present invention, a plurality of training monitoring images are obtained, data preprocessing is performed on each training monitoring image in turn, and the training monitoring images after data preprocessing are manually labeled to obtain a monitoring feature set.

[0063] It should be noted that the data preprocessing process includes flipping, rotating, scaling, cropping, brightness, contrast adjustment, random transformation, color distortion and noise addition operations on the training monitoring images. Among them, 1. Flipping: Flip the image horizontally or vertically to generate new training samples. This is particularly useful for training image classification models because objects usually still have the same category label in the mirror. 2. Rotation: Rotate the image to produce variants at different angles. This may be beneficial for dealing with some tasks with rotation invariance, such as object detection or image segmentation. 3. Scaling and Cropping: Adaptively scale or crop the image to simulate inputs at different scales and perspectives. This helps improve the model's ability to recognize targets at different locations in the image. 4. Brightness and Contrast Adjustment: Adjust the brightness and contrast of the image to make the model more robust to inputs under different lighting conditions. 5. Random Transformations: Introduce some random transformations during training, such as random translation, random rotation, etc., to increase the diversity of samples. 6. Color Distortion: Perform color changes on the image, such as randomly adjusting brightness, contrast, saturation, etc., to improve the model's ability to adapt to color changes. 7. Adding Noise: Add random noise to the image to make the model have a certain tolerance for noise interference.

[0064] Step 102: Using the monitoring feature set to input a preset initial mobile phone detection model for training, generating a target mobile phone detection model, wherein the target mobile phone detection model includes a backbone network and a detection network;

[0065] In an embodiment of the present invention, a monitoring feature set is input into a preset initial mobile phone detection model for training, and training detection data is output. A training loss function value of the monitoring feature set is calculated based on the training detection data. When the training loss function value is greater than or equal to a preset standard loss value, a gradient descent method is used to adjust the network parameters of the preset initial mobile phone detection model, and the step of inputting the monitoring feature set into the preset initial mobile phone detection model for training and outputting the training detection data is jumped to execution until the training loss function value is less than the preset standard loss value. When the training loss function value is less than the preset standard loss value, a target mobile phone detection model is generated, wherein the target mobile phone detection model includes a backbone network and a detection network.

[0066] Step 103: when receiving a surveillance image to be identified, extracting features of the surveillance image to be identified through the backbone network, and outputting a plurality of surveillance feature images in sequence;

[0067] The surveillance image to be identified refers to the surveillance image of the current frame in the surveillance video.

[0068] In an embodiment of the present invention, when a surveillance image to be identified is received, a backbone network is used to extract features of the surveillance image to be identified, and multiple surveillance feature images are output in sequence, wherein the backbone network includes a first convolution group, a second convolution group and a third convolution group.

[0069] Step 104: Use the detection network to perform target detection on all monitoring feature images to obtain detection results corresponding to the monitoring images to be identified.

[0070] In an embodiment of the present invention, target detection is performed on all monitoring feature images through a detection network to obtain detection results corresponding to the monitoring images to be identified.

[0071] In an embodiment of the present invention, a plurality of training monitoring images are obtained, data preprocessing is performed on all the training monitoring images, a monitoring feature set is generated, and the monitoring feature set is used to input a preset initial mobile phone detection model for training to generate a target mobile phone detection model, wherein the target mobile phone detection model includes a backbone network and a detection network. When a monitoring image to be identified is received, feature extraction is performed on the monitoring image to be identified through the backbone network, and a plurality of monitoring feature maps are output in sequence. The detection network is used to perform target detection on all monitoring feature maps to obtain the detection result corresponding to the monitoring image to be identified. The technical problem that the monitoring personnel rely too much on manual observation to judge whether someone is using a mobile phone in violation of the regulations when observing the monitoring video of the confidential room, and the monotonous monitoring screen is easy to make people tired, and it is easy to miss the discovery, which reduces the reliability of confidential document management. The present application integrates the target detection in deep learning technology into the traditional monitoring video, avoids the missed discovery caused by manual fatigue, and improves the reliability of confidential document management.

[0072] See also Figure 2 , Figure 2 This is a flow chart of the steps of a method for detecting illegal use of a mobile phone provided in Embodiment 2 of the present invention.

[0073] The present invention provides a method for detecting illegal use of a mobile phone, comprising:

[0074] Step 201: Acquire multiple training monitoring images, perform data preprocessing on all training monitoring images, and generate a monitoring feature set;

[0075] In an embodiment of the present invention, a plurality of training monitoring images are obtained, data preprocessing (flipping, rotating, scaling, cropping, brightness and contrast adjustment, random transformation, color distortion, and noise addition, etc.) is performed on each training monitoring image, and each training monitoring image after data preprocessing is manually labeled to obtain a monitoring feature set.

[0076] Step 202: Use the monitoring feature set to input a preset initial mobile phone detection model for training, and output training detection data;

[0077] In the embodiment of the present invention, the monitoring feature set is input into a preset initial mobile phone detection model for training, and the training detection data is output.

[0078] Step 203: Calculate the training loss function value of the monitoring feature set according to the training detection data;

[0079] In the embodiment of the present invention, based on the intersection-over-union loss function, the training loss function value of the monitoring feature set is calculated according to the training detection data.

[0080] It should be noted that the overlap between the training detection data (prediction box) and the monitoring data (real box) in the monitoring feature set is calculated through the intersection-over-union loss function, which effectively shields the interference of the bounding box size in the form of a ratio. However, there are two problems: 1. If the two boxes do not intersect, according to the definition, IoU=0, which cannot reflect the distance between the two (overlapping degree). At the same time, because IoU=0, there is no gradient backpropagation, and learning and training cannot be performed. 2. When the intersection-over-union ratio of the predicted box and the real box is the same, but the training detection data is located in different locations, because the calculated loss is the same, it is impossible to judge which prediction is more accurate. In order to ensure the accuracy of the model prediction, the ratio of the center point distance to the diagonal distance is added to avoid the generation of a larger outer box when the two boxes are far apart, and the Loss value is large and difficult to optimize; at the same time, the addition of length and width loss directly minimizes the difference in height and width between the predicted target bounding box and the real bounding box, so that it has a faster convergence speed and better positioning results.

[0081] It should be noted that the intersection-over-union loss function is specifically:

[0082]

[0083]

[0084] in, is the degree of overlap, is the prediction box, is the real frame, is the training loss function value, is the distance loss, is the length and width loss, is the coordinate of the center point of the real frame, is the coordinate of the center point of the prediction box, is the Euclidean distance between the center point coordinates, is the prediction box width, is the actual frame width, is the Euclidean distance between widths, is the predicted box height, is the real frame height, is the Euclidean distance between the heights.

[0085] Step 204: when the training loss function value is greater than or equal to the preset standard loss value, the network parameters of the preset initial mobile phone detection model are adjusted using the gradient descent method until the training loss function value is less than the preset standard loss value;

[0086] In an embodiment of the present invention, when the training loss function value is greater than or equal to the preset standard loss value, the gradient descent method is used to adjust the network parameters of the preset initial mobile phone detection model, and the step of using the monitoring feature set to input the preset initial mobile phone detection model for training and outputting the training detection data is jumped to execution until the training loss function value is less than the preset standard loss value.

[0087] Step 205: When the training loss function value is less than the preset standard loss value, a target mobile phone detection model is generated, wherein the target mobile phone detection model includes a backbone network and a detection network;

[0088] In an embodiment of the present invention, when the training loss function value is less than a preset standard loss value, a target mobile phone detection model is generated, wherein the target mobile phone detection model includes a backbone network and a detection network.

[0089] Step 206: When a surveillance image to be identified is received, feature extraction is performed on the surveillance image to be identified through the backbone network, and a plurality of surveillance feature images are output in sequence;

[0090] Further, the backbone network includes a first convolution group, a second convolution group and a third convolution group, and step 206 includes the following sub-steps:

[0091] S11. When receiving a surveillance image to be identified, extracting features of the surveillance image to be identified by using a first convolution group to generate a first surveillance feature map, wherein the first convolution group includes a first detection module and a second detection module connected in sequence;

[0092] In the embodiment of the present invention, refer to Figure 3 As shown, when a monitoring image to be identified is received, features of the monitoring image to be identified are extracted through the first convolution group to generate a first monitoring feature map, wherein the first convolution group includes a first detection module and a second detection module connected in sequence.

[0093] It should be noted that the first detection module includes a second detection module and a feature fusion layer connected in sequence; the specific processing process of the first detection module is:

[0094] A1. Use the second detection module to extract features from the input first feature map to generate a second feature map;

[0095] A2. The first feature map and the second feature map are fused through a feature fusion layer to obtain a third feature map.

[0096] In the embodiment of the present invention, refer to Figure 3 As shown, the second detection module performs feature extraction on the input first feature map to generate a second feature map, and then the feature fusion layer performs feature fusion on the first feature map and the second feature map to obtain a third feature map.

[0097] It should be noted that the second detection module includes a 3×3 standard convolution layer, a batch normalization layer, and an activation layer connected in sequence; the specific processing process of the second detection module is:

[0098] B1. Extract features from the fourth feature map input through a 3×3 standard convolutional layer to generate a fifth feature map;

[0099] In the embodiment of the present invention, a 3×3 standard convolutional layer is used to perform feature extraction on the input fourth feature map to generate a fifth feature map.

[0100] B2. Performing normalization processing on the fifth feature map through a batch normalization layer to obtain a sixth feature map;

[0101] In the embodiment of the present invention, the fifth feature map is normalized by a batch normalization layer to obtain a sixth feature map.

[0102] It is worth mentioning that after processing through the batch normalization layer, gradient disappearance and overfitting can be prevented.

[0103] B3. Use the activation layer to perform nonlinear transformation on the sixth feature map to obtain the seventh feature map.

[0104] It should be noted that the activation layer may be a SILU activation layer.

[0105] In the embodiment of the present invention, the sixth feature map is nonlinearly transformed through the SILU activation layer to obtain the seventh feature map.

[0106] S12, extracting features from the first monitoring feature map through a second convolution sub-group to generate a second monitoring feature map, wherein the second convolution group includes a second detection module, an attention mechanism module, a second detection module, and a third detection module connected in sequence;

[0107] S13. Perform feature extraction on the second monitoring feature map through a third convolution group to generate a third monitoring feature map, wherein the third convolution group includes a second detection module and a deconvolution layer connected in sequence.

[0108] In the embodiment of the present invention, the second monitoring feature map is feature extracted by the second detection module, and then the deconvolution layer (Upsample) is used to perform an upsampling operation on the second monitoring feature map after feature extraction to obtain a third monitoring feature map.

[0109] Step 207: Use the detection network to perform target detection on all monitoring feature images to obtain detection results corresponding to the monitoring images to be identified.

[0110] Further, the detection network includes a first fused convolution group, a second fused convolution group and a target detection module, and step 207 includes the following sub-steps:

[0111] S21, using a first fusion convolution group to perform feature fusion on the first monitoring feature map and the third monitoring feature map to obtain a first fusion map and a fourth monitoring feature map, wherein the first fusion convolution group includes a feature fusion layer, a second detection module, an attention mechanism module, a third detection module and a deconvolution layer connected in sequence;

[0112] Furthermore, S21 includes the following sub-steps:

[0113] S211, performing feature fusion on the first monitoring feature map and the third monitoring feature map through a feature fusion layer to obtain a first fusion map;

[0114] In the embodiment of the present invention, the first monitoring feature map and the third monitoring feature map are fused by a feature fusion layer (Concat) to obtain a first fusion map.

[0115] S212, using a second detection module to extract features from the first fusion image to obtain a first fusion feature image;

[0116] In the embodiment of the present invention, the second detection module performs feature extraction on the first fusion map to obtain a first fusion feature map.

[0117] S213, using an attention mechanism module to perform global adaptive pooling and linear transformation operations on the first fused feature map to obtain a second fused feature map;

[0118] In the embodiment of the present invention, refer to Figure 4 As shown, the first fusion feature map is subjected to global adaptive pooling, linear transformation and channel weighting operations in sequence through the intention mechanism module to obtain the second fusion feature map.

[0119] It should be noted that the specific processing process of the attention mechanism module is:

[0120] 1. The input feature map is subjected to global average pooling, and the feature map is converted from a matrix of [h,w,c] to a vector of [1,1,c].

[0121] 2. Calculate the adaptive one-dimensional convolution kernel size kernel_size according to the number of channels of the feature map.

[0122] 3. Use kernel_size for one-dimensional convolution to get the weight for each channel of the feature map.

[0123] 4. Multiply the normalized weights and the original input feature map channel by channel to generate a weighted feature map.

[0124] It is worth mentioning that the attention mechanism module can effectively capture global feature information, which is very useful for processing channel information in tasks such as small target detection such as mobile phones. It can help the network better focus on channels that contribute to the task and reduce unnecessary computational burden. The target mobile phone detection model embeds the attention mechanism module to enhance the network's perception of features at different layers.

[0125] S214, extracting features from the second fused feature map through a third detection module to obtain a third fused feature map;

[0126] In an embodiment of the present invention, a third detection module (SPP) is used to perform feature extraction on the second fused feature map to obtain a third fused feature map, wherein the third detection module includes a 1×1 maximum pooling layer, a 5×5 maximum pooling layer, a 9×9 maximum pooling layer, a 13×13 maximum pooling layer and a feature fusion layer.

[0127] It should be noted that the specific processing process of the third detection module is: use 1×1 maximum pooling layer, 5×5 maximum pooling layer, 9×9 maximum pooling layer, and 13×13 maximum pooling layer to perform pooling operations on the input feature map, respectively, to obtain 4 pooled feature maps, and then use the feature fusion layer to perform feature fusion on all the pooled feature maps and the input feature map to obtain the feature map after feature extraction.

[0128] S215. Perform an upsampling operation on the third fused feature map through a deconvolution layer to obtain a fourth monitoring feature map.

[0129] In the embodiment of the present invention, a deconvolution layer is used to perform an upsampling operation on the third fused feature map to obtain a fourth monitoring feature map.

[0130] S22, performing feature fusion on the second monitoring feature map, the third monitoring feature map, the first fusion map and the fourth monitoring feature map through a second fusion convolution group to obtain a target feature map;

[0131] Further, the second fused convolution group includes two extraction branches, a fusion branch, a feature fusion layer and a first detection module, and S22 includes the following sub-steps:

[0132] S221, extracting features from the fourth monitoring feature map through an extraction branch to obtain a fifth monitoring feature map, wherein the extraction branch includes a feature fusion layer and a second detection module connected in sequence;

[0133] In the embodiment of the present invention, the fourth monitoring feature map is connected to the second detection module through a feature fusion layer, and the second detection module is used to extract features from the fourth monitoring feature map to obtain a fifth monitoring feature map.

[0134] S222, performing feature fusion on the fourth monitoring feature map, the second monitoring feature map and the third monitoring feature map through a fusion branch to obtain a fourth fused feature map, wherein the fusion branch includes two feature fusion layers and a second detection module;

[0135] In an embodiment of the present invention, the fourth monitoring feature map and the second monitoring feature map are feature fused through a feature fusion layer to obtain an intermediate fused feature map, the third monitoring feature map is connected to the second detection module through the feature fusion layer, and features are extracted from the intermediate fused feature map and the third monitoring feature map through the second detection module to obtain a fourth fused feature map.

[0136] S223, using the extraction branch to perform feature extraction on the third monitoring feature graph to obtain a sixth monitoring feature graph;

[0137] In the embodiment of the present invention, the third monitoring feature map is connected to the second detection module through a feature fusion layer, and the second detection module is used to extract features from the third monitoring feature map to obtain a sixth monitoring feature map.

[0138] S224, using a feature fusion layer to perform feature fusion on the fifth monitoring feature map, the fourth fused feature map, and the sixth monitoring feature map to obtain a fifth fused feature map;

[0139] In the embodiment of the present invention, the fifth monitoring feature map, the fourth fused feature map and the sixth monitoring feature map are feature fused through a feature fusion layer to obtain a fifth fused feature map.

[0140] S225. Perform feature extraction on the fifth fusion feature map through the first detection module to obtain a target feature map.

[0141] In the embodiment of the present invention, refer to Figure 3 As shown, the fifth fused feature map is extracted through the first detection module (Focus) to obtain a target feature map.

[0142] S23. Perform target detection on the target feature map through the target detection module to obtain the detection result corresponding to the monitoring image to be identified.

[0143] In the embodiment of the present invention, target detection is performed on the target feature map through a target detection module (SoftMax function) to obtain a detection result corresponding to the monitoring image to be identified.

[0144] In an embodiment of the present invention, a plurality of training monitoring images are obtained, data preprocessing is performed on all the training monitoring images, a monitoring feature set is generated, and the monitoring feature set is used to input a preset initial mobile phone detection model for training to generate a target mobile phone detection model, wherein the target mobile phone detection model includes a backbone network and a detection network. When a monitoring image to be identified is received, feature extraction is performed on the monitoring image to be identified through the backbone network, and a plurality of monitoring feature maps are output in sequence. The detection network is used to perform target detection on all monitoring feature maps to obtain the detection result corresponding to the monitoring image to be identified. The technical problem that the monitoring personnel rely too much on manual observation to judge whether someone is using a mobile phone in violation of the regulations when observing the monitoring video of the confidential room, and the monotonous monitoring screen is easy to make people tired, and it is easy to miss the discovery, which reduces the reliability of confidential document management. The present application integrates the target detection in deep learning technology into the traditional monitoring video, avoids the missed discovery caused by manual fatigue, and improves the reliability of confidential document management.

[0145] See also Figure 5 , Figure 5 This is a structural block diagram of a system for detecting illegal use of mobile phones provided in Embodiment 3 of the present invention.

[0146] The present invention provides a detection system for illegal use of mobile phones, comprising:

[0147] The preprocessing module 301 is used to obtain multiple training monitoring images, perform data preprocessing on all the training monitoring images, and generate a monitoring feature set;

[0148] A training module 302 is used to use the monitoring feature set to input a preset initial mobile phone detection model for training to generate a target mobile phone detection model, wherein the target mobile phone detection model includes a backbone network and a detection network;

[0149] The extraction module 303 is used for extracting features of the surveillance image to be identified through the backbone network when receiving the surveillance image to be identified, and outputting a plurality of surveillance feature images in sequence;

[0150] The detection module 304 is used to use the detection network to perform target detection on all monitoring feature images to obtain the detection results corresponding to the monitoring images to be identified.

[0151] Furthermore, the training module 302 includes:

[0152] A training submodule, used to use the monitoring feature set to input a preset initial mobile phone detection model for training, and output training detection data;

[0153] A first analysis submodule is used to calculate the training loss function value of the monitoring feature set based on the training detection data;

[0154] When the training loss function value is greater than or equal to the preset standard loss value, the gradient descent method is used to adjust the network parameters of the preset initial mobile phone detection model until the training loss function value is less than the preset standard loss value;

[0155] When the training loss function value is less than the preset standard loss value, the target mobile phone detection model is generated.

[0156] Further, the backbone network includes a first convolution group, a second convolution group and a third convolution group, and the extraction module 303 includes:

[0157] A first extraction submodule is used for, when receiving a surveillance image to be identified, extracting features of the surveillance image to be identified by using a first convolution group to generate a first surveillance feature map, wherein the first convolution group includes a first detection module and a second detection module connected in sequence;

[0158] A second extraction submodule is used to extract features from the first monitoring feature map through a second convolution submodule to generate a second monitoring feature map, wherein the second convolution group includes a second detection module, an attention mechanism module, a second detection module and a third detection module connected in sequence;

[0159] The third extraction submodule is used to extract features from the second monitoring feature map through a third convolution group to generate a third monitoring feature map, wherein the third convolution group includes a second detection module and a deconvolution layer connected in sequence.

[0160] Furthermore, the detection network includes a first fused convolution group, a second fused convolution group and a target detection module, and the detection module 304 includes:

[0161] A first fusion submodule, used for performing feature fusion on the first monitoring feature map and the third monitoring feature map by using a first fusion convolution group to obtain a first fusion map and a fourth monitoring feature map, wherein the first fusion convolution group includes a feature fusion layer, a second detection module, an attention mechanism module, a third detection module and a deconvolution layer connected in sequence;

[0162] A second fusion submodule is used to perform feature fusion on the second monitoring feature map, the third monitoring feature map, the first fusion map and the fourth monitoring feature map through a second fusion convolution group to obtain a target feature map;

[0163] The detection submodule is used to perform target detection on the target feature map through the target detection module to obtain the detection result corresponding to the monitoring image to be identified.

[0164] Furthermore, the first fusion submodule includes:

[0165] A first fusion unit is used to perform feature fusion on the first monitoring feature map and the third monitoring feature map through a feature fusion layer to obtain a first fusion map;

[0166] A first extraction unit, configured to extract features from the first fusion image using a second detection module to obtain a first fusion feature image;

[0167] A second extraction unit is used to perform global adaptive pooling and linear transformation operations on the first fused feature map by using an attention mechanism module to obtain a second fused feature map;

[0168] A third extraction unit, configured to extract features from the second fused feature map through a third detection module to obtain a third fused feature map;

[0169] The fourth extraction unit is used to perform an upsampling operation on the third fused feature map through a deconvolution layer to obtain a fourth monitoring feature map.

[0170] Furthermore, the second fusion convolution group includes two extraction branches, a fusion branch, a feature fusion layer and a first detection module, and the second fusion submodule includes:

[0171] A fifth extraction unit, configured to extract features from the fourth monitoring feature map through an extraction branch to obtain a fifth monitoring feature map, wherein the extraction branch includes a feature fusion layer and a second detection module connected in sequence;

[0172] A second fusion unit is used to perform feature fusion on the fourth monitoring feature map, the second monitoring feature map and the third monitoring feature map through a fusion branch to obtain a fourth fused feature map, wherein the fusion branch includes two feature fusion layers and a second detection module;

[0173] a sixth extraction unit, configured to extract features from the third monitoring feature graph using an extraction branch to obtain a sixth monitoring feature graph;

[0174] A third fusion unit is used to use a feature fusion layer to perform feature fusion on the fifth monitoring feature map, the fourth fusion feature map and the sixth monitoring feature map to obtain a fifth fusion feature map;

[0175] The seventh extraction unit is used to perform feature extraction on the fifth fusion feature map through the first detection module to obtain a target feature map.

[0176] Furthermore, the first detection module includes a second detection module and a feature fusion layer connected in sequence; the specific processing process of the first detection module is:

[0177] Using a second detection module to extract features from the input first feature map to generate a second feature map;

[0178] The first feature map and the second feature map are fused through the feature fusion layer to obtain a third feature map.

[0179] Furthermore, the second detection module includes a 3×3 standard convolution layer, a batch normalization layer, and an activation layer connected in sequence; the specific processing process of the second detection module is:

[0180] Extract features from the fourth feature map of the input through a 3×3 standard convolutional layer to generate a fifth feature map;

[0181] The fifth feature map is normalized by a batch normalization layer to obtain a sixth feature map;

[0182] The activation layer is used to perform nonlinear transformation on the sixth feature map to obtain the seventh feature map.

[0183] See also Figure 6 , Figure 6 This is a structural block diagram of a computer device provided in Embodiment 4 of the present invention.

[0184] An electronic device according to an embodiment of the present invention includes: a memory 401 and a processor 402, wherein the memory 402 stores a computer program; when the computer program is executed by the processor 402, the processor 402 executes a method for detecting illegal use of a mobile phone as described in any of the above embodiments.

[0185] The memory 401 may be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk or a ROM. The memory 401 has a storage space 403 for a program code 413 for executing any method step in the above method. For example, the storage space 403 for the program code may include individual program codes 413 for implementing the various steps in the above method, respectively. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards or floppy disks. The program code may be compressed, for example, in an appropriate form. When these codes are run by a computing and processing device, the computing and processing device performs the various steps in the method described above. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards or floppy disks. The program code may be compressed, for example, in an appropriate form. When these codes are executed by a computing and processing device, the computing and processing device is caused to execute each step of the above-described method for detecting illegal use of a mobile phone.

[0186] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0187] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0188] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0189] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0190] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0191] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting illegal use of a mobile phone, characterized in that: include: Acquire multiple training monitoring images, perform data preprocessing on all of the training monitoring images, and generate a monitoring feature set; The monitoring feature set is used to input a preset initial mobile phone detection model for training to generate a target mobile phone detection model, wherein the target mobile phone detection model includes a backbone network and a detection network; When a surveillance image to be identified is received, feature extraction is performed on the surveillance image to be identified through the backbone network, and a plurality of surveillance feature images are output in sequence; The detection network is used to perform target detection on all the monitoring feature images to obtain detection results corresponding to the monitoring images to be identified.

2. The method for detecting illegal use of mobile phones according to claim 1, characterized in that: The backbone network includes a first convolution group, a second convolution group and a third convolution group. When a surveillance image to be identified is received, the backbone network is used to extract features of the surveillance image to be identified, and a plurality of surveillance feature images are sequentially outputted, comprising: When receiving a surveillance image to be identified, extracting features of the surveillance image to be identified by using the first convolution group to generate a first surveillance feature map, wherein the first convolution group includes a first detection module and a second detection module connected in sequence; Performing feature extraction on the first monitoring feature map through the second convolution sub-group to generate a second monitoring feature map, wherein the second convolution group includes a second detection module, an attention mechanism module, a second detection module and a third detection module connected in sequence; The third convolution group is used to perform feature extraction on the second monitoring feature map to generate a third monitoring feature map, wherein the third convolution group includes a second detection module and a deconvolution layer connected in sequence.

3. The method for detecting illegal use of mobile phones according to claim 2, characterized in that: The detection network includes a first fused convolution group, a second fused convolution group and a target detection module. The step of using the detection network to perform target detection on all the monitoring feature graphs to obtain the detection result corresponding to the monitoring image to be identified includes: Using the first fused convolution group to perform feature fusion on the first monitoring feature map and the third monitoring feature map to obtain a first fusion map and a fourth monitoring feature map, wherein the first fused convolution group includes a feature fusion layer, a second detection module, an attention mechanism module, a third detection module and a deconvolution layer connected in sequence; Performing feature fusion on the second monitoring feature map, the third monitoring feature map, the first fusion map, and the fourth monitoring feature map through the second fusion convolution group to obtain a target feature map; The target detection module performs target detection on the target feature map to obtain a detection result corresponding to the surveillance image to be identified.

4. The method for detecting illegal use of mobile phones according to claim 3, characterized in that: The step of using the first fused convolution group to perform feature fusion on the first monitoring feature map and the third monitoring feature map to obtain a first fused map and a fourth monitoring feature map includes: Performing feature fusion on the first monitoring feature map and the third monitoring feature map through a feature fusion layer to obtain a first fusion map; Using a second detection module to extract features from the first fusion image to obtain a first fusion feature image; Using an attention mechanism module to perform global adaptive pooling and linear transformation operations on the first fused feature map to obtain a second fused feature map; Performing feature extraction on the second fused feature map by a third detection module to obtain a third fused feature map; An upsampling operation is performed on the third fused feature map through a deconvolution layer to obtain a fourth monitoring feature map.

5. The method for detecting illegal use of mobile phones according to claim 2, characterized in that: The first detection module includes a second detection module and a feature fusion layer connected in sequence; the specific processing process of the first detection module is: Using a second detection module to extract features from the input first feature map to generate a second feature map; The first feature map and the second feature map are subjected to feature fusion through a feature fusion layer to obtain a third feature map.

6. The method for detecting illegal use of a mobile phone according to any one of claims 2 to 5, characterized in that: The second detection module includes a 3×3 standard convolution layer, a batch normalization layer and an activation layer connected in sequence; the specific processing process of the second detection module is: Extract features from the fourth feature map of the input through a 3×3 standard convolutional layer to generate a fifth feature map; Performing normalization processing on the fifth feature map through a batch normalization layer to obtain a sixth feature map; The sixth feature map is nonlinearly transformed by using an activation layer to obtain a seventh feature map.

7. The method for detecting illegal use of mobile phones according to claim 3, characterized in that: The second fused convolution group includes two extraction branches, a fusion branch, a feature fusion layer and a first detection module. The step of performing feature fusion on the second monitoring feature map, the third monitoring feature map, the first fusion map and the fourth monitoring feature map through the second fused convolution group to obtain a target feature map includes: Performing feature extraction on the fourth monitoring feature map through an extraction branch to obtain a fifth monitoring feature map, wherein the extraction branch includes a feature fusion layer and a second detection module connected in sequence; Performing feature fusion on the fourth monitoring feature map, the second monitoring feature map and the third monitoring feature map through a fusion branch to obtain a fourth fused feature map, wherein the fusion branch includes two feature fusion layers and a second detection module; Using an extraction branch to perform feature extraction on the third monitoring feature graph to obtain a sixth monitoring feature graph; Using a feature fusion layer to perform feature fusion on the fifth monitoring feature map, the fourth fused feature map and the sixth monitoring feature map to obtain a fifth fused feature map; The first detection module performs feature extraction on the fifth fused feature map to obtain a target feature map.

8. The method for detecting illegal use of mobile phones according to claim 1, characterized in that: The step of using the monitoring feature set to input a preset initial mobile phone detection model for training to generate a target mobile phone detection model includes: Using the monitoring feature set to input a preset initial mobile phone detection model for training, and outputting training detection data; Calculate the training loss function value of the monitoring feature set according to the training detection data; When the training loss function value is greater than or equal to the preset standard loss value, the network parameters of the preset initial mobile phone detection model are adjusted using the gradient descent method until the training loss function value is less than the preset standard loss value; When the training loss function value is less than the preset standard loss value, a target mobile phone detection model is generated.

9. A detection system for illegal use of mobile phones, characterized in that: include: A preprocessing module, used to obtain a plurality of training monitoring images, perform data preprocessing on all the training monitoring images, and generate a monitoring feature set; A training module, used to use the monitoring feature set to input a preset initial mobile phone detection model for training to generate a target mobile phone detection model, wherein the target mobile phone detection model includes a backbone network and a detection network; An extraction module, configured to extract features of a surveillance image to be identified through the backbone network when receiving the surveillance image to be identified, and output a plurality of surveillance feature images in sequence; The detection module is used to use the detection network to perform target detection on all the monitoring feature images to obtain the detection results corresponding to the monitoring images to be identified.

10. An electronic device, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for detecting illegal use of a mobile phone as described in any one of claims 1-8.