Intelligent privacy protection camera system applied to online car-hailing
By applying an intelligent privacy protection camera system in online car-hailing, using SwinMIX-Det object detection algorithm, improving Gaussian fuzzy algorithm and improving ArcFace algorithm, the problems of insufficient privacy protection, difficulty in detecting small targets and low facial recognition accuracy in the existing technology are solved, and more accurate privacy protection and dispute evidence acquisition are achieved.
Patent Information
- Application Number
- CN202510512148.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing online car-hailing and on-board camera devices have shortcomings in privacy protection, and cannot accurately identify the privacy areas of passengers and drivers, it is difficult to detect small targets, and the accuracy of facial recognition is low, which has security loopholes and inconveniences.
An intelligent privacy protection camera system was designed, using SwinMIX-Det object detection algorithm and improved Gaussian fuzzy algorithm. Combined with driver terminal, passenger terminal and camera device, it realizes accurate identification and blurring of the custom privacy areas of drivers and passengers, and improves facial recognition accuracy by improving ArcFace algorithm.
It realizes accurate identification and protection of driver and passenger privacy areas, improves the accuracy of small object detection, enhances the accuracy of facial recognition, improves the blur effect, and ensures the comprehensiveness of privacy protection and the reliability of dispute evidence.
Smart Images

Figure CN120050398A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to camera privacy protection, and particularly to an intelligent privacy protection camera system applied to online car-hailing services. Background Art
[0002] With the rapid development of Internet technology, online car-hailing has become one of the important choices for people's daily travel. While providing convenient travel services for people, the safety and privacy issues in its operation have become increasingly prominent. To ensure operation safety, online car-hailing vehicles usually install camera devices in the car. However, existing in-vehicle camera devices have many deficiencies in privacy protection. On the one hand, traditional camera devices cannot accurately identify the privacy areas of passengers and drivers. During actual rides, passengers may place personal items such as wallets, mobile phones, and documents on their seats, which carry a large amount of personal privacy information; drivers may also place some items such as personal documents and work materials around the driver's seat that they do not want to be randomly photographed. However, existing camera devices lack an effective privacy area recognition mechanism, resulting in these privacy information being easily exposed.
[0003] On the other hand, when it comes to small target detection, such as small personal items carried by passengers like earphones and keychains, existing camera devices are difficult to accurately identify. This not only affects the comprehensiveness of privacy protection but may also lead to the lack of evidence due to the inability to clearly identify relevant items during the handling of disputes, and thus cannot effectively safeguard the rights and interests of both drivers and passengers. In addition, in the process of handling disputes, the existing process of viewing surveillance footage has security vulnerabilities and inconveniences. For drivers to view surveillance footage, there is a lack of a reliable identity verification mechanism, which is prone to abuse of authority and infringement of passengers' privacy; while passengers often face cumbersome processes and difficulties in obtaining effective surveillance evidence when safeguarding their own rights and interests, resulting in the difficulty of fully protecting passengers' rights and interests. Summary of the Invention
[0004] In view of the above-mentioned drawbacks of the prior art, the present invention provides an intelligent privacy protection camera system applied to online car-hailing, which can effectively overcome the defects of the prior art, such as the inability to achieve effective privacy protection, the difficulty in accurately detecting small targets, and the low accuracy of face recognition.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: An intelligent privacy protection camera system applied to online car-hailing, comprising: A driver terminal, which allows the driver to preset the driver's privacy area in the surveillance footage collected by the camera device, upload it to the system terminal, and simultaneously perform face recognition on the driver; A passenger terminal for passengers to mark the passenger privacy areas they wish to protect in the in-vehicle virtual scene model and upload them to the system side, while authenticating the passengers' identities; A camera device that obtains feature data related to detecting passengers entering the online car-hailing vehicle, starts the camera when the feature data meets the boot condition, and uploads the captured surveillance video to the system side; A system side that, based on the driver privacy area and the passenger privacy area, uses the SwinMIX-Det object detection algorithm to detect the target privacy area in the surveillance video, and uses an improved Gaussian blur algorithm to blur the target privacy area. At the same time, according to the face recognition result and the identity authentication result, the corresponding original surveillance video is retrieved and sent to the driver terminal and the passenger terminal respectively; Among them, the SwinMIX-Det object detection algorithm uses Swin Transformer as the backbone network, introduces the SE module and the SOEMCM module dedicated to small object detection, and optimizes and improves the BiFPN layer, adding top-down, bottom-up, cross, and skip connections, fusing deep image processing features and RGB image processing features, and stacking multiple levels of BiFPN layers. The feature fusion structure of each level of BiFPN layer is different.
[0006] Preferably, the driver terminal allows the driver to preset the driver privacy area in the surveillance video captured by the camera device and upload it to the system side, including: The driver presets the driver privacy area in the surveillance video captured by the camera device using touch and box selection interaction methods, and the driver terminal uploads the driver privacy area to the system side; The passenger terminal allows passengers to mark the passenger privacy areas they wish to protect in the in-vehicle virtual scene model and upload them to the system side, including: Before taking an online car-hailing vehicle, the passenger marks the passenger privacy areas they wish to protect in the in-vehicle virtual scene model provided by the online car-hailing platform application, and the passenger terminal uploads the passenger privacy area to the system side; Among them, the passenger privacy area includes a specific area on the seat where personal items are placed, and an area within a certain radius around the seat to prevent personal behavior from being overly photographed.
[0007] Preferably, the camera device obtains feature data related to detecting passengers entering the online car-hailing vehicle, starts the camera when the feature data meets the boot condition, and uploads the captured surveillance video to the system side, including: Obtaining feature data related to detecting passengers entering the online car-hailing vehicle, including a spatial structure topology map representing the spatial structure layout of the vehicle interior M and a seat occupancy identification vector representing whether each seat in the vehicle is occupied S iand the illumination feature matrix that characterizes the color temperature and intensity information of the ambient light at the door opening position L i ; When a door opening action is detected in the spatial structure topology map M and at the same time the seat occupancy identification vector S i changes within a short period of time, and the changes in the part of the illumination feature matrix L i that characterize the color temperature and intensity of the ambient light exceed a preset threshold, the camera device starts shooting and uploads the captured monitoring screen to the system terminal.
[0008] Preferably, based on the driver privacy area and the passenger privacy area, the system terminal uses the SwinMIX-Det object detection algorithm to detect the target privacy area in the monitoring screen, including: The network structure corresponding to the SwinMIX-Det object detection algorithm includes: An image fusion module that fuses the input RGB image and depth image to obtain a fused image, and normalizes the fused image and then inputs it into the Swin Transformer backbone network to enrich the information source for feature extraction and help the network understand the target spatial position and structural features; The Swin Transformer backbone network contains 5 levels of Swin Transformer Block modules. The input of each level of Swin Transformer Block module is the downsampling result of the output feature map of the previous level of Swin Transformer Block module, that is, the length and width are halved and the number of features is doubled. Among them, the feature maps output by the first two levels of Swin Transformer Block modules F 1 and F 2 are input into the SE module, and the feature maps output by the last three levels of Swin Transformer Block modules F 3 、 F 4 and F 5 are input into the SOEMCM module; The SE module performs global average pooling on the feature maps F 1 and F 2 to obtain a j -dimensional vector z j where j is the number of channels of the feature map, and then the vectorz j Pass through two fully connected layers in sequence to obtain channel attention weights , and finally multiply the channel attention weights by the original input feature map channel by channel to obtain the output feature map of the SE module; The SOEMCM module performs convolution, max pooling (MaxPool), and average pooling (AvgPool) on the feature maps F 3 , F 4 , F 5 respectively to obtain the output feature map of the SOEMCM module; The BiFPN layer adjusts the number of channels of the output feature maps of the SE module and the SOEMCM module through convolution to achieve dimension adaptation, and obtains input feature maps of 5 dimensions through channel dimension concatenation. Then, the three-level BiFPN layer with different feature fusion structures is used to connect the feature points in the input feature maps in different ways to obtain the output feature map of the BiFPN layer; The output layer inputs the output feature map of the BiFPN layer into the class prediction network and the boundary prediction network respectively. The class prediction network and the boundary prediction network both process the output feature map through convolution to obtain the class prediction result and the boundary prediction result respectively, completing the detection task of the target privacy area in the input image.
[0009] Preferably, in the Swin Transformer backbone network, the feature map transmission process is represented by the following formula: ; In the above formula, F N-1 , F N are the input feature map and the output feature map of the N th Swin Transformer Block module respectively, N = 1, 2, …, 5, F 0 is the initial feature map extracted by preliminary convolution of the normalized fused image; The processing process of the SE module for the feature maps F 1 and F 2 is represented by the following formula: ; In the above formula, F j ( i , k ) is the feature mapF j The feature points in i , k ), j where \(i = 1, 2\), H and W are the height and width of the feature map F j respectively; W 1 and W 2 are the weight matrices of two fully connected layers, ReLU is the rectified linear unit function, and the sigmoid function maps the output value to the interval \((0, 1)\) as the channel attention weight , F SEj is the output feature map of the SE module; The processing process of the SOEMCM module for the feature maps F 3 , F 4 and F 5 is expressed by the following formula: ; In the above formula, K d is the convolution kernel, , , are the processing results corresponding to the feature maps F 3 , F 4 , F 5 respectively, that is, the output feature maps of the SOEMCM module; In the BiFPN layer, the first - level BiFPN layer performs top - down, bottom - up, and skip connections on the feature points in the input feature map, the second - level BiFPN layer performs cross and top - down connections on the feature points in the input feature map, and the third - level BiFPN layer performs reverse top - down, reverse bottom - up, and skip connections on the feature points in the input feature map. The output feature maps corresponding to various connection methods are expressed by the following formula: ; In the above formula, , are the high - level feature map and low - level feature map of the l th level respectively, Upsample is upsampling, is the output feature map corresponding to the top - down connection; K ( m , n) is the convolution kernel K The weight at m , n ), the size of the convolution kernel K is s * s , is the feature point ([[]] l in the low-level feature map of the is+m , ks+n ), is the output feature map corresponding to the bottom-up connection; , are respectively the feature maps of the l th level from the paths a , b , is the output feature map corresponding to the cross connection; , , are respectively the feature maps of the l , l -1, l -2 levels from the path c , is the output feature map corresponding to the skip connection; In the said output layer, the process of the category prediction network and the boundary prediction network processing the output feature map of the BiFPN layer through convolution is represented by the following formula: ; In the above formula, K e , K f are respectively the category prediction convolution kernel and the boundary prediction convolution kernel, is the output feature map of the BiFPN layer, F class , F bbax are respectively the category prediction feature map and the boundary prediction feature map.
[0010] Preferably, the system end uses an improved Gaussian blur algorithm to blur the target privacy area, including: The network structure corresponding to the improved Gaussian blur algorithm includes: Optimize the convolution calculation module, convert the original image containing the target privacy area and the Gaussian kernel to the frequency domain by accelerating convolution through FFT, multiply them in the frequency domain and then convert back to the spatial domain, or directly obtain the corresponding convolution result from the lookup table through convolution based on the lookup table, and finally obtain the initial Gaussian blurred image; The multi-scale analysis and fusion module generates multiple secondary Gaussian blurred images using Gaussian kernels of different scales based on the initial Gaussian blurred image, calculates the local edge intensity of the initial Gaussian blurred image, determines the fusion weights of the secondary Gaussian blurred images according to the local edge intensity of the initial Gaussian blurred image, and adaptively fuses the multiple secondary Gaussian blurred images to obtain a fused Gaussian blurred image; The content adaptive module based on the image adaptively adjusts the Gaussian blur parameters by calculating the local features of the fused Gaussian blurred image, thereby obtaining the final Gaussian blurred image.
[0011] Preferably, the optimized convolution calculation module converts the original image containing the target privacy area and the Gaussian kernel to the frequency domain by accelerating convolution with FFT, multiplies them in the frequency domain, and then converts back to the spatial domain, which is represented by the following formula: ; In the above formula, I ( x , y ) is the original image, G ( x , y ) is the Gaussian kernel, x and y are pixel coordinates, F represents the Fourier transform, F -1 represents the inverse Fourier transform, I b ( x , y ) is the initial Gaussian blurred image, is the standard deviation of the Gaussian distribution that controls the degree of blur; The optimized convolution calculation module directly obtains the corresponding convolution result from the lookup table through convolution based on the lookup table, including: Pre-calculate and store the Gaussian kernels and the corresponding convolution results under different standard deviations of the Gaussian distribution. When processing the original image, according to the input local features and the required standard deviation of the Gaussian distribution, directly obtain the corresponding convolution result from the lookup table to obtain the initial Gaussian blurred image I b ( x , y ); The process by which the multi-scale analysis and fusion module obtains the fused Gaussian blurred image includes: S1. Generate multiple secondary Gaussian blurred images using Gaussian kernels of different scales based on the initial Gaussian blurred image: ; In the above formula, Gp ( x , y ) is the p th Gaussian kernel, I p ( x , y ) is the p th quadratic Gaussian blurred image, p = 1, 2, …, P , P is the number of Gaussian kernels at different scales; S2. Calculate the local edge intensity of the initial Gaussian blurred image: ; In the above formula, G X ( x , y ), G Y ( x , y ) are the gradients of the initial Gaussian blurred image I b ( x , y ) along the x , y directions at the point ([[]] X , Y ), and E ( x , y ) is the local edge intensity of the initial Gaussian blurred image I b ( x , y ); S3. Determine the edge region based on the local edge intensity I b ( x , y ) of the initial Gaussian blurred image. For the edge region, assign a larger fusion weight to the quadratic Gaussian blurred image at a small scale to retain edge details; for the non-edge region, assign a larger fusion weight to the quadratic Gaussian blurred image at a large scale to achieve a better smoothing effect; E ( x , y ) of the initial Gaussian blurred image. S4. Perform adaptive fusion on multiple quadratic Gaussian blurred images to obtain a fused Gaussian blurred image: ; In the above formula, w p ( x , y ) is thep a second Gaussian blurred image I p ( x , y ) in ( x , y ) the fusion weight at the position, , I f ( x , y ) is the fused Gaussian blurred image; The image-based content adaptive module adaptively adjusts the Gaussian blur parameters by calculating the local features of the fused Gaussian blurred image, and is represented by the following formula: ; In the above formula, (2 M + 1)(2 N + 1) is the Gaussian kernel window size, I f ( x - u , y - v ) is the fused Gaussian blurred image I f ( x , y ) in ( x - u , y - v ) the pixel value at the position, is the mean pixel value of the local region in the fused Gaussian blurred image I f ( x , y ) in ( x , y ) at the position, V ( x , y ) is the local variance in the fused Gaussian blurred image I f ( x , y ) in ( x , y ) at the position, p r is the probability of pixels with a gray level of r in the local region, L is the gray level, H ( x , y ) is the fused Gaussian blurred image I f ( x , y ) in (x , y the entropy at A 、 B 、 C are all preset parameters, is the standard deviation of the Gaussian distribution corresponding to the pixel at ( x , y ); Among them, by adaptively adjusting the Gaussian blur parameters, for regions with rich textures and many details, that is, V ( x , y ) and H ( x , y ) with larger values, a smaller blur degree is adopted, that is, is smaller; for flat regions with more noise, that is, V ( x , y ) and H ( x , y ) with smaller values, a larger blur degree is adopted, that is, is larger.
[0012] Preferably, when a dispute occurs or a passenger safeguards their own rights and interests, the driver performs face recognition through the driver terminal. The driver terminal uses an improved ArcFace algorithm to perform face recognition on the driver. If the face recognition is successful, the system terminal retrieves the corresponding original surveillance video and sends it to the driver terminal; When a dispute occurs or a passenger safeguards their own rights and interests, the passenger performs identity verification through the passenger terminal. The passenger terminal uses the reserved mobile phone number to perform identity verification on the passenger. If the identity verification is successful, the system terminal retrieves the corresponding original surveillance video and sends it to the passenger terminal.
[0013] Preferably, in terms of feature extraction, the improved ArcFace algorithm introduces an adaptive multi-scale feature fusion mechanism: Let the input original face image be I , and through a multi-branch convolutional neural network structure, a set of feature maps at different scales is obtained. Each feature map contains information within a specific spatial frequency range, i = 1, 2, …, n ; Then, an adaptive weight generation module g is designed to generate an adaptive weight for each feature map according to the image content, that is, ; The fused features are represented as , and after being processed by subsequent network layers, feature vectors for recognition are obtained , and normalization is performed ; In the angular constraint part of the improved ArcFace algorithm, angular boundaries are added to the traditional ArcFace algorithm m . On this basis, a dynamic adjustment factor related to the sample difficulty is introduced k : Let the weight vector corresponding to the sample with class label y be w y , and normalization is also performed , the feature vector and the weight vector w y The included angle between them is , that is ; The sample difficulty is measured by calculating the average distance between the features of this sample and the features of samples in the same class, as well as the average distance between the features of samples in other classes, and a dynamic adjustment factor is generated k : For simple samples, the dynamic adjustment factor k is smaller, and the increase in the angular boundary is smaller; for difficult samples, the dynamic adjustment factor k is larger, and the increase in the angular boundary is larger to enhance the discrimination ability of the model for difficult samples; The loss function of the improved ArcFace algorithm is: ; In the above formula, is the feature vector corresponding to the fused features of the i th sample, is the weight vector corresponding to the sample with class label y i , is the feature vector and the weight vector The included angle between them, k i is the i th sample's dynamic adjustment factor k , is the weight vector corresponding to the sample with class label j , is the feature vector and the weight vector The included angle between them, N is the number of training samples.
[0014] Compared with the prior art, the intelligent privacy protection camera system applied to online car-hailing provided by the present invention has the following beneficial effects: 1) Accurately identify privacy areas. For the privacy areas customized by drivers and passengers, it can accurately identify and blur them, effectively avoiding the exposure of privacy information and protecting the privacy of drivers and passengers; 2) Improve the small target detection ability. The SwinMIX-Det target detection algorithm proposed by the present invention integrates multiple convolutional modules, which can effectively improve the accuracy of small object detection and improve privacy protection and the acquisition of dispute evidence; 3) Improve the blurring effect. The improved Gaussian blurring algorithm proposed by the present invention can perform Gaussian blurring according to the local features of the image by optimizing convolution calculation, multi-scale analysis and fusion, and adaptive adjustment of Gaussian blurring parameters on the image, achieving a better blurring effect; 4) Improve the accuracy of face recognition. The improved ArcFace algorithm proposed by the present invention introduces adaptive multi-scale feature fusion and an angle dynamic adjustment factor, strengthens the ability to distinguish similar images and complex situations, and realizes accurate face recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1 It is a schematic diagram of the system operation process of the present invention; Figure 2 It is a schematic diagram of the network structure corresponding to the SwinMIX-Det target detection algorithm in the present invention; Figure 3 It is a schematic diagram of the process of blurring the target privacy area by using the improved Gaussian blurring algorithm in the present invention; Figure 4 It is an example diagram of image processing using the SwinMIX-Det target detection algorithm and the improved Gaussian blurring algorithm proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0018] An intelligent privacy protection camera system applied to online car-hailing, such as Figure 1 shown, a driver terminal for the driver to preset the driver privacy area in the monitoring screen collected by the camera device, upload it to the system terminal, and perform face recognition on the driver at the same time; a passenger terminal for the passenger to mark the passenger privacy area to be protected in the in-vehicle virtual scene model, upload it to the system terminal, and perform identity verification on the passenger at the same time; a camera device for obtaining feature data related to detecting the entry of a passenger into an online car-hailing, starting the camera when the feature data meets the power-on condition, and uploading the collected monitoring screen to the system terminal; a system terminal for detecting the target privacy area in the monitoring screen using the SwinMIX-Det target detection algorithm based on the driver privacy area and the passenger privacy area, and performing blur processing on the target privacy area using an improved Gaussian blur algorithm. At the same time, according to the face recognition result and the identity verification result, the corresponding original monitoring screen is retrieved and sent to the driver terminal and the passenger terminal respectively; Among them, the SwinMIX-Det target detection algorithm uses Swin Transformer as the backbone network, introduces the SE module and the SOEMCM module focusing on small target detection, optimizes and improves the BiFPN layer, adds top-down, bottom-up, cross, and skip connections, fuses the depth image processing features and the RGB image processing features, and stacks multiple levels of BiFPN layers. The feature fusion structures of each level of BiFPN layer are different.
[0019] The driver terminal is for the driver to preset the driver privacy area in the monitoring screen collected by the camera device and upload it to the system terminal, including: The driver presets the driver privacy area in the monitoring screen collected by the camera device by using touch and box selection interaction methods, and the driver terminal uploads the driver privacy area to the system terminal; The passenger terminal is for the passenger to mark the passenger privacy area to be protected in the in-vehicle virtual scene model and upload it to the system terminal, including: Before taking an online car-hailing, the passenger marks the passenger privacy area to be protected in the in-vehicle virtual scene model provided by the online car-hailing platform application, and the passenger terminal uploads the passenger privacy area to the system terminal; Among them, the passenger privacy area includes a specific area on the seat for placing personal items and an area within a certain radius around the seat to prevent excessive filming of personal behavior.
[0020] The camera device acquires feature data related to detecting the entry of a passenger into the online car-hailing vehicle. When the feature data meets the startup condition, the camera is started, and the captured monitoring video is uploaded to the system side, including: Acquiring feature data related to detecting the entry of a passenger into the online car-hailing vehicle, including a spatial structure topology map representing the layout of the interior space structure M and a seat occupancy identification vector representing whether each seat in the vehicle is occupied S i and a lighting feature matrix representing the ambient light color temperature and intensity information at the door opening position L i ; When a door opening action is detected in the spatial structure topology map M , and at the same time the seat occupancy identification vector S i changes within a short period of time, and the change in the part representing the ambient light color temperature and intensity in the lighting feature matrix L i exceeds a preset threshold, the camera device starts filming and uploads the captured monitoring video to the system side.
[0021] Based on the driver privacy area and the passenger privacy area, the system side uses the SwinMIX-Det object detection algorithm to detect the target privacy area in the monitoring video, including: As Figure 2 shown, the network structure corresponding to the SwinMIX-Det object detection algorithm includes: An image fusion module that fuses the input RGB image and depth image to obtain a fused image, and normalizes the fused image before inputting it into the Swin Transformer backbone network to enrich the information source for feature extraction and help the network understand the target spatial position and structural features; The Swin Transformer backbone network contains 5 levels of Swin Transformer Block modules. The input of each level of Swin Transformer Block module is the downsampling result of the output feature map of the previous level of Swin Transformer Block module, that is, the length and width are halved and the number of features is doubled. The feature maps output by the first two levels of Swin Transformer Block modules F 1 and F 2Input to the SE module, and the feature maps output by the last three Swin Transformer Block modules F 3 、 F 4 and F 5 are input to the SOEMCM module; The SE module performs global average pooling on the feature maps F 1 and F 2 to obtain a j -dimensional vector z j , where j is the number of channels of the feature map. Then, the vector z j is successively passed through two fully connected layers to obtain the channel attention weights. Finally, the channel attention weights are multiplied by the original input feature map channel by channel to obtain the output feature map of the SE module; The SOEMCM module performs convolution, max pooling (MaxPool), and average pooling (AvgPool) on the feature maps F 3 、 F 4 、 F 5 respectively to obtain the output feature map of the SOEMCM module; The BiFPN layer adjusts the number of channels of the output feature maps of the SE module and the SOEMCM module through convolution to achieve dimensional adaptation, and obtains the input feature maps of 5 dimensions through channel dimension concatenation. Then, the feature points in the input feature maps are connected in different ways using three BiFPN layers with different feature fusion structures to obtain the output feature map of the BiFPN layer; The output layer inputs the output feature map of the BiFPN layer into the class prediction network and the bounding box prediction network respectively. The class prediction network and the bounding box prediction network both process the output feature map through convolution to obtain the class prediction result and the bounding box prediction result respectively, completing the detection task of the target privacy area in the input image.
[0022] 1) In the Swin Transformer backbone network, the feature map transmission process is represented by the following formula: ; In the above formula, F N-1 、 F N are the input feature map and the output feature map of the N -th Swin Transformer Block module respectively,N = 1, 2, …, 5, F 0 is the initial feature map extracted by the preliminary convolution of the normalized fused image; 2) The processing process of the SE module for the feature map F 1 and F 2 is represented by the following formula: ; In the above formula, F j ( i , k ) is the feature point ( F j ) in the feature map i , k ), j = 1, 2, H 、 W are the height and width of the feature map F j respectively, W 1 、 W 2 are the weight matrices of two fully connected layers, ReLU is the rectified linear unit function, and the sigmoid function maps the output value to the interval (0, 1) as the channel attention weight , F SEj is the output feature map of the SE module; 3) The processing process of the SOEMCM module for the feature maps F 3 、 F 4 and F 5 is represented by the following formula: ; In the above formula, K d is the convolution kernel, 、 、 are the corresponding processing results of the feature maps F 3 、 F 4 、 F 5 respectively, that is, the output feature maps of the SOEMCM module; 4) In the BiFPN layer, the first - level BiFPN layer performs top - down, bottom - up, and skip connections on the feature points in the input feature map. The second - level BiFPN layer performs cross and top - down connections on the feature points in the input feature map. The third - level BiFPN layer performs reverse top - down, reverse bottom - up, and skip connections on the feature points in the input feature map. The output feature maps corresponding to various connection methods are represented by the following formula: ; In the above formula, and are the high - level feature map and low - level feature map of the l th level respectively. Upsample is upsampling, is the output feature map corresponding to the top - down connection; K ( m , n ) is the weight at the position of( K in the convolutional kernel m , n ). The size of the convolutional kernel K is s * s . is the feature point( l in the low - level feature map of the is+m , ks+n )th level. is the output feature map corresponding to the bottom - up connection; and are the feature maps of the l th level from the paths a and b respectively. is the output feature map corresponding to the cross connection; and and are the feature maps of the l , l - 1, l - 2th levels from the path c respectively. is the output feature map corresponding to the skip connection; 5) In the output layer, the process of the class prediction network and the bounding box prediction network both processing the output feature map of the BiFPN layer through convolution is represented by the following formula: ; In the above formula, K e and K f are the class prediction convolutional kernel and the bounding box prediction convolutional kernel respectively. is the output feature map of the BiFPN layer, F class and F bbax are the class prediction feature map and the bounding box prediction feature map respectively.
[0023] The system side uses an improved Gaussian blur algorithm to blur the target privacy area, including: As Figure 3 shown, the network structure corresponding to the improved Gaussian blur algorithm includes: An optimized convolution calculation module that converts the original image containing the target privacy area and the Gaussian kernel to the frequency domain through FFT-accelerated convolution, multiplies them in the frequency domain, and then converts back to the spatial domain, or directly obtains the corresponding convolution result from a lookup table through convolution based on the lookup table, and finally obtains the initial Gaussian blurred image; A multi-scale analysis and fusion module that generates multiple secondary Gaussian blurred images using Gaussian kernels of different scales based on the initial Gaussian blurred image, calculates the local edge intensity of the initial Gaussian blurred image, determines the fusion weights of each secondary Gaussian blurred image according to the local edge intensity of the initial Gaussian blurred image, and adaptively fuses the multiple secondary Gaussian blurred images to obtain a fused Gaussian blurred image; An image-based content adaptive module that adaptively adjusts the Gaussian blur parameters by calculating the local features of the fused Gaussian blurred image, thereby obtaining the final Gaussian blurred image.
[0024] 1) The optimized convolution calculation module converts the original image containing the target privacy area and the Gaussian kernel to the frequency domain through FFT-accelerated convolution, multiplies them in the frequency domain, and then converts back to the spatial domain, which is represented by the following formula: ; In the above formula, I ( x , y ) is the original image, G ( x , y ) is the Gaussian kernel, x and y are pixel coordinates, F represents the Fourier transform, F -1 represents the inverse Fourier transform, I b ( x , y ) is the initial Gaussian blurred image, is the standard deviation of the Gaussian distribution that controls the degree of blur; The optimized convolution calculation module directly obtains the corresponding convolution result from the lookup table through convolution based on the lookup table, including: Pre-calculate and store different standard deviations of the Gaussian distribution The Gaussian kernel below and the corresponding convolution result. When processing the original image, according to the input local features and the standard deviation of the required Gaussian distribution , directly obtain the corresponding convolution result from the look-up table to obtain the initial Gaussian blurred image I b ( x , y ); 2) The process of the multi-scale analysis and fusion module to obtain the fused Gaussian blurred image includes: S1. Generate multiple secondary Gaussian blurred images using Gaussian kernels of different scales based on the initial Gaussian blurred image: ; In the above formula, G p ( x , y ) is the p th Gaussian kernel, I p ( x , y ) is the p th secondary Gaussian blurred image, p = 1, 2, …, P , P is the number of Gaussian kernels of different scales; S2. Calculate the local edge intensity of the initial Gaussian blurred image: ; In the above formula, G X ( x , y ), G Y ( x , y ) are the gradients of the initial Gaussian blurred image I b ( x , y ) at ([[]] x , y ) along the X , Y directions, E ( x , y ) is the local edge intensity of the initial Gaussian blurred image I b ( x , y ); S3. According to the initial Gaussian blurred image I b ( x ,y Local edge strength of E ( x , y ) Determine the edge region. For the edge region, assign a larger fusion weight to the small-scale second Gaussian blurred image to retain edge details; for the non-edge region, assign a larger fusion weight to the large-scale second Gaussian blurred image to achieve a better smoothing effect; S4. Perform adaptive fusion on multiple second Gaussian blurred images to obtain a fused Gaussian blurred image: ; In the above formula, w p ( x , y ) is the p rd second Gaussian blurred image I p ( x , y ) at ([[]] x , y ) is the fusion weight, , I f ( x , y ) is the fused Gaussian blurred image; 3) The content adaptive module based on the image adaptively adjusts the Gaussian blur parameters by calculating the local features of the fused Gaussian blurred image, which is represented by the following formula: ; In the above formula, (2 M + 1)(2 N + 1) is the Gaussian kernel window size, I f ( x - u , y - v ) is the pixel value of the fused Gaussian blurred image I f ( x , y ) at ([[]] x - u , y - v ), is the pixel mean of the local region of the fused Gaussian blurred image I f ( x , y ) at ([[]] x , y ), V ( x ,y ) is for fusing Gaussian blurred images I f ( x , y ) at the local variance at ( x , y ), where the local variance is p r the probability of pixels with gray value r in the local area, L is the gray level, H ( x , y ) is for fusing Gaussian blurred images I f ( x , y ) at the entropy at ( x , y ), A , B , C are all preset parameters, is the standard deviation of the Gaussian distribution corresponding to the pixel at ( x , y ); Among them, by adaptively adjusting the Gaussian blur parameters, for areas with rich textures and many details, that is, V ( x , y ) and H ( x , y ) with larger values, a smaller blur degree is adopted, that is, is smaller; for flat areas with more noise, that is, V ( x , y ) and H ( x , y ) with smaller values, a larger blur degree is adopted, that is, is larger.
[0025] As Figure 1 shown, when a dispute occurs or a passenger defends their own rights and interests, the driver performs face recognition through the driver terminal. The driver terminal uses the improved ArcFace algorithm to perform face recognition on the driver. If the face recognition is successful, the system terminal retrieves the corresponding original surveillance video and sends it to the driver terminal; When a dispute occurs or a passenger defends their own rights and interests, the passenger performs identity verification through the passenger terminal. The passenger terminal uses the reserved mobile phone number to perform identity verification on the passenger. If the identity verification is successful, the system terminal retrieves the corresponding original surveillance video and sends it to the passenger terminal.
[0026] 1) In terms of feature extraction of the improved ArcFace algorithm, an adaptive multi-scale feature fusion mechanism is introduced: Let the input original face image be I , and through a multi-branch convolutional neural network structure, a set of feature maps at different scales is obtained. Each feature map contains information in a specific spatial frequency range, i = 1, 2, …, n ; Then, an adaptive weight generation module g is designed to generate an adaptive weight for each feature map according to the image content, that is, ; The fused feature is expressed as . After being processed by subsequent network layers, a feature vector for recognition is obtained and normalized ; 2) In the angle constraint part of the improved ArcFace algorithm, on the basis of adding an angle boundary m in the traditional ArcFace algorithm, a dynamic adjustment factor k related to the sample difficulty is introduced: Let the weight vector corresponding to the sample with class label y be w y , which is also normalized . The angle between the feature vector and the weight vector w y is , that is, ; The sample difficulty is measured by calculating the average distance between the features of this sample and those of the same-class samples, as well as the average distance between the features of this sample and those of other-class samples, and a dynamic adjustment factor k is generated: For easy samples, the dynamic adjustment factor k is small, and the increase in the angle boundary is small; for difficult samples, the dynamic adjustment factor k is large, and the increase in the angle boundary is large, so as to enhance the discrimination ability of the model for difficult samples; 3) The loss function of the improved ArcFace algorithm is: ; In the above formula, is the feature vector corresponding to the fused feature of the i -th sample, is the class label y i The weight vector corresponding to the sample of is the feature vector and the weight vector The included angle between k i is the i Dynamic adjustment factor of the individual sample k , is the class label j The weight vector corresponding to the sample of is the feature vector and the weight vector The included angle between N is the number of training samples
[0027] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements will not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention
Claims
1. An intelligent privacy protection camera system for online car-hailing, characterized by: include: The driver terminal allows the driver to pre-set the driver's privacy area in the monitoring image captured by the camera device and upload it to the system end, while performing facial recognition on the driver; Passenger terminal, which allows passengers to mark the privacy areas they wish to protect in the virtual scene model inside the car, upload the information to the system, and authenticate the passengers at the same time; The camera device obtains characteristic data related to detecting passengers entering the online car-hailing vehicle, starts the camera when the characteristic data meets the startup conditions, and uploads the collected monitoring images to the system end; On the system side, based on the driver privacy area and the passenger privacy area, the SwinMIX-Det target detection algorithm is used to detect the target privacy area in the monitoring image, and the improved Gaussian blur algorithm is used to blur the target privacy area. At the same time, the corresponding original monitoring image is retrieved according to the face recognition result and the identity verification result, and sent to the driver terminal and the passenger terminal respectively; Among them, the SwinMIX-Det target detection algorithm uses Swin Transformer as the backbone network, introduces the SE module and the SOEMCM module focusing on small target detection, and optimizes and improves the BiFPN layer, adding top-down, bottom-up, cross and jump connections, integrating deep image processing features and RGB image processing features, and superimposing multiple levels of BiFPN layers. The feature fusion structure of each level of BiFPN layer is different.
2. The intelligent privacy protection camera system for online car-hailing according to claim 1 is characterized in that: The driver terminal allows the driver to pre-set the driver's privacy area in the monitoring picture collected by the camera device and upload it to the system end, including: The driver pre-sets the driver's privacy area in the monitoring image captured by the camera device by using touch and frame selection interaction, and the driver terminal uploads the driver's privacy area to the system end; The passenger terminal allows passengers to mark the privacy areas they wish to protect in the virtual scene model inside the vehicle and upload the marks to the system end, including: Before taking an online ride-hailing car, the passenger marks the privacy area that he or she wishes to protect in the virtual scene model of the car provided by the online ride-hailing platform application, and the passenger terminal uploads the privacy area to the system end; Among them, the passenger privacy area includes a specific area on the seat where personal belongings are placed, and an area within a certain radius around the seat to prevent personal behavior from being excessively photographed.
3. The intelligent privacy protection camera system for online car-hailing according to claim 1 is characterized in that: The camera device acquires characteristic data associated with detecting a passenger entering an online car-hailing vehicle, starts the camera when the characteristic data meets the startup conditions, and uploads the collected monitoring images to the system end, including: Obtain feature data associated with detecting passengers entering online ride-hailing vehicles, including a spatial structure topology diagram that characterizes the spatial structure layout inside the vehicle M , a seat occupancy identification vector indicating whether each seat in the car is occupied S i and the illumination feature matrix representing the ambient light color temperature and intensity information at the door opening position L i ; When in the spatial structure topology M The door opening action is detected in the S i Changes in a short time, and the light feature matrix L i When the change of the part representing the color temperature and intensity of the ambient light exceeds the preset threshold, the camera device starts recording and uploads the collected monitoring images to the system end.
4. The intelligent privacy protection camera system for online car-hailing according to claim 1 is characterized in that: The system uses the SwinMIX-Det target detection algorithm to detect the target privacy area in the monitoring screen based on the driver privacy area and the passenger privacy area, including: The network structure corresponding to the SwinMIX-Det target detection algorithm includes: Image fusion module: It fuses the input RGB image and depth image to obtain a fused image, normalizes the fused image and inputs it into the Swin Transformer backbone network to enrich the information source for feature extraction and help the network understand the target spatial position and structural characteristics; The Swin Transformer backbone network consists of five levels of Swin Transformer Block modules. The input of each level of Swin Transformer Block module is the downsampling result of the output feature map of the previous level of Swin Transformer Block module, that is, the length and width are halved and the feature quantity is doubled. The feature maps output by the first two levels of Swin Transformer Block modules are F 1 and F 2 Input to SE module, feature map output by the last three levels of Swin Transformer Block modules F 3. F 4 and F 5 Input to SOEMCM module; SE module, for feature maps F 1 and F 2 Perform global average pooling and get j Dimensional vector z j ,in j is the number of feature map channels, and then the vector z j Pass through two fully connected layers in sequence to obtain the channel attention weight , and finally the channel attention weight Multiply the original input feature map channel by channel to obtain the output feature map of the SE module; SOEMCM module, feature map F 3. F 4. F 5. Perform convolution, maximum pooling MaxPool, and average pooling AvgPool respectively to obtain the output feature map of the SOEMCM module; The BiFPN layer adjusts the number of channels of the output feature maps of the SE module and the SOEMCM module through convolution to achieve dimensional adaptation, and obtains a 5-dimensional input feature map through channel dimension splicing. Then, the three-level BiFPN layer with different feature fusion structures is used to connect the feature points in the input feature map in different ways to obtain the output feature map of the BiFPN layer; At the output layer, the output feature map of the BiFPN layer is input into the category prediction network and the boundary prediction network respectively. Both the category prediction network and the boundary prediction network process the output feature map through convolution to obtain the category prediction results and the boundary prediction results respectively, thus completing the detection task of the target privacy area in the input image.
5. The intelligent privacy protection camera system for online car-hailing according to claim 4 is characterized in that: In the Swin Transformer backbone network, the feature map transmission process is expressed as follows: ; In the above formula, F N-1 , F N Respectively N The input feature map and output feature map of the Swin Transformer Block module, N =1,2,…,5, F 0 is the initial feature map extracted by the initial convolution of the normalized fused image; The SE module performs F 1 and F The processing of 2 is expressed as follows: ; In the above formula, F j ( i , k ) is the feature map F j The feature points in i , k ), j =1,2, H , W The feature maps are F j The height and width of W 1. W 2 are the weight matrices of the two fully connected layers, ReLU is the linear rectification function, and the sigmoid function maps the output value to the (0,1) interval as the attention weight. , F SEj is the output feature map of the SE module; The SOEMCM module has a characteristic map F 3. F 4 and F The processing of 5 is expressed as follows: ; In the above formula, K d is the convolution kernel, , , The feature maps are F 3. F 4. F 5 The corresponding processing result, i.e., the output feature map of the SOEMCM module; In the BiFPN layer, the first-level BiFPN layer performs top-down, bottom-up and jump connections on the feature points in the input feature map, the second-level BiFPN layer performs cross and top-down connections on the feature points in the input feature map, and the third-level BiFPN layer performs reverse top-down, reverse bottom-up and jump connections on the feature points in the input feature map. The output feature maps corresponding to various connection methods are expressed by the following formula: ; In the above formula, , Respectively l The high-level feature map and low-level feature map are Upsampled. is the output feature map corresponding to the top-down connection; K ( m , n ) is the convolution kernel K middle( m , n ), the weight at the convolution kernel K The size is s * s , For the l The feature points in the low-level feature map ( is+m , ks+n ), is the output feature map corresponding to the bottom-up connection; , Respectively l Level from path a , b The feature map of is the output feature map corresponding to the cross connection; , , Respectively l , l -1. l -2 levels from the path c The feature map of is the output feature map corresponding to the skip connection; In the output layer, the process in which the category prediction network and the boundary prediction network process the output feature map of the BiFPN layer through convolution is expressed as follows: ; In the above formula, K e , K f They are the category prediction convolution kernel and the boundary prediction convolution kernel respectively. is the output feature map of the BiFPN layer, F class 、F bbax They are category prediction feature map and boundary prediction feature map respectively.
6. The intelligent privacy protection camera system for online car-hailing according to claim 1 is characterized in that: The system uses an improved Gaussian blur algorithm to blur the target privacy area, including: Improve the network structure corresponding to the Gaussian fuzzy algorithm, including: Optimize the convolution calculation module, convert the original image containing the target privacy area and the Gaussian kernel into the frequency domain through FFT accelerated convolution, and convert them back to the spatial domain after multiplication in the frequency domain, or directly obtain the corresponding convolution result from the lookup table through lookup table-based convolution, and finally obtain the initial Gaussian blurred image; The multi-scale analysis and fusion module generates multiple secondary Gaussian blurred images based on the initial Gaussian blurred image using Gaussian kernels of different scales, calculates the local edge strength of the initial Gaussian blurred image, determines the fusion weight of each secondary Gaussian blurred image according to the local edge strength of the initial Gaussian blurred image, and adaptively fuses multiple secondary Gaussian blurred images to obtain a fused Gaussian blurred image. Based on the image content adaptation module, the Gaussian blur parameters are adaptively adjusted by calculating the local features of the fused Gaussian blurred image, so as to obtain the final Gaussian blurred image.
7. The intelligent privacy protection camera system for online car-hailing according to claim 6 is characterized in that: The optimized convolution calculation module converts the original image containing the target privacy area and the Gaussian kernel into the frequency domain through FFT accelerated convolution, and converts them back to the spatial domain after multiplication in the frequency domain, which is expressed by the following formula: ; In the above formula, I ( x , y ) is the original image, G ( x , y ) is the Gaussian kernel, x and y is the pixel coordinate, F represents Fourier transform, F -1 stands for inverse Fourier transform, I b ( x , y ) is the initial Gaussian blurred image, The standard deviation of the Gaussian distribution to control the blur level; The optimized convolution calculation module directly obtains the corresponding convolution result from the lookup table through convolution based on the lookup table, including: Precompute and store different Gaussian distribution standard deviations The Gaussian kernel and the corresponding convolution result under the original image are processed according to the local features of the input and the required Gaussian distribution standard deviation. , directly obtain the corresponding convolution result from the lookup table to get the initial Gaussian blurred image I b ( x , y ); The process of obtaining the fused Gaussian blurred image by the multi-scale analysis and fusion module includes: S1. Generate multiple secondary Gaussian blurred images using Gaussian kernels of different scales based on the initial Gaussian blurred image: ; In the above formula, G p ( x , y ) is the p Gaussian kernel, I p ( x , y ) is the p A quadratic Gaussian blurred image, p =1,2,…, P , P is the number of Gaussian kernels of different scales; S2. Calculate the local edge strength of the initial Gaussian blurred image: ; In the above formula, G X ( x , y ), G Y ( x , y ) are the initial Gaussian blurred images I b ( x , y )middle( x , y ) X , Y The gradient of the direction, E ( x , y ) is the initial Gaussian blurred image I b ( x , y )’s local edge strength; S3, according to the initial Gaussian blurred image I b ( x , y )'s local edge strength E ( x , y ) judge the edge area, for the edge area, assign a larger fusion weight to the small-scale quadratic Gaussian blurred image to retain the edge details; for the non-edge area, assign a larger fusion weight to the large-scale quadratic Gaussian blurred image to achieve a better smoothing effect; S4, adaptively fuse multiple secondary Gaussian blurred images to obtain a fused Gaussian blurred image: ; In the above formula, w p ( x , y ) is the p Quadratic Gaussian blurred image I p ( x , y )middle( x , y ), , I f ( x , y ) is a fused Gaussian blurred image; The image-based content adaptation module adaptively adjusts the Gaussian blur parameters by calculating the local features of the fused Gaussian blur image, which is expressed by the following formula: ; In the above formula, (2 M +1)(2 N +1) is the Gaussian kernel window size, I f ( x - u , y - v ) is the fused Gaussian blurred image I f ( x , y )middle( x - u , y - v ), To fuse Gaussian blurred images I f ( x , y )middle( x , y ), V ( x , y ) is the fused Gaussian blurred image I f ( x , y )middle( x , y ), p r The gray value in the local area is r The probability of a pixel is L is the gray level, H ( x , y ) is the fused Gaussian blurred image I f ( x , y )middle( x , y ), A , B , C All are preset parameters. for( x , y ) The standard deviation of the Gaussian distribution corresponding to the pixel at ; Among them, by adaptively adjusting the Gaussian blur parameters, it is possible to blur the area with rich texture and more details, that is, V ( x , y )and H ( x , y ) larger area, a smaller blur level is used, i.e. Smaller; for flat and noisy areas, that is, V ( x , y )and H ( x , y ) smaller area, a larger blur level is used, i.e. Larger.
8. The intelligent privacy protection camera system for online car-hailing according to claim 1 is characterized in that: When a dispute occurs or a passenger defends his or her rights, the driver performs face recognition through the driver terminal. The driver terminal uses an improved ArcFace algorithm to perform face recognition on the driver. If the face recognition is successful, the system retrieves the corresponding original monitoring screen according to the face recognition result and sends it to the driver terminal; When a dispute occurs or a passenger protects his or her own rights, the passenger authenticates the identity through the passenger terminal, and the passenger terminal uses the reserved mobile phone number to authenticate the passenger. If the identity authentication is successful, the system retrieves the corresponding original monitoring screen according to the identity authentication result and sends it to the passenger terminal.
9. The intelligent privacy protection camera system for online car-hailing according to claim 8 is characterized in that: The improved ArcFace algorithm introduces an adaptive multi-scale feature fusion mechanism in feature extraction: Assume the input original face image is I , through the multi-branch convolutional neural network structure, we can obtain a set of feature maps at different scales , feature map of each scale All contain information in a specific spatial frequency range. i =1,2,…, n ; Then, design an adaptive weight generation module g , according to the image content, the feature map of each scale Generate Adaptive Weights ,Right now ; The fused features Expressed as , after subsequent network layer processing, the feature vector used for recognition is obtained , and normalize ; The improved angle constraint part of the ArcFace algorithm adds an angle boundary to the traditional ArcFace algorithm. m Based on this, a dynamic adjustment factor related to sample difficulty is introduced k : Set category label y The weight vector corresponding to the sample is w y , and normalize , the eigenvector With the weight vector w y The angle between ,Right now ; Sample difficulty is measured by calculating the average distance between sample features and features of similar samples, as well as the average distance between sample features and features of other types of samples, generating a dynamic adjustment factor k : For simple samples, dynamic adjustment factor k Smaller, smaller increase in angle boundary; for difficult samples, dynamic adjustment factor k Larger, the angle boundary increases more to enhance the model's ability to distinguish difficult samples; The loss function of the improved ArcFace algorithm for: ; In the above formula, For the i The feature vector corresponding to the features after the fusion of samples is For category labels y i The weight vector corresponding to the sample of is the feature vector With the weight vector The angle between k i For the i Dynamic adjustment factor of samples k , For category labels j The weight vector corresponding to the sample of is the feature vector With the weight vector The angle between N is the number of training samples.
Citation Information
Patent Citations
Intelligent video monitoring equipment module and intelligent video monitoring system with privacy protection and intelligent video monitoring method
CN101610396A
Security and protection service network based on Internet of Things
CN103578240A
Screen recording method and mobile terminal
CN110446097A
Lightweight face recognition model-based privacy protection method
CN112766422A
Data privacy protection method based on adversarial learning
CN115936958A