PRIVACY CAMERA WITH THERMAL SENSOR-ASSISTED IMAGE PROCESSING.
Patent Information
- Authority / Receiving Office
- TR · TR
- Patent Type
- Patents
- Current Assignee / Owner
- ORTA DOGU TEKNIK UNIVERSITESI
- Filing Date
- 2021-12-31
- Publication Date
- 2026-06-22
AI Technical Summary
Existing surveillance cameras lack effective privacy protection mechanisms, leading to unauthorized recording and potential violations of personal privacy, and require significant human oversight, which is costly and inefficient.
A thermal sensor-assisted image processing privacy camera system that uses dual cameras and thermal imaging to detect and filter sensitive data, employing depth and optical flow detection, high dynamic range image processing, thermal flow analysis, and deep learning algorithms to anonymize images in real-time.
The system efficiently anonymizes sensitive data, reducing the need for human oversight and minimizing privacy violations while maintaining effective surveillance, thereby lowering operational costs and enhancing security.
Smart Images

Figure 00000018_0000 
Figure 00000019_0000 
Figure 00000019_0001
Abstract
Description
TARIFF PRIVACY CAMERA WITH THERMAL SENSOR-ASSISTED IMAGE PROCESSING. Technical Area This invention relates to the development of a thermal sensor-assisted image processing privacy camera that automatically detects all faces entering the image through automatic human detection and masks them reversibly by applying different filters and / or blurring, with the aim of preventing the violation of personal privacy caused by the unauthorized recording of images of all individuals in any environment by security cameras. Previous Technique With the advancement of camera and lens technologies, their use in daily life has become very widespread. In addition to their active use by security forces in surveillance and detection situations, their use is also increasing every day in homes and workplaces for crime protection and deterrence purposes. The use of closed-circuit television cameras (CCTV cameras) is increasing day by day. The global market for CCTV systems is worth over $3 billion, and it is projected to exceed $4 billion by 2025 (Research and Markets report). In Türkiye, the market is worth over $120 million and is expected to surpass $150 million by 2025. However, these systems bring about problems regarding the violation of personal privacy. In this context, in European Union countries, visual data is also included within the scope of personal information security, which is guaranteed by the Data Protection Directive, the EU Convention on Human Rights, and the EU Privacy Protection Directive. In our country, visual data is protected under the Personal Data Protection Law (KVKK). While the Personal Data Protection Law (KVKK) prohibits sending SMS messages without permission, we are being monitored and recorded by dozens of cameras every day. The need for 24 / 7 / 365 uninterrupted security prevents the necessary attention from being paid to privacy, and violations of our private lives occur at every point in our lives, from shopping malls to subways, from corner markets to traffic lights. Furthermore, hundreds, even thousands, of cameras operate, especially in industrial facilities, military zones, and high-security areas. A security operator can simultaneously monitor a maximum of 4 screens, and a maximum of 16 meaningful images can be displayed on a single screen. Consequently, the maximum number of cameras a single operator can monitor is 64. Various studies have shown that the attention span of operators, even with almost identical images, drops to 80% after 15 minutes, 60% after 1 hour, and below 40% after 2 hours. For these reasons, in a system with, for example, 320 cameras, the number of operators required to work simultaneously is 5, and considering 3 shifts and leave, a total of 20 people are needed for monitoring. Considering that the monthly cost per operator is $750 USD per person at today's value, and that the economic lifespan of these systems is 7 years according to TÜBİTAK's valuation, the total cost is ($750 x 20 people x 12 months x 7 years) $1.26 million USD. In other words, the cost per camera is $4000 USD. The only way to reduce operator costs is to create smart systems, and the only way to achieve this is with IP cameras that can successfully run smart video analytics software. However, simply processing images from cameras to make intelligent inferences is insufficient. This is because most cameras on the market lack privacy protection. While privacy protection is attempted outside the camera through image processing, it can be copied during transmission to the processing point and is vulnerable to cybersecurity attacks. Therefore, even though these cameras can offer significant benefits, they cannot be used in patient rooms, prison cells, or for monitoring elderly and vulnerable individuals. However, directly protecting privacy on the camera itself enables applications that were previously impossible. Because the privacy software is embedded in the camera system, the transmission of any privacy-violating images over the network is prevented, eliminating the need and cost of high-end computer processing units (CPUs), random access memory (RAM), and graphics processing units (GPUs) required for image processing. Thanks to the efficient video analytics obtained through the use of dual cameras, much more efficient, event-oriented, analytics-based monitoring operations can be performed with 1 / 3 of the operators. The invention described in PCT application WO2017132074A1, which is a known application of the art, involves a system that performs autonomous photography while recording only the images of the relevant target. The system determines whether autonomous operation of a camera on an unmanned aerial vehicle (UAV) is permitted in the relevant area and records only the relevant images along with location and geographic information. While the developed system essentially uses a single camera to capture the relevant image, it does not perform continuous recording. Furthermore, it has shortcomings in preventing the recording of images that may be needed later. Another known application of the technique is the United States patent application US20070286520A1. This application describes a system for use in a desktop video conferencing application that blurs the background image, keeping only the participant's view visible. The system is designed for real-time use and does not involve recording or autonomous detection. Therefore, it does not aim to prevent the capture of anonymous images of individuals or a single camera. Similarly, the invention described in the Chinese patent application CN1111524060A also describes background blurring and blurring of unwanted images. Brief Description of the Invention The aim of this invention is to develop a thermal sensor-assisted image-processing privacy camera to ensure the protection of personal data in images taken of individuals in public and private spaces. Another aim of this invention is the development of a thermal sensor-assisted image-processed privacy camera that allows for the re-encryption of relevant obscured images during any social event, while taking public safety into consideration. Another objective of this invention is the development of a thermal sensor-assisted image processing privacy camera capable of applying the most appropriate filtering to the environment using different filtering techniques. Another aim of this invention is the development of a thermal sensor-assisted image-processed privacy camera that, with the support of thermal cameras, provides clear identification of people in the environment. Another objective of this invention is the development of a thermal sensor-supported image processing privacy camera that uses artificial intelligence and deep learning algorithms to improve the efficiency of the model's results by processing personal data received from thermal cameras. Another goal of this invention is to develop a privacy camera with thermal sensor-assisted image processing that uses two cameras simultaneously to provide depth perception along with the most accurate filtering. Another goal of this invention is to develop a thermal sensor-supported image-processed privacy camera that, using relevant artificial intelligence and deep learning algorithms, can not only identify individuals but also filter out various details that can personalize the data in the environment. Definitions of the Figures Illustrating the Invention The figures and accompanying explanations used to better illustrate the thermal sensor-assisted image processing privacy camera developed with this invention are given below. Figure 1 shows a flowchart illustrating the operating algorithm of the thermal sensor-assisted image processing privacy camera according to the invention. Figure 2 is a schematic view of the flowchart illustrating the depth and optical flow detection algorithm of the thermal sensor-assisted image processing privacy camera according to the invention. Figure 3 shows a schematic view of the flowchart illustrating the high dynamic range image generation algorithm for a thermal sensor-assisted image processing privacy camera, according to the invention. Figure 4 shows a schematic view of the flowchart illustrating the high dynamic range tone matching algorithm of the thermal sensor-assisted image processing privacy camera according to the invention. Figure 5 is a schematic view of the flow showing the thermal flow algorithm of the thermal sensor-assisted image processing privacy camera according to the invention. Figure 6 shows a schematic view of the flowchart illustrating the deep learning module algorithm for the thermal sensor-assisted image processing privacy camera, according to the invention. Figure 7 shows a schematic view of the flowchart illustrating the encoding and broadcasting algorithm for a thermal sensor-assisted image-processed privacy camera, according to the invention. Figure 8 is a schematic view illustrating the filtering process of a thermal sensor-assisted image-processed privacy camera, according to the invention. Figure 9 shows a schematic view of the subsystems of the thermal sensor-assisted image processing privacy camera according to the invention. The elements shown in the figures are numbered, and their corresponding elements are given below. 1. Privacy system 2. Camera system 3. First camera 4. Second camera 5. Thermal camera 6. First chip 7. Second chip 8. Lens 9. Sensor 10. Digital signal processor 11. Integrated circuit 12. Reading memory 13. Dynamic memory 14. Network card 15. Access point Detailed Description of the Invention The invention is essentially a privacy system (1) that filters the images it receives with a camera system (2) by processing them in a way that hides the sensitive data of individuals. The camera system (2) contains 3 cameras: the first camera (3), the second camera (4) and the thermal camera (4). The first camera (3) and the second camera (4) operate with two different microprocessors, the first chip (6) and the second chip (7). The privacy system (1) performs the filtering by completing the main algorithm in seven steps. Each step also contains algorithm steps within itself. Essentially, the privacy system includes the following steps: (1) main algorithm, (101) Depth detection, (102) Optical flow detection, (103) LDA image generation, (104) LDA tone matching algorithm, (105) Thermal flow, (106) Encoding and publishing within the deep learning module. The system ensures the acquisition of a filtered image as a result of the main algorithm. The privacy system (1) uses a dual camera system, the first camera (3) and the second camera (4), to create a depth map and heat map. More cameras can also be used within the privacy system (1). In addition, a low-resolution thermal camera (5) is used to detect areas such as the human face and body where privacy needs to be ensured, on high-resolution image data with the help of cameras. Depth detection (101) and optical flow detection (102), (201) Images are taken from the dual camera system at different exposures. (202) The distortion effect in the images is removed by using calibrated calculated values. (203) Distortion-reduced images are oriented horizontally so that a depth map can be created. (204) Optical flow and depth map are extracted from the corrected and depth map prepared images using total variation optical flow. (205) Noise in the depth map is removed by using resolution-calculated wear and expansion filters. It is carried out in steps. Depth detection (101) and optical flow detection (102) steps are carried out with the help of the first camera (3) and the second camera (4). Images from the dual camera system are edited to correct any errors. Once all edits are complete, an optical flow and depth map is generated. After this map is created, any distortions in the depth map are corrected to achieve the final result. The cameras have the ability to capture high dynamic range images. They achieve this by using both stereo (dual aperture) and variable exposure (exposure bracketing) to produce a high dynamic range image stream, enabling them to capture images even in challenging lighting conditions. High dynamic range image creation (103), (301) Images with different exposure values are taken from the camera. (302) Future exposure values are calculated according to changing light conditions, using the image with the medium exposure value. (303) The images taken at different exposure values are converted into RGB (Red-Green-Blue) images. (304) Exposure matching is performed using the available exposure values for combining the images. (305) A separate normalized weight map is generated for each position according to the pre-calculated table. (306) A single high dynamic range image is created using calculated weight maps and different exposures. It consists of steps. In the high dynamic range image creation (103) stage, the images transmitted from the camera are adjusted according to the variable light conditions based on the values of the determined reference exposure. These exposure values are also used in the images to be taken later. The images converted to RGB images are adjusted according to the values of the acquired images and combined. Errors on the exposures are corrected and mapped using pre-calculated data and weight values that do not have a unit. All data are collected to create a dynamic range image. The tone matching process is then performed with the created image. The high dynamic range tone matching (104) algorithm, (401) The generated high dynamic range image is converted into a single-channel image by multiplying it with predetermined and different coefficients for each channel. (402) The average logarithmic brightness value of the generated single-channel image and the minimum and maximum values are calculated, excluding extreme %1Ίik values. (403) The calculated values are used together with a predetermined curve value, saturation coefficient and single channel image to create a tone matching filter. (404) The tone matching filter is multiplied by the original high dynamic range image and the tone matching result is created. (405) The resulting tone matching is reduced to the 0-1 range and made ready by making gamma correction according to the determined gamma value. It is carried out within the framework of these steps. After the high dynamic range image generation (103) stage, the acquired image is converted into a single-channel image by editing it with certain coefficients. Maximum and minimum values are calculated independently of the average logarithmic brightness and extreme values. Using these values, the curve value, image saturation coefficient and the resulting single-layer image are used to create a tone matching filter. The resulting tone matching filter and the image obtained in the high dynamic range image generation (103) stage are multiplied. The resulting tone matching result is normalized to obtain a gamma value in the range of 0-1, and the high dynamic range tone matching (104) step is performed by removing errors from this value. Thermal cameras (5) enable the operation of all video analytics, including face detection and unauthorized access, with a high success rate. In this context, thermal flow (105) steps are carried out with the use of thermal cameras (5). Thermal flow (105), (501) Thermal images are obtained from the thermal camera as a result of correct configuration. (502) The acquired thermal image is resized so that it can be used together with the image from the main camera. (503) A binary filter is generated from the newly resized thermal image using a predetermined threshold value. (504) Noises in the filter are eliminated by using wear and expansion filters calculated according to resolution. (505) The image is filtered with a filter created before face detection and identification. This is carried out through steps. With the correct positioning of the thermal camera (5), the thermal images obtained are matched and edited with the images obtained from the first camera (3) and the second camera (4). The final version of the image is filtered and edited with the threshold value determined by the calculations beforehand. This image obtained at the end of the thermal flow (105) is used as the first filter before artificial intelligence filtering. After the thermal flow (105) step, the images obtained are improved with the relevant artificial intelligence and deep learning algorithms, and the existing filtering is also regulated. In addition, the deep learning module (106) ensures that the system learns continuously in the future and prevents errors. Within the deep learning module (106), (601), the generated high dynamic range image is resized according to the determined sizing factor. (602) The newly created image is fed to the deep learning module for face detection. (603) The positions of the faces in the picture are taken. (604) A certain number of landmark coordinates are taken for each face. (605) The specified face regions are enlarged in such a way as not to distort the coordinate points they contain. (606) The parts of the original image that correspond to the newly calculated facial regions are calculated and fed to the deep learning module. (607) A certain number of pairs of values are taken to determine the facial features for each face. (608) The values determined for each face are compared with the face values recorded in the database. (609) If it exceeds the specified threshold value, it is considered to be the most similar face. It includes the steps. The image created within the deep learning module (106) and filtered in the thermal flow (105) step is resized with a specified coefficient. In addition to this resized image, the original image continues to be stored. The image defined in the deep learning module first determines the positions of the faces for face detection. Different numbers of signal coordinates are obtained for each face depending on the use case. Faces whose coordinates and positions have been determined are enlarged to the desired and necessary dimensions without distorting their coordinates. The original image and the calculated image are matched and input into the deep learning module. At this stage, values are determined for each face, preferably 64 pairs, but which may vary depending on the use case, indicating the face line. These values are compared with the learning results obtained in the database. If the threshold value is exceeded, the face that most closely resembles the original is accepted. Encoding and publishing step (107), (701) The filtered image is converted into a different color model and encoding format for encoding. (702) The image, converted to a different format, is encoded into a compression format using the embedded device's hardware encoder. (703) Extra information is encrypted and added to the data encoded in the compression format. (704) Encoded compressed data is parsed for publication along with the encrypted data. (705) The parsed data is broadcast over the network with a multimedia software and a specified network protocol. It consists of steps. The filtered image is converted to a format with lower bandwidth. YUV420 is preferred for this format, although other formats can be used. The converted image is encoded into a compression format using the embedded system's hardware encoder. H264 is preferred for this compression format, but other formats can also be used. Filtered facial regions, high dynamic range image parameters, and similar extra information are encrypted into the compressed data. The encrypted information is then parsed as SEI (Self-Encrypting Information) data, and the compressed data is then broadcast. The parsed data is broadcast over the network using a multimedia software via Real-Time Streaming Protocol (RTSP). Ideally, an RTSP server library using the Gstreamer multimedia software structure can be used. The privacy system (1) uses 5 different filters in the filtering process. These are applied as (801) Mask filter (802) Reverse color filter (803) Bilateral filter (804) Icing filter (805) Pixelation filter. Different filters can also be applied depending on different purposes and usage points. Face regions and hand-drawn regions detected within the mask filter (801) are taken. Pixel values in the regions are recalculated according to the selected color and opacity level. The inverse color filter (802) is a filter that takes the detected face regions and hand-drawn regions. It is mapped according to the new values of the existing pixel using predefined tables according to the selected palette and compression. The bilateral filter (803) takes the detected facial regions and hand-drawn regions. A custom edge-protecting filter is created for the determined values and the regions taken. The image is filtered using a filter in this way. The frosting filter (804), like other filters, takes the detected face regions and hand-drawn regions. To eliminate noise, an automatic Gaussian filter is calculated according to the resolution. The image is filtered by changing the kernel size used according to the selected amount of frosting. In the pixelation filter (805), the detected face regions and hand-drawn regions are taken in the same way. The pixel size in the specified areas is calculated according to the selected pixelation amount and the available resolution. The specified pixel areas are equated to the average pixel value in that region. Privacy and security issues are essentially two sides of a scale, balancing elements. Increased privacy naturally leads to decreased security, and increased security naturally leads to decreased privacy. The privacy software embedded in the cameras within the privacy system (1) will allow monitoring of all events occurring on the scene while preventing the faces of individuals from being viewed by anyone other than authorized personnel. This will enable both increased privacy and security simultaneously. The privacy system (1) uses artificial intelligence and deep learning algorithms to filter the faces and selected areas of individuals. However, with the help of artificial intelligence and deep learning algorithms, all sensitive data of individuals can be filtered. It can be used in offices and similar environments to filter confidential information of companies in security logs. Due to the nature of deep learning models, the well-known continuous learning model can be applied here. Furthermore, the system possesses the ability to make decisions on security matters, such as evaluating images in situations involving criminal activity and transmitting information to relevant authorities. The camera system (2) consists of three cameras, the first camera (3), the second camera (4) and the thermal camera (5), and a digital signal processor (10) that processes their data. The first camera (3) and the second camera (4) are placed to capture images and add depth to the image. The thermal camera (5) is used to perform the first filtering when identifying people in the captured image. The thermal camera (5) monitors the body temperature of individuals and detects the exposed areas of their skin. In this way, it can filter out features such as skin color, face shape, and hair structure, which are sensitive data that distinguish individuals. The first camera (3) and the second camera (4) contain two microchips, the first chip (6) and the second chip (7). The microchips have different functions and are integrated into the camera. The cameras have one lens (8). The lens (8) transmits the image it receives to the system, and after the image is processed, it is sent to the system via an access point (15). From there, the image is processed on the server. Communication is done according to the LAN / Internet protocol. Different communication methods can also be used depending on the design and implementation. The first chip (6) contains a sensor (9) and a digital signal processor (10) (Digital Signal Processor DSP). Sensor (9) is an image sensor and detects the received image. The received image is transmitted to the processor (10) for the relevant image processing. Once the image is finalized on the first chip (6), the relevant data is transmitted to the second chip (7). The second chip (7) contains a network card (12), an integrated circuit (11), and read memory (12) and dynamic memory (13) connected to this integrated circuit (11). The integrated circuit (11) is a system that can contain all kinds of sub-components of a computer element on a system-on-a-chip (SoC). The integrated circuit (11) system performs the compression and decompression of video and audio. Although H264 compression format is preferred, it is also possible to perform different operations. The integrated circuit (11) has a read memory (12) for storing the operations it performs and a dynamic memory (13) called Random Access Memory (RAM) to increase the speed of operations. All data processed on the integrated circuit (11) is transmitted to the server via the relevant access point (15) with the help of a network card in the system with any type of connection.
Claims
REQUESTS 1. To prevent the violation of personal privacy caused by the unauthorized recording of the images of all persons in any environment by security cameras, to process the images obtained by the camera system (2) through the first camera (3) and the second camera (4); to obtain images from the dual camera system in different positions (201), to remove the distortion effect in the images using calibrated and calculated values (202), to straighten the distortion-removed images horizontally so that a depth map can be created (203), to create optical flow and depth map from the corrected and depth map prepared images using total variation optical flow (204), to remove noise in the depth map using erosion and expansion filters calculated according to the resolution (205), and to determine the depth in the images (101) as a result of these steps.The deep learning module (106) is completed as a result of the following steps: - the image is fed into the deep learning module for face detection (602), the positions of the faces in the image are obtained (603), a certain number of landmark coordinates are obtained for each face (604), the parts where the face regions correspond in the original image are calculated and fed into the deep learning module (606), - the image, which is filtered with a mask filter (801) or inverse color filter (802) or bilateral filter (803) or frosting filter (804) or pixelation filter (805), is converted into a different color model and encoding format for encoding (701), - the image converted to a different format is encoded into a compression format using the hardware encoder of the embedded device (702), extra information is added to the encoded data in the compression format by encrypting it (703), the encoded compressed data together with the encrypted data is parsed for publication (704),a privacy camera system characterized by a digital signal processor (10) that operates its steps.
2. To process the images obtained by the first camera (3) and a thermal camera (5) in order to filter sensitive data in the image data received by the cameras and to enable access to them again when needed; taking images at different exposures from the dual camera system (201), removing the distortion effect in the images using calibrated and calculated values (202), straightening the distortion-removed images horizontally so that a depth map can be created (203), creating optical flow and depth map from the corrected and depth map prepared images using total variation optical flow (204), removing noise in the depth map using resolution-calculated wear and expansion filters (205), and as a result of these steps, depth detection in the images (101) is made, and thermal image is taken from the thermal camera (501).The thermal flow (105) is completed as a result of the following steps: resizing the acquired thermal image so that it can be used together with the image from the main camera (502), generating a binary filter from the newly resized thermal image using a predetermined threshold value (503), removing noise in the filter created using resolution-calculated wear and expansion filters (504), converting the image filtered with a mask filter (801) or inverse color filter (802) or bilateral filter (803) or icing filter (804) or pixelation filter (805) into a different color model and encoding format for encoding (701), encoding the converted image into a compression format using the embedded device's hardware encoder (702), adding extra information to the encoded data in the compression format by encrypting (703),a privacy camera system characterized by a digital signal processor (10) which runs the steps of parsing the encoded compressed data together with the encoded data for broadcasting (704).