Detecting windshield reflection
The method and system generate averaged frames to distinguish and filter windshield reflections, improving detection accuracy and privacy in outward-facing vehicle cameras by removing or blurring sensitive content.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NETRADYNE INC
- Filing Date
- 2025-10-14
- Publication Date
- 2026-04-23
AI Technical Summary
Outward-facing vehicle cameras capture windshield reflections, leading to false positives in computer vision systems, degrading detection accuracy and privacy risks by revealing sensitive content.
A method and system that utilize an image capturing device to generate an averaged frame from multiple frames, detect static windshield reflections, and apply privacy filters to mitigate privacy risks by removing, blurring, or blacklisting such reflections.
Enhances detection accuracy by distinguishing external features from internal reflections and effectively filters out sensitive content, ensuring privacy compliance and reliable data sharing.
Smart Images

Figure US2025050798_23042026_PF_FP_ABST
Abstract
Description
DETECTING WINDSHIELD REFLECTIONCROSS-REFERENCE TO RELATED APPLICATIONS[0001.1] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 707,238, filed October 15, 2024, the entirety of which is incorporated by reference herein.FIELD OF THE INVENTION
[0001] Certain aspects of the present disclosure generally relate to processing and sharing of images collected by cameras in vehicles, and more particularly to systems and methods of filtering imagery to mitigate privacy risks that may be associated with objects that are inside a vehicle or attached to the vehicle and that may be detectable by an outward-facing camera of the vehicle.BACKGROUND
[0002] The information disclosed in this background section is only for the enhancement of understanding of the general background of the disclosure and should not be taken as an acknowledgement or any form of suggestion that this information forms the prior art already known to a person skilled in the art.
[0003] Modem vehicles are increasingly equipped with outward-facing cameras for a variety of applications, including advanced driver assistance systems (ADAS), fleet monitoring, driver coaching, and large-scale mapping. The outward-facing cameras, despite being mounted to capture external scenery, are often positioned behind the windshield or in other locations where reflective surfaces are present. The positioning of the outward-facing cameras may thus introduce a technical challenge such as visual reflections of the interior of the vehicle may be captured in the camera’s field of view (FOV). The reflections include dashboard elements, objects placed on the dashboard, or materials attached to the windshield. Depending on lighting conditions and other factors, the reflections may appear superimposed on the captured road scene and thus introduce visual content that is unrelated to the intended purpose of environmental data collection.
[0004] The presence of reflections introduces multiple technical problems in downstream processing pipelines. For instance, computer vision systems that attempt to detect and classify road signage, lane markings, or other external objects may produce false positives or degraded accuracy when windshield reflections are present. Reflections of papers, placards, or electronic devices can be misinterpreted as legitimate environmental features. This not only affects the precision of2detection algorithms but also reduces the reliability of data used for mapping and safety-critical applications.
[0005] Beyond degraded algorithmic performance, a further technical problem arises with respect to privacy. Windshield reflections may reveal sensitive content, such as personal documents, identification cards, handwritten notes, partial images of a driver or passenger, and the like. The unintended capture of such content creates a privacy risk, particularly when vehicle-sourced imagery is shared with third-party services. Even if external scenes are anonymized, the presence of reflections can reintroduce personally identifiable information (PII), undermining the expectation of driver anonymity in shared datasets. Thus, systems and methods that can effectively mitigate privacy risks associated with outward-facing imagery containing interior reflections solves a critical technical problem.
[0006] Accordingly, certain aspects of the present disclosure are directed to reliable detection of the reflections and filtering out associated imagery from data sharing.SUMMARY
[0007] This summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the invention. This summary is neither intended to identify key or essential inventive concepts of the invention nor is it intended for determining the scope of the invention.
[0008] According to one embodiment of the present disclosure, a method for detecting windshield reflections of an object in a vehicle using an image capturing device installed inside the vehicle is disclosed. The method includes capturing a media sequence from the image capturing device while the vehicle is in motion. Further, the method includes generating an average frame based on computing, for each pixel location, an average of pixel values across a plurality of frames of the media sequence. The plurality of frames is captured while the vehicle is in motion. Further, the method includes detecting a static object in the vehicle that is reflected in a windshield of the vehicle based on processing the average frame. Furthermore, the method includes generating at least one privacy filter tag in response to detecting the static object. The at least one privacy filter tag indicates that the static object is subject to privacy filtering and comprising instructions to at least one, remove the static object from the media sequence, blur the static object in the media sequence, and blacklist the image capturing device.
[0009] According to one embodiment of the present disclosure, a system for detecting windshield reflections of an object in a vehicle is disclosed. The system includes an image capturing device installed inside the vehicle and a processing unit in communication with the image capturing device. The processing unit is configured to capture a media sequence from the image capturing device while the vehicle is in motion. The processing unit is configured to generate an averaged frame based on computing, for each pixel location, an average of pixel values across a plurality of frames of the media sequence. The plurality of frames is captured while the vehicle is in motion. Further, the processing unit is configured to detect a static object in the vehicle that is reflected in a windshield of the vehicle based on processing the averaged frame. Further, the processing unit is configured to generate at least one privacy filter tag in response to detecting the static object. The at least one privacy filter tag indicates that the static object is subject to privacy filtering and comprising instructions to at least one of remove the static object from the media sequence, blur the static object in the media sequence, and blacklist the image capturing device.
[0010] To further clarify the advantages and features of the present invention, a more particular description of the invention will be rendered by reference to specific embodiments thereof, which are illustrated in the appended drawings. It is appreciated that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting of its scope. The invention will be described and explained with additional specificity and detail in the accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] These and other features, aspects, and advantages of the present invention will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:
[0012] FIGURE la illustrates an example implementation environment of a system for detecting windshield reflections of an object in a vehicle, in accordance with an embodiment of the present invention;
[0013] FIGURE lb illustrates a block diagram of the example implementation environment of the system, in accordance with an embodiment of the present invention;
[0014] FIGURE 2 illustrates a block diagram of the system, in accordance with an embodiment of the present invention;
[0015] FIGURE 3 illustrates a functional block diagram of the system with modules for detecting the windshield reflections of the object in the vehicle, in accordance with an embodiment of the present invention;
[0016] FIGURE 4 illustrates a flow chart of a method for detecting the windshield reflections of the object in the vehicle, in accordance with an embodiment of the present invention; and
[0017] FIGURES 5-7 illustrates example use-cases depicting the detection of the windshield reflections of the object in the vehicle, in accordance with an embodiment of the present invention.
[0018] Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help to improve understanding of aspects of the present invention. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present invention so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.DETAILED DESCRIPTION OF FIGURES
[0019] For the purpose of promoting an understanding of the principles of the invention, reference will now be made to the embodiment illustrated in the drawings and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the invention is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the invention as illustrated therein being contemplated as would normally occur to one skilled in the art to which the invention relates.
[0020] It will be understood by those skilled in the art that the foregoing general description and the following detailed description are explanatory of the invention and are not intended to be restrictive thereof.5
[0021] Reference throughout this specification to “an aspect”, “another aspect” or similar language means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrase “in an embodiment”, “in one embodiment”, “in another embodiment”, and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.
[0022] The terms “comprise”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process or method that comprises a list of steps does not include only those steps but may include other steps not expressly listed or inherent to such process or method. Similarly, one or more devices or sub-systems or elements or structures or components proceeded by “comprises... a” does not, without more constraints, preclude the existence of other devices or other sub-systems or other elements or other structures or other components or additional devices or additional sub-systems or additional elements or additional structures or additional components.
[0023] The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as to not unnecessarily obscure the embodiments herein. Also, the various embodiments described herein are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments. The term “or” as used herein, refers to a non-exclusive or unless otherwise indicated. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein can be practiced and to further enable those skilled in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein.
[0024] As is traditional in the field, embodiments may be described and illustrated in terms of blocks that carry out a described function or functions. These blocks, which may be referred to herein as units or modules or the like, are physically implemented by analog or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits, or the like, and may optionally be driven by firmware and software. The circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuitboards and the like. The circuits constituting a block may be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiments may be physically separated into two or more interacting and discrete blocks without departing from the scope of the invention. Likewise, the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the invention.
[0025] The accompanying drawings are used to help easily understand various technical features and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. As such, the present disclosure should be construed to extend to any alterations, equivalents, and substitutes in addition to those which are particularly set out in the accompanying drawings. Although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are generally only used to distinguish one element from another.
[0026] Certain road mapping and infrastructure monitoring systems may rely on image and video sequences captured by vehicles in motion to provide real-time updates on road conditions, traffic signage, and environmental context. To enable scalability, such systems are designed to collect imagery from thousands of vehicles and aggregate the data for use in centralized applications such as high-frequency map updating.
[0027] One approach to mitigating reflection risks relies on manual review of individual frames or video sequences. Human reviewers are tasked with examining captured frames to identify and discard those that contain visible reflections that violate an anonymization objective. This approach is neither scalable nor efficient. Further, due to how images or video sequences may be sampled, human review may may be expected to miss windshield reflections that are only apparent in some frames and not others. At fleet or map-update scale, the sheer volume of imagery precludes human-only review. Moreover, manual review itself poses privacy risks, since it exposes sensitive information to human operators before any anonymization is applied. Consequently, manual review is not only resource-intensive but may also exacerbate the problem and be counterproductive from a privacy standpoint.
[0028] Single frame automated detection techniques for windshield reflections is an approach involving the application of deep learning classifiers or object detection pipelines to individual frames. Such techniques may suffer from technical limitations. In single frames, reflections are7often subtle and confounded with environmental features. Lighting conditions, shadow patterns, and roadside clutter can mask or mimic reflections, making it difficult for techniques to reliably distinguish between genuine external objects and interior reflections. For example, reflections of a rectangular document may visually resemble a road sign, while reflections of dashboard textures may resemble crosswalk lines or lane markings. These ambiguities may lead to an unacceptable percentage of misclassifications.
[0029] The technical difficulty is exacerbated by variability in environmental conditions. Reflection visibility changes dramatically with factors such as time of day, sun angle, artificial lighting, and weather conditions. A reflection that is nearly invisible in one frame may become prominent in the next due to a small change in illumination. This temporal inconsistency further challenges frame-based detection techniques, since single image sampling may miss reflections that are discernible in some frames of a short video sequence but not frames that were sampled.
[0030] Another layer of complexity arises from the interaction between reflections and vehicle motion. While the external scene shifts continuously as the vehicle moves, objects inside the vehicle remain fixed relative to the camera. Consequently, resulting in a peculiar visual artifact, including the stationary appearance of internal reflections, while external features shift. However, in single frames, this distinction is not obvious, especially when external features such as vehicles travelling at similar speeds remain relatively static. As a result, single-frame approaches cannot reliably exploit motion cues to isolate reflections from the environment.
[0031] Accordingly, the inability to separate internal reflections from external features introduces several downstream challenges. First, erroneous detections may waste computational resources by introducing irrelevant objects into vision pipelines (such as vision pipelines that apply additional processing for each detected object). . Second, the failure to identify and suppress sensitive reflections risks the dissemination of PII at scale, undermining compliance with privacy regulations and damaging user trust, ultimately limiting the utility of collected data for mapping and safety applications.
[0032] FIGURE la illustrates an example implementation environment for detecting windshield reflections of an object in a vehicle, in accordance with an embodiment of the present invention. FIGURE lb illustrates a block diagram of the example implementation environment, in accordance with an embodiment of the present invention.
[0033] Referring now to FIGURES la and lb combined, there is shown an example implementation environment 100 for detecting windshield reflections of an object in a vehicle. A vehicle 102 includes a windshield 106 through which an image capturing device 108 views an outward field of view (FOV). The image capturing device 108 such as a front-facing camera, optical sensor, or dash-mounted video recorder is installed inside the vehicle 102 and oriented toward the roadway. Although FIGURE la depicts the image capturing device 108 on the interior surface of the windshield 106, other locations inside the cabin (i.e., the vehicle) that provide an outward view are equally suitable. The optical axis of the image capturing device 108 faces outside so as to record the external environment.
[0034] During normal use, a driver or occupant may place various objects, such as documents, identification cards, notebooks, or placards, on the dashboard of the vehicle. One such object 104a (i.e., a document) may reflect from the inner surface of the windshield 106 and form a reflected image 104b (also referred to as the windshield reflection) within the FOV of the image capturing device 108. When the image capturing device 108 captures outward-facing imagery for use in mapping, fleet-management, or driver-safety applications, the reflected image 104b becomes embedded in the recorded imagery. Consequently, the reflected image 104b may introduce a privacy risk in any dataset that is later shared externally, because the reflected object 104a may contain personally identifiable or private information.
[0035] In an embodiment, the image capturing device 108 (also referred to as camera 108 for the sake of brevity) continuously captures a plurality of frames (referred to as the “frame” for the sake of brevity) while the vehicle 102 is in motion, thus producing a media sequence. Each frame represents an optical snapshot of the external scenery plus any reflections from within the cabin of the vehicle 102. In an example, capturing occurs at a predetermined frame rate, such as 15, 30, or 60 frames per second and may be triggered by motion data available on the vehicle’s network or sensors (for example, wheel rotation sensors, an accelerometer, or a GPS unit). In an advantageous aspect, capturing while the vehicle 102 is moving ensures that stationary external objects shift position in the camera FOV between frames, whereas interior reflections remain spatially fixed with respect to the camera 108. Thus, the relative motion difference forms the foundation for distinguishing between dynamic external features and static reflective content.
[0036] In an embodiment, the media sequence may cover a predefined duration or predefined distance, such as five to fifteen seconds of travel or a segment of ten to one hundred meters. The predefined duration or predefined distance is selected so that sufficient visual diversity is capturedin the external environment to allow moving scenery to be averaged away for later processing steps. Vehicle motion signals may further be logged so that a processing unit (shown in Figure 2) may confirm that the sequence corresponds to continuous travel and not stationary conditions, since no useful motion-based separation occurs when the vehicle 102 is parked.
[0037] In an embodiment, in response to capturing the media sequence, an averaged frame is computed based on averaging pixel values across the frames in the sequence. For each pixel location, the system calculates a mean or another statistical aggregate of corresponding pixel intensities over the selected number of frames. The selected set can be defined by a predetermined frame count (for example, 100 frames) or a time window (for example, 17 seconds). Advantageously, the averaged frame is a production of a synthetic image that visually emphasizes stationary content and suppresses motion.
[0038] As the vehicle moves forward, external features such as road markings, vehicles, or trees appear in different positions in successive frames. When those frames are averaged, the changing pixels blend together, forming a blurred background. In contrast, interior reflections such as the reflection 104b of the document 104a remain fixed relative to the camera 108, causing their corresponding pixel values to reinforce one another in the average. Consequently, the averaged frame highlights static reflections while muting the moving scenery. The resulting frame is therefore better suited to the task of detecting stationary reflective artifacts associated with incabin objects.
[0039] In an embodiment, averaging employs one or more techniques, such as temporal mean, weighted mean emphasizing recent frames, or adaptive averaging that stops when pixel variance converges. Regardless of implementation, the effect remains consistent, i.e., dynamic external content becomes visually smoothed, while static reflections persist. In an example, averaging approximately 100 frames representing roughly five to fifteen seconds of travel yields a clear distinction between blurred external scenery and sharp in-cabin reflections.
[0040] Further, because lighting conditions vary, the averaged frame may be normalized for brightness and contrast. In an example, at night or in tunnels, raw averaged frames may appear underexposed conversely, in bright daylight the averaged frames may saturate. In the example, normalization using techniques such as histogram equalization, gamma correction, or adaptive contrast stretching advantageously ensures that reflective regions remain visible even under challenging lighting. In an advantageous aspect, for nighttime imagery, the normalization mayamplify faint reflections caused by interior lighting or dashboard glow, thereby improving downstream detection accuracy.
[0041] Thus, the generation of the average frame advantageously solves a key technical limitation of conventional single-frame reflection detection. In single frames, environmental shadows and roadside features often resemble reflections, leading to false detections or missed reflections. Conversely, the average frame creates a temporal composite, eliminating ambiguities and provides a stable image on which reliable detection of the reflection is performed.
[0042] In an embodiment, the averaged frame is analyzed to detect a static object in the vehicle 102 that is reflected in the windshield 106. The detection proceeds through a structured computervision (CV) pipeline that includes pre-processing, segmentation, and content-specific analysis for identifying sensitive content such as documents, text, logos, or human faces that persist as the reflection.
[0043] In an embodiment, pre-processing includes grayscale conversion and gradient-field computation to emphasize sharp intensity transitions that may be characteristic of certain reflection edges. As the external scenery appears blurred in the frames, resulting in weak gradients; in contrast, the reflection edges remain well defined. Thus, analysis of the gradient magnitude field, consequently, isolates candidate zones likely to contain the reflective content (i.e., the reflection 104b).
[0044] In an embodiment, the averaged frame is then subjected to semantic segmentation, dividing the average frame into regions based on visual similarity or spatial context. A segmentation machine learning (ML) model trained for automotive imagery may identify areas corresponding to windshield zones versus the external background. The output of the segmentation ML model is a segmented frame in which each region is labeled, allowing subsequent ML models to focus on high-confidence reflective regions while ignoring irrelevant areas like sections of clear sky or the road surface which may appear normal due to their regularity.
[0045] Once the reflective regions are identified, specialized detection ML models are applied, for instance, optical character recognition (OCR) to identify legible text, face-detection networks to identify human faces, text-at-any-orientation models to capture mirrored or rotated writing, and object-detection neural networks for general reflective artifacts. Each ML model outputs potential detections along with confidence scores.
[0046] In one example, if using the OCR, prior to processing the averaged image with OCR, the averaged image may be rotated and / or flipped. The pre-processing may be helpful if the model or CV algorithm (such as the OCR), was trained on properly oriented text.
[0047] In some windshield reflection situations, there may be no discernible text, but there may be a reflection of a document, notebook, or the like, on which text or other personal or private information may be typically written. In the example shown in FIGURE 1, the reflection 104b may be from a spiral bound notebook (i.e., the object 104a) and is discernible as a reflection. Such examples may be detectable by processing the averaged image with an object detection model capable of detecting notebooks or paper. Likewise, as discussed in more detail below, the spiralbound notebook may be detected using an image-to-text model that can describe elements that appear to be contained in an image. Given that disclosure of the contents of notebook may be considered a privacy risk, such detections may be flagged for privacy filtering, as described above for discernible faces or legible writing.
[0048] In some other example scenarios, the text associated with a visual element inside the vehicle 102 and reflected by the windshield 106 may be blurry in the averaged frame but may be less blurry and / or legible in single frames. For example, in an example in which a package is left on the vehicle dashboard, it may be apparent from the processing of the averaged image that there is a reflection of a package, the package containing a shipping label, but the textual contents of the label itself are not discernible. The textual contexts may be more legible in single images, such as may be requested by a mapping consumer of such imagery. For this reason, it may be more fruitful to process an averaged image to detect a package rather than to attempt to detect text, as the text may not be legible in the averaged image due to small movements of the package.
[0049] As another example, a clipboard on the dashboard containing a document that contains text, might vibrate or move slightly as the vehicle moves, making the text blurry in the averaged image. In such a case, if one frame is selected, the text on the document may be visible and discernible. If situations like these are expected or possible, then selection of a computer vision model that is capable of detecting documents, clipboards, and / or blurry text may be preferable to reliance on a computer vision model that is only capable of detecting regular (unblurred) text.
[0050] Other pre-processing steps may be beneficial based on environmental factors. Nighttime conditions, for example, may present additional challenges, such that windshield reflections may be difficult to distinguish against an otherwise dark external scene. In this case, image normalization may be applied to enhance brightness and contrast. After averaging and optionally12normalizing, the set of images, a computer vision algorithm, a neural network, or the like, may be applied. Various image normalization processing may be performed before or after the averaging. For example, in one embodiment, the steps may comprise instructing a processor to average the frames, normalize the averaged image for brightness and contrast, then processing the averaged, normalized image with one or more computer vision models to detect things that are potentially concerning for privacy.
[0051] In an embodiment, the detection process may also reference metadata such as vehicle speed or sun angle. Reflection visibility is known to change with illumination direction therefore; detections may be correlated with timestamps and solar position to better distinguish reflections from external objects that share similar geometry. In an advantageous aspect, such context awareness enhances robustness under variable conditions.
[0052] Consequently, a static object is identified, for example, the reflection 104b of the object 104a that persists across frames and remains sharp in the averaged frame. Thus, resulting in a set of detections representing potential privacy-risk reflections located within the windshield region of the media sequence.
[0053] In an embodiment, in response to the detection, a privacy filter tag associated with the static object is generated. The privacy filter tag may be a structured metadata element describing both the presence of sensitive content and the corrective or risk mitigating action to be applied. Each privacy filter tag indicates that the static object is subject to privacy filtering and may contain executable or interpretable instructions specifying the type of filtering.
[0054] In an example, the privacy filter tag may include removal instructions to exercise or replace the reflection region across the original media sequence.
[0055] In an example, the privacy filter tag may include blurring instructions to obscure the reflection using segmentation-based regions of interest (ROIs) or Al-generated masks, ensuring that underlying details remain illegible.
[0056] In an example, the privacy filter tag may include blacklist instructions to temporarily flag the image capturing device 108 so that subsequent imagery from that device is excluded from sharing for one or more specified data-use contexts.
[0057] The privacy filter tag may also include ancillary metadata such as timestamp, location, detection confidence, and object type (document, face, logo, or other). Downstream privacy -filtering systems or cloud-based aggregation servers can use this metadata to apply appropriate enforcement actions. For example, in a mapping application, a high-confidence reflection detection may trigger automatic exclusion of the corresponding video segment; in a drivercoaching application, the same detection may trigger selective blurring or cropping, preserving the rest of the content.
[0058] When a blacklist instruction is issued, the system records the event in a policy database. A blacklist manager enforces temporary suspension of data contribution from that particular camera or vehicle for the relevant application domain such as mapping or fleet analysis. This isolation prevents repeated privacy violations from the same source without interrupting other benign data uses.
[0059] The privacy filter tag, therefore, becomes a unifying control element connecting reflection detection with policy enforcement. In an example, the privacy filter tag may be stored locally with the camera 108 or transmitted via a data-sharing interface to centralized remote servers.
[0060] In an embodiment, the ML model training techniques may be employed to improve detection accuracy for the averaged frame rather than for a single frame. In an example, training data is generated by labeling averaged frames that either contain or do not contain reflections. Machine-learning models such as convolutional neural networks (CNN) trained on such examples learn to recognize reflection patterns distinct from motion-blur artifacts.
[0061] In an example, the pre-trained ML models (OCR, face detectors) may be combined with specialized reflection classifiers trained on averaged data. Further, fine-tuning the ML models using feedback from human reviewers may advantageously allow the detection of the reflection to evolve as more driving data is processed.
[0062] In an embodiment, expected visual features that might remain partially visible in averaged frames, such as lane markings, telephone poles, or other vehicles traveling at similar speeds, are also detected. The ML models (image processing models), by classifying the external objects into non-sensitive categories, prevent them from being misinterpreted as reflections. Furthermore, filtering rules automatically discard detections associated with known non-sensitive object classes.
[0063] In an embodiment, selection of appropriate image processing models may be based on a consideration of the scale of data consumption. Detectable objects that do not contain personal information may become a privacy risk on a large scale. For example, in the context of large-scale data consumption of anonymized imagery, it may be important to detect elements that are visible 14in windshield reflections (or outside of the windshield on the hood, etc.). Such elements may be a privacy risk by virtue of enabling a downstream data consumption process to link geolocated imagery collected at different times to the same vehicle or person. As with models that are capable of detecting faces, text, and / or documents, the processing applied to the averaged image can be augmented to detect other visual elements. Additional relevant elements may include visible air vent elements that are visible in the reflection, the cropping of which may be unique to a particular vehicle. Likewise, a visible hood or portions of a side-view mirror may be elements that are discernible in the averaged image by virtue of appearing at a fixed location relative to the camera 108.
[0064] For elements that are not themselves privacy risks, but which may become so at very large scale of data collection, privacy filtering may result in cropping of the shared imagery. For example, a small random portion of the lower portion of each shared image may be cropped and / or the left or right side of the image, so these elements from a single vehicle appear differently in different images shared from that vehicle.
[0065] The example implementation of the present invention enables a transformed role for human reviewers. In one example scenario, short video segments may be shared with a data consumer who is trying to solve the problem of timely updates to mapping data. In one approach, every short video segment that is shared may be first reviewed by a human labeler.
[0066] In a system that incorporates the aforementioned systems and methods, the role of human labelers may be transformed and / or simplified. In one example, the role of human labelers will be focused on two areas, (a) selective audits, and (b) reviewing examples for which the automated processing disclosed herein yields inconsistent or low confidence results.
[0067] Selective auditing may be applied to a portion, such as 1%, of shared imagery, for monitoring and improvement purposes. These selective audits will aid understanding of model performance for privacy filtering and in relation to other objectives and metrics, on a large and diverse corpus of real-world data.
[0068] Human review of inconsistent and / or low confidence results may relate to particular model outputs generated by the ML models that are utilized to process the averaged frames. If the model does not yield confident enough predictions, such as by inferring that the averaged image contains a blurry document, but with a self-report of a confidence threshold that is low (e.g. 60%), then such imagery may be put into a queue for human review.
[0069] The confidence threshold that is utilized may vary based on the potential sensitivity of the object type. For example, for a face detection algorithm, the confidence threshold may not be high enough (may not exceed the threshold), if there is only a 10% certainty that there is a face in the image. Based on these, one or more thresholds, single images or short video clips may be put into a queue for human review.
[0070] In certain embodiments, there may be multiple different filtering steps, and further, there may be an interaction between privacy filtering steps. For example, low confidence images or video clips may be chosen for human review because the image or short video clip met the parameters of the data consumption task, and additionally, there was not a replacement available (e.g. from the same location and pointing in the same direction). There may be one or more queues of images or short videos for human review. Elements from these queues may be removed prior to being reviewed by a human. For example, if there are 1000 images to review, there might be 800 of these for which a distributed system of image search may be able to source another image of the same location, and for which the automated processing revealed no privacy risk problems. If the data consumption task may be satisfied with these alternative images, then the images in the queue that they replace do not need to be reviewed. These images can then be deleted from this portion of a pipeline, thereby focusing human review resources on the remaining images in the queue for which no replacement has been found.
[0071] Similarly, for a camera / vehicle for which imagery is in a queue for human review, subsequently detected imagery may yield a repeated detection of a privacy risk, but with a higher confidence score. The subsequent higher confidence score can be applied to imagery collected from the same vehicle for a time-period surrounding the event, both before and after. Images in a human review queue may be identified as being associated with a vehicle / camera that is temporarily excluded, and such imagery may be removed on the basis that a human reviewer would likely confirm that the imagery poses a privacy risk.
[0072] Imagery that is removed from a human labeling queue, wherein the queue contains imagery having a low confidence score, may still be entered into a different human labelling queue that is directed to creating training data for the image processing algorithms and models. For images that were processed to yield a low confidence score, but for which human reviewers would confidently label the images as containing a sensitive reflection or not, such images may be labelled and added as training data to improve the image processing algorithms and models.
[0073] The above descriptions of labelling queues may be specifically applied to eliminating or significantly reducing privacy risks from windshield reflections in the context of sharing imagery from certain locations in an anonymized fashion, avoiding privacy risks. The disclosed systems and methods may be applied in more contexts, however. For example, there may be contexts and applications for which imagery is collected and shared with the driver and / or the fleet that employs or contracts with the driver. In one example, these fleets may be interested in coaching their drivers by reviewing video recordings of instances of safe or unsafe driving with each driver. In this context, the driver and fleet may prefer to jointly review video imagery that does not contain extraneous personal details, as may be reflected by documents or objects inside the vehicle 102.
[0074] In a driver safety application, such as one for which imagery may be shared with a driver’s safety manager and used to coach the driver on safer driving practices, it may be considered privacy-enhancing if coachable examples could be selected for which there is no extraneous PII or other sensitive information regarding the interior of the vehicle made visible to an outwardfacing camera.
[0075] A fleet using a camera-based driver safety system enabled with certain aspects of the present invention may be interested in obtaining additional outward-facing camera imagery that is collected from vehicles in its fleet but not otherwise associated with a detected safety event.
[0076] The fleet may attempt to collect this additional outward-facing camera imagery while still minimizing the sharing of personal information regarding the driver of the vehicle. In one example, the fleet may provide infrastructure equipment that is deployed near roadways. The fleet may want to collect imagery relating to these infrastructure assets so that they may review the state of their infrastructure. In this scenario, the image / video processing and privacy flagging techniques may be employed so that, for example, video for which a reflection of a driver’s medical records laying on the dashboard is visible may be automatically avoided or ignored.
[0077] Another potential application may make use of detected reflections to enforce a company policy, which may include a prohibition on leaving loose items on the front dashboard while the vehicle 102 is moving. Thus, based on the detection of objects 104a in the windshield 106 reflection, the safety manager might initiate a conversation with a driver to remind the driver that leaving certain things on the dashboard may be unsafe, and / or against the policy of the company. Such an application may involve computing a percentage of time that an object is visible in the windshield as a proportion of driving time when lighting conditions make visible windshield reflections more likely, such as when the vehicle is traveling in the direction of the sun, or when17driving at night with interior lighting in the cab, or when an empty dashboard or dashboard features are visible.
[0078] In an embodiment, the averaged frame is processed with a chatbot that is capable of describing images or videos. For example, the chatbot may be prompted with: “describe the lower part of the image.” There may then be a number of responses which may be considered a detection of a sensitive reflection. For example, the chatbot may reply that: “There is a document or paper attached to the inside of the windshield, possibly containing some form of information, like instructions, or data sheets.” Or: “It is a document or paper attached inside of the windshield.” In such examples, the chatbot might actually incorrectly describe the image. For example, the document or paper visible in the windshield reflection may not be “attached” to the inside of the windshield. However, because the response indicates that a document or paper is visible, given that the chatbot was describing an averaged image, it can be inferred that a document that is detectable in the averaged image must be a windshield reflection and / or that it should be flagged for future privacy filtering.
[0079] A chatbot interface, or other general-purpose visual detection model, may be suited to detect a logo of the vehicle’s hood. Because the vehicle hood will also appear crisp in an averaged frame, and further, may be capable of associating the frame with a particular vehicle, privacy filtering may likewise be applied to the frames for which there is a visible hood. In some embodiments, the privacy filtering may be more thorough (more likely to be replaced, discarded) for the frames having visible hoods if there is a recognizable logo or other marking on the hood.
[0080] In an embodiment, the chatbot or similar image-to-language model may be fine-tuned or prompt engineered. For example, the averaged frame may be presented to the image-to-language model with a prompt that explains, for example, “This is a blurry image created by averaging many images together that were captured by a camera attached to a moving vehicle. I want to know if there are any documents that are visible in this image, or any objects that may be considered sensitive to or relatable to the driver.” The chatbot interface may respond that there is a circular logo with words written on it. If the description of the logo is similar to a description of the logo of the company that owns the vehicles, then such imagery may be flagged so that it is not shared with other entities, but in such a way that it may still be used for purposes that are internal to the company that owns the vehicles, since this would not constitute new information in that latter context.
[0081] In an embodiment, for a given data sharing application, there may be one or more considerations that guide the number of frames included in the averaging process to yield an averaged frame. First, if the vehicle is not moving at all, then the averaging will not have the effect of blurring objects that are in the external environment. Accordingly, at least some of the frames that are selected for averaging should be captured while the vehicle is moving, as may be ascertained by a GPS sensor, inertial sensor, or data available on the CAN bus, such as vehicle speed or wheel rotation.
[0082] Second, it may be observed that windshield reflection intensity changes as a function of the direction that the camera is pointed relative to the sun. For example, where the driver is travelling on an off ramp, and making a circular path, the presence or intensity of reflections may change as the vehicle moves because the sunlight is coming from a different direction. In such a case, it may be helpful to consider multiple different time periods over which to apply image averaging and further processing. This processing may identify which vehicle orientations (relative to the sun) are problematic in a given location, weather condition, and / or time-of-day.
[0083] Third, it may be beneficial to average over a relatively short period so that the privacy risks may be properly localized in time, if, for example, the drivers frequently move their documents and other personal materials, which may appear as the windshield reflection.
[0084] On the other hand, it may be helpful to average over longer periods of time or a number of images, such as over one minute at 30 frames per second. With longer averaging, items that are visible in the external environment may become more and more blurred, leaving the reflected objects relatively clear and distinguishable.
[0085] FIGURE 2 illustrates a block diagram of a system 200, in accordance with an embodiment of the present invention.
[0086] As shown in FIGURE 2, the system 200 includes the image capturing device 108. The image capturing device 108 may include a bus 201, a processing unit 202, a memory unit 204, a communication unit 206, an I / O interface 208, and an output unit 210. In another embodiment, the system 200 implemented at least in the image capturing device 108 acts as a virtual machine running on host hardware.
[0087] The processing unit 202 may include one or more processors as a single processing unit or several units. The processing unit 202 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state 19machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. Among other capabilities, the one or more processors are configured to fetch and execute computer-readable instructions and data stored in the memory unit 204.
[0088] The memory unit 204 includes one or more computer-readable storage media that can communicate via the bus 201. The memory unit 204 may include non-volatile storage elements. Examples of such non-volatile storage elements may include magnetic hard discs, optical discs, floppy discs, flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories. In addition, the memory may, in some examples, be considered a non-transitory storage medium. The term “non-transitory” may indicate that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term “non-transitory” should not be interpreted to mean that the memory is nonmovable. In certain examples, a non-transitory storage medium may store data that can, over time, change (e.g., in Random Access Memory (RAM) or cache.
[0089] The memory unit 204 may further include any non-transitory computer-readable medium known in the art including, for example, volatile memory, such as static random-access memory (SRAM) and dynamic random-access memory (DRAM), and / or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes.
[0090] The memory unit 204 includes a database 212 and modules 214. The database 212 is configured to be accessed by the processing unit 202 and stores information as required by processing unit 202 to perform the one or more functions. The database 212 may store the media sequence of the ML models. The database 212 may also store historical data including information of past key scores, value scores, and the user interactions. Additionally, the database 212 may serve dual functions in storing different types of information.
[0091] The modules 214, amongst other things, include routines, programs, objects, components, data structures, etc., which perform particular tasks or implement data types. The modules 214 may also be implemented as, signal processor(s), state machine(s), logic circuitries, and / or any other device or component that manipulates signals based on operational instructions.
[0092] Further, the modules 214 can be implemented in hardware, instructions executed by the processing unit 202, or by a combination thereof. The processing unit 202 can comprise a computer, a processor, a state machine, a logic array, or any other suitable devices capable ofprocessing instructions. The processing unit 202 can be a general-purpose processor which executes instructions to cause the general-purpose processor to perform the required tasks or, the processing unit can be dedicated to performing the required functions. In another embodiment of the present disclosure, the modules 214 may be machine-readable instructions (software) which, when executed by the processing unit 202, perform any of the described functionalities / methods, as discussed throughout the present disclosure.
[0093] Furthermore, the modules 214 may be implemented through the Al model. A function associated with Al may be performed through the non-volatile memory, the volatile memory, and the processing unit 202.
[0094] The modules 214 may store the one or more AI / ML models with learning capabilities. Each Al model may include a plurality of neural network layers. Examples of neural networks include, but are not limited to, CNN, DNN, RNN, and RBM. The learning technique for training each Al model uses learning data to cause, allow, or control the system 200 to make a determination or analysis. Examples of learning techniques include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. At least one of a plurality of CNN, DNN, RNN, RMB models and the like may be implemented to thereby achieve execution of the mechanism, of the present disclosure, through the Al models. A function associated with the AI / ML module may be performed through the non-volatile memory, the volatile memory, and the processing unit 202.
[0095] The communication unit 206 is configured to communicate sensor data, or any other content over a communication network via a communication port or interface or using the bus 201. Further, the communication unit 206 may include a communication port or a communication interface for sending and receiving simulation datasets via the communication network. The communication port or the communication interface may be a part of the processing unit 202 or may be a separate component. The communication port may be created in software or may be a physical connection in hardware. The communication port may be configured to connect with the communication network, external media, the display, or any other components in the system 200, or combinations thereof. The connection with the communication network may be a physical connection, such as a wired Ethernet connection, or may be established wirelessly as discussed above. Likewise, the additional connections with other components of the system 200 may be physical or may be established wirelessly. The communication unit 206 may include the Wi-Fi21module or Bluetooth module for enabling wireless communication capability and data exchange capability between various modules of the system 200.
[0096] The I / O interface 208 refers to hardware or software components that enable communication between various modules of the system 200. The I / O interface 208 serves as a communication medium for exchanging information, commands, signals, or query responses with other devices or systems. The I / O interface 208 may be a part of the processing unit 202 or maybe a separate component. The I / O interface 208 may be created in software or maybe a physical connection in hardware. The I / O interface 208 may be configured to connect with an external network, external media, the display, or any other components, or combinations thereof. The external network may be a physical connection, such as a wired Ethernet connection, or may be established wirelessly.
[0097] The output device 210 which is preferably an output unit comprises of a display device. The display device may be an Augmented Reality / Virtual Reality (AR / VR) device to display a virtual environment to the user. The display device may include a display screen. As a non-limiting example, the display screen may be Light Emitting Diode (LED), Liquid Crystal Display (LCD), Organic Light Emitting Diode (OLED), Active Matrix Organic Light Emitting Diode (AMOLED), or Super Active Matrix Organic Light Emitting Diode (AMOLED) screen. The display screen may be of varied resolutions. The remote device, while operating as the output unit, is further configured for presenting the selected frames.
[0098] In an embodiment, the present invention also contemplates a computer-program product, having machine-readable instructions stored therein, when executed by the processing unit 202, which causes the processing unit 202 to perform the method for detecting the windshield reflections 104b of the object 104a in the vehicle 102 using the image capturing device 108 installed inside the vehicle 102. The details on the method(s) performed by the processing unit 202 have been elaborated in subsequent paragraphs at least with reference to FIGURE 3.
[0099] Further, the present invention also contemplates a non-transitory computer-readable medium encoded with executable instructions. The executable instructions, when executed by the processing unit 202, cause the processing unit 202 to perform a method detecting the windshield reflections 104b of the object 104a in the vehicle 102 using the image capturing device 108 installed inside the vehicle 102. The details on the method(s) performed by the processing unit 202 have been elaborated in subsequent paragraphs at least with reference to FIGURE 3.
[0100] FIGURE 3 illustrates a functional block diagram of the system 200 configured for detecting the windshield reflections 104b of the object 104a in a vehicle 102. The system includes the modules 214 i.e., a receiving module 302, an averaging module 304, a detecting module 306, and a generating module 308. Each module represents a logical or physical component implemented by the processing unit 202 executing instructions or operating through dedicated hardware circuitry. The modules 214 communicate with each other through data buses or shared memory within the system architecture.
[0101] In an embodiment, the receiving module 302 is configured to interface with the image capturing device 108 installed inside the vehicle 102 and oriented outward through the windshield 106. The image capturing device 108 provides a live video or image feed representing the media sequence of frames recorded while the vehicle 102 is in motion. The receiving module 302 captures and stores the media sequence (alternatively referred to as the images in the detailed description) for subsequent processing. The receiving module 302 is configured to also verify that the vehicle 102 is moving by reading signals from motion sensors, GPS, or the onboard diagnostic (OBD) system to ensure that external scenery changes between frames. As stationary capture would not produce the relative motion difference necessary to separate internal reflections from external features. The receiving module 302 is in communication with the averaging module 304.
[0102] In an embodiment, the averaging module 304 is configured to operate on the media sequence received by the receiving module 302. The averaging module 304 is configured to compute the averaged frame based on calculating, for each pixel location, an average of the pixel values across the frames within the media sequence. The averaging module 304 may employ temporal mean, weighted averaging, or adaptive variance-based averaging techniques. In an example, the averaging period may be determined by a predefined number of frames (for example, 100 frames) or a time window (for example, 17 seconds). Thus, the averaging images (i.e., the media sequence) captured while the vehicle 102 is moving, the averaging module 304 causes the dynamic external scenery to appear blurred, while stationary reflections of interior objects, such as the reflection 104b of the object 104a remain sharp and clearly visible.
[0103] In an embodiment, the averaging module 304 is configured to select a predefined number of frames or a predefined duration from the media sequence to compute the averaged frame. The averaging module 304 dynamically adjusts the averaging window based on vehicle speed and frame rate to maintain optimal separation between moving and stationary elements. For instance, at higher speeds, fewer frames may be sufficient to induce external blur, while at slower speeds,23more frames are averaged. The resulting averaged frame presents blurred external scenery relative to static interior objects. Advantageously, improving robustness under varying driving conditions and ensuring consistent reflection detection performance. Further, the receiving module 302 is configured to determine capture coverage based on vehicle movement metrics, for example, a predefined travel distance or frame count. Capturing only while covering at least a specific distance ensures that the external scene changes sufficiently to produce reliable motion-induced blur. The advantage of distance-based control is that it provides consistent image diversity independent of frame rate or travel duration, enhancing the effectiveness of subsequent averaging.
[0104] The receiving module 302 and the averaging module 304 are in communication with the detecting module 306.
[0105] In an embodiment, the averaged frame generated by the averaging module 304 is transmitted to the detecting module 306, which identifies the static objects within the averaged frame. The detecting module 306 is configured to process the averaged frame using the combination of CV models and pattern-recognition algorithms to identify reflections within the windshield 106. The detecting module 306 is configured to apply the pre-processing steps, such as the grayscale conversion, the normalization, and the gradient-field analysis, to emphasize sharp regions that are indicative of the reflection 104b. The detecting module 306 may employ multiple specialized ML models, including the optical character recognition (OCR) for text, face detection for human imagery, document-detection networks for rectangular paper-like structures, and general object-detection neural networks for other reflective artifacts.
[0106] In an example, the grayscale gradient-field is applied to pre-process the averaged frame before applying reflection classifiers. Sharp gradient transitions highlight reflection contours. The detecting module 306 then targets its classifiers specifically on these sharp regions, optimizing computation and improving accuracy. For instance, by focusing only on high-gradient regions, the system 200 avoids analyzing blurred external scenery, reducing false detections and processing time.
[0107] The detecting module 306 thus distinguishes between the reflections originating from interior static objects and blurred visual content representing external scenery. The detecting module 306 analyzes sharpness, spatial stability, and geometric consistency across the frames to confirm that a region corresponds to the reflection 104b. For example, if the same contour appears stationary across multiple frames while the background scene shifts, the module flags it as a static reflection. Advantageously, ensuring that the detected object is truly internal to the vehicle 102 24and not an external static feature like a roadside pole or sign. The receiving module 302, the averaging module 304, and the detecting module 306 are in communication with the generating module 308.
[0108] In an embodiment, the generating module 308 is configured to generate the privacy filter tag based on detections received from the detecting module 306. The privacy filter tag indicates that the detected static object is subject to privacy filtering and includes one or more executable instructions specifying treatment of the corresponding content (i.e., sensitive content). Depending on configuration or policy rules, the generating module 308 is configured to insert instructions into the privacy filter tag for removing, blurring, or blacklisting imagery associated with the static object.
[0109] In an embodiment, the system 200 includes the image capturing device 108 and the processing unit 202 configured to execute the functionality of the present invention. The image capturing device 108 records the media sequence while the vehicle is in motion, and the processing unit 202, via the averaging module 304 and the detecting module 306, generates the averaged frame and detects the reflection 104b. The generating module 308 produces the privacy filter tag containing instructions to either remove, blur, or blacklist the imagery.
[0110] For example, in a fleet vehicle equipped with multiple cameras, each camera continuously streams image data to the processing unit 202. When the system 200 detects a dashboard document reflection in the windshield 106, the privacy filter tag may command immediate blurring in the recorded frames before uploading them to a central server. Thus, in an advantageous aspect, ensuring that sensitive reflections are never propagated outside the vehicle 102 without appropriate anonymization.[OHl] In an embodiment, the detecting module 306 is configured to determine that the static object corresponds to the sensitive content which may include a document, text, logo, or human face by processing the averaged frame and assessing the discemability or legibility of such content. The term “discernability” refers to the ability to identify facial or shape features, while “legibility” pertains to readable text or symbols. The detecting module 306 thus evaluates clarity metrics, such as edge contrast or text recognition confidence, to confirm the sensitivity of the detected reflection.
[0112] In an embodiment, the detecting module 306 uses the OCR to identify readable text. For instance, if the reflection 104b shows a printed address label on the document (i.e., the object 104a), the OCR engine recognizes alphabetic or numeric characters. The discerned charactersconfirm that the reflection represents sensitive textual information. Thereby, allowing the generating module 308 to prioritize the removal of the sensitive content. In an advantageous aspect, this leads to early, automatic recognition of textual reflections that may contain PII before any human review.
[0113] In an embodiment, the detecting module 306 may be configured to apply face detection algorithms to locate human facial patterns within the averaged frame. Faces may appear as reflections, such as those of the driver or passengers, and are treated as privacy sensitive. The detecting module 306 is configured to measure facial feature confidence (e.g., eye and nose landmarks) even if the reflection is partial or mirrored. In an advantageous aspect, the human facial patterns detection capability ensures that faces are detected and filtered consistently.
[0114] In an embodiment, the detecting module 306 is further configured to apply text detection at any orientation. In an example, the reflection 104b on the windshield 106 may appear rotated, mirrored, or skewed relative to the camera 108. The text-detection ML model uses rotationinvariant convolutional kernels or transformer-based detectors to capture text regardless of its orientation. Advantageously, the system 200 identifies printed or handwritten information appearing diagonally or reversed on reflective glass, an improvement over ordinary horizontal OCR techniques.
[0115] In an embodiment, the detecting module 306 is configured to detect documents and identify paper-like structures even when text is not readable. For example, the detecting module 306 is configured to detect rectangular boundaries with uniform textures corresponding to forms, IDs, or placards placed on the dashboard. Even if the OCR fails due to lighting glare, the geometric cues of the document still trigger generation of the privacy filter tags.
[0116] In an embodiment, the detecting module 306 is configured to incorporate the objectdetection neural network capable of recognizing sensitive content such as corporate logos, brand markings, or device screens. For instance, if a smartphone display or laptop screen reflects in the windshield 106, the object-detection neural network classifies the pattern and marks it for filtering. Advantageously, the recognition of the sensitive content allows the system 200 to capture all classes of the sensitive content, not just text or faces, thereby achieving broad privacy protection.
[0117] In an embodiment, the processing unit 202, via the averaging module 304 and the detecting module 306 may adjust the averaged frame to enhance brightness and contrast, followed by26normalization prior to reflection detection. The image normalization may reduce variance between day and night imagery, allowing the detection module 306 to function with uniform thresholds.
[0118] In an example, during nighttime conditions, the processing unit 202 further enhances brightness and contrast adaptively using gain control or exposure fusion. Consequently, low- visibility reflections such as dashboard lights or illuminated papers may remain detectable even under dark conditions.
[0119] In an embodiment, the detecting module 306 is configured to generate a segmented frame of the averaged frame, dividing it into semantic regions such as windshield, dashboard, sky, and road. The detecting module 306 is configured to then apply the object-detection neural network to the semantic regions to identify the static reflections. Thus, based on combining segmentation with neural inference outputs, the system 200 distinguishes true reflective artifacts from false positives. For example, a segmentation label identifying the “windshield” region limits detection scope, reducing computation and improving precision. Advantageously, the combined approach ensures that the detected static objects are accurately located and contextually relevant.
[0120] In an embodiment, the generating module 308 associates each detected static object with the removal instruction in the privacy filter tag. The removal instruction directs subsequent processing to delete or replace the reflection region across all frames in the media sequence. The removal instructions may include inpainting or frame interpolation to fill the removed region using nearby pixels. For example, a document reflection can be seamlessly replaced by an interpolated region representing the external background. In an advantageous aspect, the complete elimination of the sensitive content does not visually impair the usefulness of the remaining data.
[0121] In an embodiment, the generating module 308 is configured to associate a blurring instruction with the privacy filter tag. The blurring instruction applies Gaussian, median, or AI- based segmentation blurs to the specific region where the reflection 104b was detected in the frames. The segmentation ensures that only the reflection region is blurred while the surrounding visual context remains intact. The blurred imagery conceals personal information while maintaining contextual continuity for analytic algorithms such as lane or object detection that rely on the external scene.
[0122] In an embodiment, the generating module 308 is configured to generate the privacy filter tag, including a temporary blacklist state for the image capturing device 108. Upon activation, the temporary blacklist state restricts or halts data transmission from the camera 108 for certainspecified data uses. For instance, the system 200 may permit internal storage but block data sharing with mapping servers. Advantageously, the temporary blacklist state helps prevent repeated privacy violations from persistent reflections, such as an always-visible document taped to the dashboard, without disabling the entire camera permanently.
[0123] In an example, if a vehicle participates in crowd-sourced road mapping, the system 200 temporarily excludes that camera’s imagery from upload to the mapping server until the reflection issue is resolved. Preventing private dashboard content from being embedded in public map datasets.
[0124] In another example, data-use context is coaching with video, wherein fleet operators analyze driver behavior. If a camera repeatedly records driver reflections, the privacy filter tag blacklists its output from the coaching data feed. This targeted restriction mitigates privacy risks without affecting other non-sensitive analytics, such as telematics or mechanical diagnostics.
[0125] In an advantageous aspect, the system 200 provides flexibility and scalability. The module 214 can be implemented as firmware, software, or hardware, allowing deployment on embedded edge processors within the vehicle, the camera 108, or on cloud-based servers receiving uploaded imagery.
[0126] FIGURE 4 illustrates a flow chart of a method 400 for detecting the windshield reflections 104b of the object 104a in the vehicle 102 using the image capturing device 108 installed inside the vehicle 102, in accordance with an embodiment of the present invention. The method 400 includes a series of operation steps 402 through 408 performed by the processing unit 202 of the system 200.
[0127] The various actions, acts, blocks, steps, or the like in the flow diagrams may be performed in the order presented, in a different order, or simultaneously. Further, in some embodiments, some of the actions, acts, blocks, steps, or the like may be omitted, added, modified, skipped, or the like without departing from the scope of the invention.
[0128] At step 402, the method 400 includes capturing the media sequence from the image capturing device 108 while the vehicle 102 is in motion.
[0129] At step 404, the method 400 includes generating the averaged frame based on computing, for each pixel location, the average of pixel values across the plurality of frames of the media sequence. The plurality of frames is captured while the vehicle 102 is in motion.
[0130] At step 406, the method 400 includes detecting the static object in the vehicle 102 that is reflected in the windshield 106 of the vehicle 102 based on processing the averaged frame.
[0131] At step 408, the method 400 includes generating the privacy filter tag in response to detecting the static object. The privacy filter tag indicates that the static object is subject to privacy filtering and includes privacy risk mitigating instructions which may be to remove the static object from the media sequence, blur the static object in the media sequence, and blacklist the image capturing device 108.
[0132] FIGURES 5-7 illustrates example use-cases depicting the detection of the windshield reflections of the object in the vehicle, in accordance with an embodiment of the present invention.
[0133] The images depicted in FIGURE 5 illustrate various real-world examples of the windshield reflections that may be captured by the camera 108 in the vehicle 102. Each of the examples in FIGURE 5 are based images that are produced by averaging pixel values over many frames. In averaged image 502, a portion of an identification card (having an outline of a head) and two pens are visible as windshield reflections. While the pens would generally not be considered a privacy risk, the identification card could be a serious privacy risk if the reflection were larger and / or lighting conditions made the text legible or face discernible. A document that appears to be a work order is visible as a reflection in the bottom center of averaged image 504. A clipboard with a business document containing printed and handwritten text occupies over half of averaged image 506. A spiral notebook is visible as a windshield reflection in averaged image 508. A handicap parking placard is visible at the bottom of averaged image 510. A booklet and various other documents are visible as reflections in averaged image 512.
[0134] As illustrated in FIGURE 6, when viewing single images, it is often difficult to notice the windshield reflections. A first image frame 602 depicts a road scene captured by a forward-facing camera mounted inside the vehicle 102. There is a speed limit sign visible in the distance on the right side of the road. Immediately in front of the vehicle is a complex pattern of shadows from roadside trees. There are also reflections of features of the dashboard, but they are difficult to discern relative to the complex pattern of shadows on the road.
[0135] The windshield reflections are more easily apparent in the second image frame 604, which was captured shortly after the first image frame 602. Several features of the dashboard and the outlines of a box are visible as windshield reflections in the second image frame 604. The third image frame 606 was captured moments later at a time that the vehicle is within a large shadow29cast by a building. Within this large shadow, the windshield reflections of the dashboard elements are faint, but the reflections of the box are somewhat clear, perhaps due to the lighting conditions caused by the large shadow. Finally, the windshield reflections are again difficult to discern in a fourth image frame 608, captured a few seconds later. In the fourth image frame 608, the reflection of the box interacts with road markings of a crosswalk on a side street.
[0136] Given that it is often difficult for humans to see windshield reflections in single image frames, processing of single image frames by automated techniques to detect windshield reflections, such as by application of deep learning techniques to detect a visible reflection in a windshield, cannot be relied upon to yield adequate detection results to robustly ensure anonymization of shared imagery from vehicle-mounted cameras.
[0137] In accordance with certain embodiments of the present disclosure, the vehicle 102 may be passing a traffic sign. The series of images depicted in FIGURE 6 (through 602-610) illustrates a vehicle passing a speed limit sign. The traffic sign may be one for which anonymized mapping imagery is requested. As can be seen in averaged image 610, the traffic sign itself will not be discernible in the averaged image. When the averaged image is computed over many frames, each frame captured while the vehicle is moving, there may be many signs that were passed by the vehicle. It is expected that all signs will be undetectable in the motion blur that results from the averaging across frames. Likewise, text that may be incorporated into the traffic signs are not expected to be legible in the averaged image, due to the motion of the vehicle during the time that images to be used for averaging were captured.
[0138] The single image 606 may be a target image for a mapping application, where the mapping application is concerned with updating a map of speed limit signs. As explained above, the windshield reflection may be difficult to discern in image 606. However, the images 604 and 608, as well as averaged image 610 reveal that there is a windshield reflection discernible around the time of the speed limit sign detection, as depicted in image 606.
[0139] FIGURE 7 depicts examples of averaged images via images 702, 704, and 706 for which vehicle hoods may be discernible, in accordance with certain aspects of the present disclosure.
[0140] Each averaged image shown in FIGURE 7 is generated by capturing the media sequence i.e., the plurality of consecutive frames while the vehicle 102 is in motion, and then computing, for each pixel location, the average of the pixel values across those frames. This averaging process causes external scenery such as passing vehicles, trees, or signs to become blurred because thoseelements move relative to the camera between frames. In contrast, items that are stationary with respect to the vehicle such as the vehicle hood, dashboard, or reflections of internal objects on the windshield remain sharp and discernible in the averaged result.
[0141] In averaged image 702, the outline of the vehicle hood is sharply defined along in the image. Since the vehicle hood occupies a constant position relative to the camera throughout the capture period, its visual features persist through the averaging, resulting in a clearly visible static region. This demonstrates one of the characteristic effects of the system’s 200 averaging technique i.e., static or semi-static visual elements attached to the vehicle remain identifiable, whereas the external environment becomes blurred. The image 702 thus visually separates internal or vehicle- fixed components from the outside world.
[0142] In averaged image 704, both the hood and a portion of a side-view mirror appear distinctly while the background is significantly blurred. Such examples illustrate that reflections or surfaces belonging to the vehicle itself can also appear as persistent, high-contrast areas in the averaged frame. Even though these regions do not typically reveal personal information, they can, at scale, function as vehicle identifiers for example, by exposing distinctive shapes, colors, or manufacturer logos. The system 200 allows such regions to be recognized and tagged for privacy filtering before the imagery is shared externally.
[0143] In averaged image 706, the hood surface may include reflective highlights or logos that identify the vehicle brand. These markings, although part of the vehicle’s exterior, constitute a privacy consideration when imagery from multiple vehicles is aggregated for mapping or analytics. When the logo or other text on the hood becomes legible in the averaged image, the system 200 designates that area as the sensitive content and applies filtering operations such as blurring, masking, or cropping. In instances where the static element is non-textual e.g., a shape or contour distinctive to a specific vehicle model, the filtering process introduces small random offsets or crops so that the same vehicle is not visually traceable across different datasets.
[0144] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one ordinary skilled in the art to which this invention belongs. The system, methods, and examples provided herein are illustrative only and not intended to be limiting.
[0145] While specific language has been used to describe the present subject matter, any limitations arising on account thereto, are not intended. As would be apparent to a person in theart, various working modifications may be made to the method in order to implement the inventive concept as taught herein. The drawings and the foregoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment.
[0146] The foregoing description of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and / or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of preferred embodiments, those skilled in the art will recognize that the embodiments herein can be practised with modification within the scope of the embodiments as described herein.
Claims
CLAIMSWhat is claimed is:
1. A method for detecting windshield reflections of an object in a vehicle using an image capturing device installed inside the vehicle, the method comprising: capturing a media sequence from the image capturing device while the vehicle is in motion; generating an averaged frame based on computing, for each pixel location, an average of pixel values across a plurality of frames of the media sequence, wherein the plurality of frames is captured while the vehicle is in motion; detecting a static object in the vehicle that is reflected in a windshield of the vehicle based on processing the averaged frame; and generating at least one privacy filter tag in response to detecting the static object, wherein the at least one privacy filter tag indicates that the static obj ect is subj ect to privacy filtering.
2. The method of claim 1, further comprising: determining the static object as sensitive content based on processing the averaged frame, the sensitive content comprising at least one of a document, text, a logo, or a human face; and detecting the static object based on discernability or legibility of the sensitive content in the averaged frame while the visual features of external scenery are blurred in the averaged frame.
3. The method of claim 2, wherein processing the averaged frame comprises applying optical character recognition for determining the sensitive content; and detecting the static object based on legibility of the text in the averaged frame.
4. The method of claim 2, wherein processing the averaged frame comprises applying face detection for determining the sensitive content; and detecting the static object based on discernability of the face in the averaged frame.
5. The method of claim 2, wherein processing the averaged frame comprises applying text detection at any orientation for determining the sensitive content; and detecting the static object based on legibility of the text in the averaged frame.
336. The method of claim 2, wherein processing the averaged frame comprises applying document detection for determining the sensitive content; and detecting the static object based on legibility of document text in the averaged frame.
7. The method of claim 2, wherein processing the averaged frame comprises applying an object-detection neural network for determining the sensitive content; and detecting the static object based on discernability of the face in the averaged frame.
8. The method of claim 1, wherein generating the at least one privacy filter tag comprises: associating the privacy filter tag with a removal instructions indicating removal of the static object from the plurality of frames of the media sequence.
9. The method of claim 1, wherein generating the at least one privacy filter tag comprises: associating the privacy filter tag with a blurring instructions to blur the static object in one or more frames of the media sequence using segmentation regions of interest.
10. The method of claim 1, wherein generating the at least one privacy filter tag comprises: associating the privacy filter tag with a temporary blacklist state for a specified data use; and preventing repeated capture or sharing by the image capturing device for the specified data use in response to the detection.
11. The method of claim 10, wherein the specified data use is a mapping use.
12. The method of claim 10, wherein the specified data use is coaching with video.
13. The method of claim 1, wherein generating the averaged frame comprises: selecting a predefined number of frames or a predefined duration of the media sequence for averaging; and generating an averaged frame based on averaging the selection, the averaged frame comprises visual features of external scenery blurred relative to static objects in the vehicle.
14. The method of claim 1, wherein recording the plurality of frames comprises:34covering a vehicle movement of a predefined distance or a predefined frames count; and capturing the media sequence based on the covered vehicle movement.
15. The method of claim 1, further comprising: adjusting pixel intensity of the averaged frame for brightness and contrast; and normalizing the averaged frame prior to detecting the static object.
16. The method of claim 15, wherein adjusting the pixel intensity comprises: enhancing brightness and contrast of the averaged frame during night-time conditions; and normalizing the averaged frame prior to object detection during the night-time conditions.
17. The method of claim 1, wherein detecting the static object comprises: generating a segmented frame for the averaged frame, wherein the segmented frame indicates a division into semantic regions; applying an object-detection neural network to the averaged frame; and detecting the static object based on the segmented frame and outputs of the objectdetection neural network.
18. The method of claim 1, further comprising: pre-processing the averaged frame with grayscale gradient field analysis; and detecting the static object based on targeting a classifier on sharp regions of the pre- processed averaged frame.
19. A system for detecting windshield reflections of an object in a vehicle, the system comprising: an image capturing device installed inside the vehicle; and a processing unit in communication with the image capturing device configured to: capture a media sequence from the image capturing device while the vehicle is in motion;generate an averaged frame based on computing, for each pixel location, an average of pixel values across a plurality of frames of the media sequence, wherein the plurality of frames is captured while the vehicle is in motion; detect a static object in the vehicle that is reflected in a windshield of the vehicle based on processing the averaged frame; and generate at least one privacy filter tag in response to detecting the static object, wherein the at least one privacy filter tag indicates that the static object is subject to privacy filtering.
20. The system of claim 19, wherein the processing unit is further configured to: determine the static object as sensitive content based on processing the averaged frame, the sensitive content comprising at least one of a document, text, a logo, or a human face; and detect the static object based on discernability or legibility of the sensitive content in the averaged frame while the visual features of external scenery are blurred in the averaged frame.
21. The system of claim 20, wherein to process the averaged frame the processing unit is configured to: apply optical character recognition for determining the sensitive content; and detect the static object based on legibility of the text in the averaged frame.
22. The system of claim 20, wherein to process the averaged frame the processing unit is configured to: apply face detection for determining the sensitive content; and detect the static object based on discemability of the face in the averaged frame.
23. The system of claim 20, wherein to process the averaged frame the processing unit is configured to: apply text detection at any orientation for determining the sensitive content; and detect the static object based on legibility of the text in the averaged frame.
24. The system of claim 20, wherein to process the averaged frame the processing unit is configured to:apply document detection for determining the sensitive content; and detect the static object based on legibility of document text in the averaged frame.
25. The system of claim 20, wherein to process the averaged frame the processing unit is configured to: apply an object-detection neural network for determining the sensitive content; and detect the static object based on discemability of the face in the averaged frame.
26. The system of claim 19, wherein to generate the at least one privacy filter tag the processing unit is configured to: associate the privacy filter tag with a removal instructions indicating removal of the static object from the plurality of frames of the media sequence.
27. The system of claim 19, wherein to generate the at least one privacy filter tag the processing unit is configured to: associate the privacy filter tag with a blurring instructions to blur the static object in one or more frames of the media sequence using segmentation regions of interest.
28. The system of claim 19, wherein to generate the at least one privacy filter tag the processing unit is configured to: associate the privacy filter tag with a temporary blacklist state for a specified data use; and prevent repeated capture or sharing by the image capturing device for the specified data use in response to the detection.
29. The system of claim 28, wherein the specified data use is a mapping use.
30. The system of claim 28, wherein the specified data use is coaching with video.
31. The system of claim 19, wherein to generate the averaged frame the processing unit is configured to: select a predefined number of frames or a predefined duration of the media sequence for averaging; and37generate an averaged frame based on averaging the selection, the averaged frame comprises visual features of external scenery blurred relative to static objects in the vehicle.
32. The system of claim 19, wherein to record the plurality of frames the processing unit is configured to: cover a vehicle movement of a predefined distance or a predefined frames count; and capture the media sequence based on the covered vehicle movement.
33. The system of claim 19, wherein the processing unit is configured to: adjust pixel intensity of the averaged frame for brightness and contrast; and normalize the averaged frame prior to detecting the static object.
34. The system of claim 33, wherein to adjust the pixel intensity the processing unit is configured to: enhance brightness and contrast of the averaged frame during night-time conditions; and normalize the averaged frame prior to object detection during the night-time conditions.
35. The system of claim 19, wherein to detect the static object the processing unit is configured to: generate a segmented frame for the averaged frame, wherein the segmented frame indicates a division into semantic regions; apply an object-detection neural network to the averaged frame; and detect the static object based on the segmented frame and outputs of the objectdetection neural network.
36. The system of claim 19, wherein the processing unit is configured to: pre-process the averaged frame with grayscale gradient field analysis; and detect the static object based on targeting a classifier on sharp regions of the pre- processed averaged frame.