Self-diagnostic stereo camera system using semantic map and confidence map of stereo images

The self-diagnostic system for stereo cameras uses confidence and semantic maps to improve depth estimation accuracy by distinguishing between environmental and system-related errors, facilitating efficient and accurate diagnostics and corrections.

WO2025243248A1PCT designated stage Publication Date: 2025-11-27STEREOLABS SAS
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/055327
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-24
Filing Date
2025-05-22
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Stereo camera systems face challenges in maintaining consistent performance and reliability due to issues such as low texture areas, environmental factors, and physical obstructions, which affect depth estimation accuracy and are difficult to diagnose effectively using traditional error detection methods.

Method used

A self-diagnostic system that generates confidence, semantic, and knowledge maps using stereo camera images, leveraging artificial neural networks to identify and adjust confidence values based on semantic analysis, allowing for both system-wide and localized error detection and correction.

Benefits of technology

Enhances the reliability and accuracy of depth estimation by differentiating between environmental factors and system failures, providing nuanced diagnostics and enabling automated corrective actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025055327_27112025_PF_FP_ABST
    Figure IB2025055327_27112025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides methods for self-diagnostics and error detection in stereo camera systems. A confidence map is generated for a stereo camera image captured using at least two camera images. A semantic map is generated using a semantic detector and at least one camera image. A knowledge map is generated using the semantic and confidence maps. The confidence map may be adjusted using the semantic map. The methods include using the knowledge map and confidence map to conduct self-diagnostic actions. These actions may include generating an image area alert identifying a portion of the stereo camera image and a root cause of an error. The methods also include using the adjusted confidence map to conduct self-diagnostic actions. These actions may include discarding a portion of a depth map or point cloud or generating a system error alert indicating the entire stereo camera image is deemed insufficiently accurate.
Need to check novelty before this filing date? Find Prior Art

Description

SELF-DIAGNOSTIC STEREO CAMERA SYSTEM USING SEMANTIC MAP AND CONFIDENCE MAP OF STEREO IMAGESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 651,919, filed May 24, 2024.BACKGROUND

[0002] Stereo camera systems have become increasingly prevalent in various applications, including autonomous vehicles, robotics, and computer vision. These systems utilize two or more cameras to capture images from slightly different perspectives, enabling the generation of depth information and three- dimensional representations of scenes. By analyzing the disparities between corresponding points in the captured images, stereo vision algorithms can estimate distances to objects and create depth maps or point clouds.

[0003] As stereo camera technology has advanced, the accuracy and reliability of depth estimation have improved significantly. However, several challenges remain in ensuring the consistent performance and robustness of these systems across diverse environments and operating conditions. One such challenge is the presence of areas with low texture or high homogeneity in captured images, such as sky regions or uniform surfaces. These areas can lead to difficulties in establishing reliable correspondences between images, potentially resulting in inaccurate depth estimates.

[0004] Another issue faced by stereo camera systems is the susceptibility to various environmental factors and physical obstructions. Lens contamination, such as water droplets, dirt, or condensation, can significantly impact image quality and subsequently affect depth estimation accuracy. Similarly, partial or complete occlusions of camera lenses can lead to erroneous depth calculations or system failures. Detecting and diagnosing these issues in real-time is crucial for maintaining the reliability of stereo vision systems, particularly in safety-critical applications.

[0005] Furthermore, the complexity of stereo matching algorithms and the large amount of data processed in real-time pose challenges for implementing effective self-diagnostic capabilities. Traditional error detection methods often rely on simple thresholding techniques or predefined rules, which may not adequately capture the nuanced nature of potential issues in stereo vision systems. Additionally, these approaches may struggle to differentiate between genuine system malfunctions and challenging but valid scene conditions.

[0006] It has been appreciated that a system is needed that overcomes one or more of these problems.SUMMARY

[0007] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0008] The present disclosure provides methods for self-diagnostics and error detection in stereo camera systems. A confidence map may be generated for a stereo camera image captured using at least two camera images. The confidence map may be generated using approaches disclosed herein to represent a confidence level that a match has been achieved between a given point as represented in the at least two camera images. A semantic map may be generated using a semantic detector and at least one of the at least two camera images. The semantic map can be used as a filter for later steps (e.g., filtering out portions of an image that do not correspond to a semantically identified portion). A knowledge map may be generated using the semantic and confidence maps. The confidence map may be adjusted using the semantic map to generate an adjusted confidence map. In some aspects, the knowledge map is generated using the semantic and confidence maps in that the confidence map is first used to generate an adjusted confidence map, and the adjusted confidence map is then used to generate the knowledge map. The methods may include using the knowledge map and confidence map to conduct self-diagnostic actions. These actions may include generating an image area alert identifying a portion of the stereo camera image and a root cause of an error. The methods may also include using the adjusted confidence map to conduct self-diagnostic actions. These actions may include determining if a system failure alert should be generated or determining if an image area alert should be generated. The self-diagnostic actions may include sending error messages to a user or taking an action to correct the error. The self-diagnostic action may lead to the initiation of an automated corrective action based on the identified root cause of an error.

[0009] In a first aspect, a method for stereo camera system self-diagnostics is provided. The method may include: generating a confidence map for a stereo camera image, wherein the stereo camera image is captured using a stereo camera system that captures at least two camera images to form the stereo camera image; generating a semantic map using a semantic detector and at least one of the at least two camera images; adjusting the confidence map using the semantic map to produce an adjusted confidence map; generating a knowledge map using the semantic map and the confidence map; determining if a system failure alert should be generated using the adjusted confidence map; and determining if an image area alert should be generated using the knowledge map, wherein the image area alert identifies a portion of the stereo camera image and a root cause of an error with the portion of the stereo camera image. This method may enable comprehensive self-diagnostics for stereo camera systems by combining confidence mapping, semantic analysis, and knowledge mapping. The approach may allow for both system- wide failure detection and localized error identification, improving the overall reliability and performance of stereo vision applications.

[0010] In a second aspect, a method for stereo camera system self-diagnostics is provided. The method may include: generating a confidence map for a stereo camera image, wherein the stereo camera image is captured using a stereo camera system that captures at least two camera images to form the stereo camera image; generating a semantic map using a semantic detector and at least one of the at least two camera images; generating a knowledge map using the semantic map and the confidence map; selecting a portion of the stereo camera image using the knowledge map and based on an associated portion of the confidence map having a confidence value below a threshold; and conducting a self-diagnostic action for the stereo camera system with respect to the portion of the stereo camera image.

[0011] This method may provide a targeted approach to stereo camera system diagnostics by focusing on specific areas of concern identified through the combination of confidence mapping, semantic analysis, and knowledge mapping. This targeted approach may allow for more efficient and effective diagnostic processes. In some embodiments, the portion of the stereo camera image may be selected using only the confidence map and the semantic map.

[0012] In a third aspect, a method for stereo camera system self-diagnostics is provided. The method may include: generating a confidence map of a stereo camera image and a knowledge map of the stereo camera image using an artificial neural network, wherein the artificial neural network includes an encoding of a training data set and the training data set includes confidence values based on semantics of the stereo camera image; selecting a portion of the knowledge map based on the portion having a confidence value below a threshold in the confidence map; and conducting a self-diagnostic action for the stereo camera system with respect to the portion of the stereo camera image.

[0013] This method may leverage the power of artificial neural networks to simultaneously generate confidence and knowledge maps, incorporating semantic understanding directly into the diagnostic process. This integrated approach may lead to more nuanced and accurate system diagnostics. The artificial neural network may include an encoding of the training data set in that the parameters of the artificial neural network were modified using a training routine in response to the training data set.BRIEF DESCRIPTION OF FIGURES

[0014] Non-limiting and non-exhaustive examples are described with reference to the following figures.

[0015] FIG. 1 illustrates a flowchart for stereo camera image processing and self-diagnostic analysis, according to aspects of the present disclosure.

[0016] FIG. 2 depicts a process flow for stereo camera depth estimation and confidence analysis, in accordance with example embodiments.

[0017] FIG. 3 shows a camera image and corresponding semantic map for stereo camera system analysis, according to an embodiment.

[0018] FIG. 4 illustrates a stereo camera system capturing a scene with environmental factors, according to aspects of the present disclosure.

[0019] FIG. 5 shows camera images captured by the stereo camera system of FIG. 4, according to an embodiment.

[0020] FIG. 6 depicts a variation of the flowchart in FIG. 1 using an artificial neural network in place of alternative approaches for matching points in stereo images, in accordance with example embodiments.

[0021] FIG. 7 illustrates another variation of the flowchart in FIG. 1 using a neural network for generation of the confidence values, according to aspects of the present disclosure.

[0022] FIG. 8 shows a flowchart of a self-diagnostic analysis for a stereo camera system, according to an embodiment.

[0023] FIG. 9 depicts a flowchart for stereo camera image processing and self-diagnostic analysis, in accordance with example embodiments.

[0024] FIG. 10 illustrates another flowchart for stereo camera image processing and self-diagnostic analysis, according to aspects of the present disclosure.

[0025] FIG. 11 shows a flowchart for a self-diagnostic process for a stereo camera system, according to an embodiment.DETAILED DESCRIPTION

[0026] The following description sets forth exemplary aspects of the present disclosure. It should be recognized, however, that such description is not intended as a limitation on the scope of the present disclosure. Rather, the description also encompasses combinations and modifications to those exemplary aspects described herein.

[0027] The present disclosure relates to a self-diagnostic stereo camera system. Stereo camera systems can capture multiple images of a scene from different perspectives, allowing for depth estimation and three- dimensional reconstruction. However, these systems may encounter various challenges that affect their performance and reliability. The disclosed self-diagnostic stereo camera system addresses these challenges by incorporating mechanisms to assess its own performance, identify potential issues, and provide feedback or take corrective actions as needed.

[0028] The stereo camera system described herein can utilize various types of signals for image capture and depth estimation. These signals may include electromagnetic, optical, or acoustic signals. Thesystem can employ both passive and active capture technologies, allowing for flexibility in different environments and applications.

[0029] By integrating self-diagnostic capabilities, the stereo camera system can enhance its reliability and provide more accurate depth information. This can be particularly beneficial in applications where precise depth estimation is crucial, such as in autonomous vehicles, robotics, or industrial automation.OVERVIEW OF STEREO CAMERA SYSTEM SELF-DIAGNOSTIC ANALYSIS

[0030] FIG. 1 illustrates a flowchart depicting a process for stereo camera image processing and stereo camera system self-diagnostic analysis. The process begins with inputs from a camera image (30) and a camera image (40), which together form a stereo camera image (10).

[0031] The flowchart shows two parallel processing paths. In the first path, a pixel matching disparity analysis (215) is conducted to generate a disparity map (225). The disparity map (225) then undergoes two separate operations: a disparity transformation (226) and a match score extraction that generates a confidence map (55).

[0032] The disparity transformation (226) produces either a depth map (35) or point clouds (245). The depth map (35) or point clouds (245) can be further processed to create a modified depth map (36) using a self-diagnostic action (165) of the system.

[0033] In the parallel path, a semantic detector (60) processes the camera images to generate a semantic map (50). The semantic map (50) is used to produce two outputs: semantic information (185) for area error identification to adjust the confidence map (55) and a knowledge map (80).

[0034] The confidence map (55) undergoes adjustment based on the semantic information to create an adjusted confidence map (70). For example, the semantic information could identify an area of the scene as having high element-wise homogeneity and therefore a proclivity to cause errors with a pixel matching system. The confidence map could then be adjusted by increasing the confidence values associated with that area of the scene because the low confidence in the original confidence map is an artifact of the scene itself and not an error with the stereo camera system.

[0035] The adjusted confidence map (70), along with the knowledge map (80), can be used to generate a self-diagnostic action (165). The self-diagnostic action can be directed to areas of the stereo camera image that have a confidence value below a threshold as identified by the adjusted confidence map (70). The selfdiagnostic action (165) can be informed by the knowledge map (80). For example, the knowledge map can include information regarding the root cause of any error identified in the confidence map or adjusted confidence map.

[0036] The self-diagnostic action (165) can result in either a user alert (175), an automated corrective action (135), or both. These outputs can then influence the system's trust in the depth map and produce a modified depth map (36), as indicated by the connecting arrow.

[0037] FIG. 6 illustrates how an artificial neural network can conduct the disparity matching for disparity estimation. The neural network can be trained to output a depth map and a confidence map based on ground truth and confidence knowledge that can be done manually (when tagging images with a lot of occlusions for example) or using a classical computer vision method to estimate confidence on trained images.

[0038] This artificial neural network can thereby output, besides the disparity or depth map output, a confidence map that is associated with the disparity or depth map. As such, for each pixel in the stereo image, the disparity and the confidence would be provided.

[0039] FIG. 7 shows how an artificial neural network in the form of a stereo image and confidence map generation network can generate both the adjusted confidence map (70) and the semantic map (50) mentioned above. Such a neural network may have a more efficient compute time than traditional approaches and systems with split modules. The neural network can be trained to recognize specific features on the image and directly output the confidence map. The network can be trained using data with adjusted confidence values based on the same factors listed above (e.g., the training data can include null values or altered values for areas with high element-wise homogeneity) such that the network is trained to produce those values for those areas.

[0040] In some examples, the semantic detector (60) comprises an artificial neural network. This neural network can be trained to identify specific features or areas in the camera images that may affect the accuracy of depth estimation.

[0041] The depth map (35) can be generated from the stereo camera image (10) using a focal length of the stereo camera system. This process involves transforming the disparity values in the disparity map (225) into depth values using the known geometry of the stereo camera system.

[0042] The knowledge map (80) is generated using the semantic map (50) and the confidence map (55). This knowledge map (80) combines information about the scene content (from the semantic map) with confidence levels in the depth estimation (from the confidence map) to provide a comprehensive understanding of the stereo camera system's performance and potential issues.

[0043] FIG. 2 illustrates a process flow for stereo camera depth estimation and confidence analysis. The diagram shows three main stages connected by directional arrows.

[0044] The first stage shows two input frames: a camera image 30 and a camera image 40, each containing identical scenes depicting a potted plant and a bench or chair. The two camera images 30, 40 combine to form a stereo camera image.

[0045] The middle stage shows a three-dimensional box which represents a depth map 35. The depth map 35 represents the disparity between the two camera images 30, 40 throughout a volume captured by the stereo camera image.

[0046] The final stage displays a modified depth map 36 which is the disparity value along a single dimension marked by the parallelogram in depth map 35. The graph contains two specific disparity measurements marked as disparity value 26 and disparity value 38.

[0047] For each pixel of the camera image 30, an algorithm searches for the matched pixel in the same pixel line on the camera image 40. For each pixel in the search area, a matching score is computed. The minimum or maximum local match is extracted as the candidate pixel for the camera image 40. Based on the estimation and the strength of this minimum local, a first confidence value is extracted (i.e., Matching score).

[0048] In classical computer vision matching, and based on the previous figures, the best matching points may be on disparity value 38, and the confidence may be the ratio between this cost point and the other second local minimum (e.g. disparity value 26). When the ratio is low, the confidence may be low.

[0049] An artificial neural network can be used for disparity estimation. The artificial neural network can be trained to output a depth map 35 and a confidence map based on ground truth and confidence knowledge. The training data set includes confidence values for training images based on semantics of the training images. This allows the network to learn to recognize specific features in the camera images 30, 40 that can cause failures when matching points in the camera images 30, 40.

[0050] The artificial neural network includes an encoding of a training data set. This encoding allows the network to generalize from the training examples to new, unseen images. The network can output, besides the depth map 35 or disparity map output, a confidence map that is associated with the depth map 35 or disparity map. As such, for each pixel in the stereo image, the disparity and the confidence can be provided.

[0051] The confidence map can be used to identify areas of the stereo image where depth estimation may be less reliable. These areas can then be further analyzed or processed to improve the overall accuracy of the depth estimation.

[0052] FIG. 3 illustrates a process for generating a semantic map (50) from a camera image (40). The semantic map (50) can have the same dimensions as the original camera image (40) and can classify all of the pixels of the image as belonging to different classes (e.g., trees, street, sky, etc.).

[0053] A semantic detector (60) processes the camera image (40) to generate the semantic map (50). The semantic detector (60) may be an artificial neural network trained to recognize specific features in the camera image (40) that can cause errors when generating depth maps (35). These specific features may comprise areas with high element-wise homogeneity, such as sky regions (85).

[0054] Based on the classification in the semantic map (50), specific areas of the camera image (40) can be extracted and associated portions of the confidence map (55) can be adjusted. The semantic detector (60) may identify one or more elements that have a high homogeneity at the element level, such as sky regions (85). For example, sky or untextured areas may give poor results for pixel matching, as do low- light areas. Therefore, the pixels that cover such areas may have a reduced confidence, based on the semantic output.

[0055] Adjusting the confidence map (55) may involve modifying confidence values for portions of the stereo camera image (10) identified in the semantic map (50) as having high element-wise homogeneity. For instance, portions of the stereo camera image (10) that match with sky regions (85) may have a reduced confidence based on the low-matching capabilities in a stereo system for such areas. In response, these areas of the confidence map (55) may be removed or have their scores modified to produce an adjusted confidence map (70).

[0056] The semantic map (50) or knowledge map (80) can be directed to a subset of the overall stereo camera image (10) or to individual elements in the stereo camera image (10), regardless of whether they are contiguous or separate. This allows for focused analysis on specific areas of interest or concern.

[0057] In addition to identifying areas with high element-wise homogeneity, the semantic detector (60) can also identify blurry parts of the camera image (40). These blurry areas may be due to water, condensation, oil, or dirt on the lens of the stereo camera system (20). The semantic detector (60) can identify these portions of the camera image (40) as indicating that the stereo camera system (20) is not operating nominally. In specific examples, the semantic detector (60) may also identify the specific cause of the blurred image.

[0058] The information from the semantic map (50) can be used to create a knowledge map (80) for the associated stereo camera image (10) and ultimately be used to self-diagnose the stereo camera system (20) and provide user feedback. For example, if the semantic detector (60) identifies a blurred area caused by water droplets on the lens, this information can be included in the knowledge map (80) and used to generate a user alert or initiate an automated cleaning process.

[0059] By incorporating semantic analysis into the stereo camera system (20), the system can better understand the content of the scene and adjust its confidence in depth estimations accordingly. This can lead to more accurate and reliable depth maps (35) and improved overall performance of the stereo camera system (20).

[0060] The stereo camera system (20) can encounter various environmental factors that affect the confidence values in depth estimation. These factors are often exogenous to the system itself and do not necessarily indicate a failure of the stereo camera system (20). Understanding these factors is crucial for interpreting the confidence map and making appropriate adjustments.

[0061] FIG. 4 illustrates a stereo camera system (20) capturing a scene with a skier moving down a slope amid snowflakes. The diagram shows a pair of camera units that form the stereo camera system (20). Dotted lines extend upward from each camera unit in a triangular pattern, indicating the field of view or capture area of each camera. These lines intersect with the scene being captured above, demonstrating how the dual cameras work together to capture stereoscopic images of the scene.

[0062] In this scenario, depending on the intensity and pattern of the snowfall, the view of the skier may be partially occluded by snowflakes in the camera image (30) captured by one camera of the stereo camera system (20), while it may not be occluded in the camera image (40) captured by the other camera. This discrepancy can lead to low confidence values in the depth estimation for certain areas of the scene.

[0063] FIG. 5 further illustrates this concept by showing the camera images (30, 40) captured by the stereo camera system (20) in FIG. 4. Both images show the skier, but the camera image (40) has been occluded by a snowflake, creating a blurred area. This occlusion can result in low confidence values for the affected region when generating the depth map (35).

[0064] The knowledge map (80) can include tags defining the cause of confidence values for each pixel. In the case of FIG. 5, the knowledge map (80) could identify the blurred area as being caused by a snowflake, which is an example of a blocked line of sight. This information can be used to adjust the confidence map, as the low confidence in this area is due to an exogenous factor rather than a failure of the stereo camera system (20) itself.

[0065] Other environmental factors that can lead to low confidence values include:

[0066] 1. Low light conditions: In areas with insufficient illumination, the stereo camera system (20) may struggle to accurately match points between the two camera images (30, 40), resulting in low confidence values.

[0067] 2. Occlusion from elements on the camera: Droplets of water or condensation on the lens of one or both cameras can create distortions in the camera images (30, 40), leading to difficulties in point matching and low confidence values, imilar to water or condensation, dirt on a lens can obscure parts of the scene and create inconsistencies between the two camera images (30, 40), resulting in low confidence values for the affected areas.

[0068] 3. Occlusions that are only present in the line of sight of one camera: Snow, rain, smoke, and particulate matter can obscure parts of the scene for one camera and create inconsistencies between the two camera images (30, 40), resulting in low confidence values for the affected areas.

[0069] 4. High element-wise homogeneity: Certain elements in a scene, such as clear sky or large uniform surfaces, can present challenges for stereo matching algorithms due to the lack of distinct features. These areas often result in low confidence values despite not indicating any issue with the stereo camera system (20) itself.

[0070] The presence of these environmental factors highlights the importance of using an adjusted confidence map that incorporates semantic information. By using a semantic detector to identify areas of the scene that are likely to cause low confidence values due to exogenous factors, the stereo camera system (20) can produce a more accurate representation of its true performance.

[0071] For example, in the case of FIG. 5, the semantic detector could identify the blurred area caused by the snowflake. This information could then be used to adjust the confidence map by either ignoring the confidence scores for this region or adjusting them upwards, recognizing that the low confidence is due to a temporary occlusion rather than a systemic issue with the stereo camera system (20).

[0072] Similarly, for areas of high element-wise homogeneity like sky regions, the adjusted confidence map could account for the inherent difficulty in matching these areas, preventing them from unduly influencing the overall assessment of the stereo camera system's (20) performance.

[0073] By incorporating this semantic understanding into the confidence assessment, the stereo camera system (20) can provide a more nuanced and accurate self-diagnostic capability. This approach allows the system to differentiate between low confidence values caused by environmental factors and those that may indicate actual issues with the stereo camera system (20), leading to more reliable depth estimation and improved overall performance.

[0074] FIG. 8 illustrates a flow chart 800 of a self-diagnostic system for a stereo camera system. The system performs a self-diagnostic based on confidence values or adjusted confidence values to provide feedback to a user. Different diagnostics can be outputted.

[0075] As illustrated, the inputs are a depth map 35, a knowledge map 80, and an adjusted confidence map 70. In specific examples, the adjusted confidence map 70 can be replaced with a confidence map 55 in this diagram.

[0076] One self-diagnostic action involves the generation of an area-specific alert 38. As illustrated, this can involve using the knowledge map 80 and the adjusted confidence map 70 to identify areas where there is a failure attributable to internal or exogenous factors and generating an alert where the root cause of the error is identified for the user. This can be done by comparing the confidence value for specific areasof the stereo camera image to a threshold, as in step 801, and generating an error message if the area is large enough. The area-specific alert 38 can include information from the knowledge map 80 identifying the root cause of the error and a recommended action to alleviate the issue.

[0077] Another self-diagnostic action involves computing an average value of the adjusted confidence map 70, as in step 802, and comparing the average value to a threshold, as in step 803. This can be a different threshold from the one used for the area alert. If the average value is below the threshold, a system wide alert can be issued to indicate that the camera system overall is failing as opposed to there being a localized error. The adjusted confidence map 70 can be used for this purpose as semantic information may have already been used to remove low confidence values that are attributable to exogenous factors from consideration at this stage.

[0078] If instead the average value is above the threshold, another self -di agnostic action that can be conducted is removing elements associated with low confidence scores from the depth map or point cloud that is being generated. This is illustrated by step 804 which involves generating a filtered depth map 36 using the original depth map 35. The original confidence map can be used for this purpose or the adjusted confidence map 70. Additionally, different adjustments to the confidence map can be conducted for purposes of determining a system wide error (i.e., exogenous errors can have their confidence values changed so they do not impact a judgement of the performance of the system) as opposed to for purposes of filtering the depth map or point cloud (e.g., exogenous errors can remain because the source of the error does not matter when you are trying to produce an error free representation of the scene).

[0079] FIG. 9 illustrates a flowchart depicting a process 900 for stereo camera image processing and self-diagnostic analysis. The process 900 begins with step 901, where a confidence map is generated for a stereo camera image. From step 901, the process 900 moves to step 902, where a semantic map is generated using a semantic detector. The semantic detector can be a neural network. Following step 902, the process 900 proceeds to step 903, where the confidence map is adjusted using the semantic map to produce an adjusted confidence map 70.

[0080] The process 900 then continues to step 904, where a knowledge map 80 is generated using the semantic and confidence maps. The knowledge map 80 can be generated using the confidence map directly or the knowledge map 80 can use the confidence map in that the adjusted confidence map 70 is utilized to generate the knowledge map 80 and the confidence map was in turn used to generate the adjusted confidence map 70.

[0081] After step 904, the process 900 reaches a decision point 905, which determines if a system failure alert should be generated using the adjusted confidence map 70. From decision point 905, if a system failure alert should be generated (Yes branch), the process 900 moves to step 909, where a system failurealert is generated. If no system failure alert is needed (No branch), the process 900 proceeds to another decision point 906 to determine if an image area alert should be generated.

[0082] If an image area alert should be generated (Yes branch from step 906), the process 900 moves to step 910, where an image area alert is generated. The image area alert can identify a portion of the stereo camera image and a root cause of an error with the portion of the stereo camera image. The root causes of the error can include at least one of: a blocked line of sight, water or condensation on a lens, or dirt on a lens.

[0083] The process 900 can then proceed to step 908, where a self-corrective action is initiated. The self-corrective action can be an automated corrective action based on the identified root cause of the error. For example, if the root cause were dirt on the lens an automatic cleaning process could be conducted. As another example, if the root cause was condensation an automatic heating process could be conducted. As another example, the self-corrective action can include selecting data to discard when generating the depth map 35.

[0084] From step 908, the process 900 moves to step 907, where a depth map 35 is generated. The depth map 35 can be generated from the stereo camera image using a focal length of the stereo camera system. Additionally, the depth map 35 can be generated using the self-corrective action of discarding or adjusting specific points in the stereo camera images before combining them. After generating the depth map 35, or if no image area alert was needed (No branch from step 906), the process 900 proceeds to step 911, where the process ends.

[0085] FIG. 10 illustrates a flowchart depicting a process 1000 for stereo camera image processing and self-diagnostic analysis. The process 1000 can begin with step 1006, where a pixel matching disparity analysis is performed to generate a disparity map. From step 1006, the process 1000 branches into two parallel paths. In the first path, step 1007 involves generating a depth map 35 from the disparity map. This can be done using a focal length of the stereo camera system using computer vision processing techniques. In the second path, step 1001 involves generating a confidence map for a stereo camera image. The confidence map can be generated in step 1001 independently of step 1006 such as by an artificial neural network. Alternatively, the confidence map can be generated as part of the disparity analysis conducted in step 1006.

[0086] Following step 1001, the process 1000 proceeds to step 1002, where a semantic map is generated using a semantic detector. In specific examples, the semantic detector is a neural network trained to tag specific portions of the stereo camera image. The neural network can be trained to recognize specific features in the at least two camera images that can cause failures when matching points in the at least two camera images.

[0087] The process 1000 then moves to step 1003, where a knowledge map 80 is generated using the semantic map and confidence map. The knowledge map 80 can be focused on specific areas of the image that were identified as having low confidence values using the confidence map and can include information derived from or copied from the semantic map that indicate a root cause of the errors that lead to the low confidence values.

[0088] After step 1003, the process 1000 continues to step 1004, where a portion of the stereo camera image is selected. The portion can be selected using the knowledge map 80 and based on an associated portion of the confidence map having a confidence value below a threshold. The process 1000 then proceeds to step 1005, where a self-diagnostic action is conducted.

[0089] From step 1005, the process 1000 moves to a decision point 1010, which determines if a user alert is needed. If a user alert is needed (Yes branch), the process 1000 proceeds to step 1006, where a user alert with semantic information is generated. The user alert can include semantic information related to the portion of the stereo camera image that was selected in step 1004. For example, the semantic information can identify an error source which caused the alert to issue. The user alert can include recommended actions to resolve an error source which caused the alert to be issued as identified in the knowledge map 80. If no user alert is needed (No branch), the process 1000 moves to step 1020, where the process ends.

[0090] FIG. 11 illustrates a flowchart 1100 for a self-diagnostic process for a stereo camera system. The process begins with step 1101, where a confidence map and knowledge map 80 are generated using an artificial neural network. The confidence map generated in step 1101 can be similar to the adjusted confidence map 70 described above in that the confidence map adjusts confidence values that would be found by traditional computer vision algorithms based on semantic information in the stereo camera image. To achieve this, the neural network can include an encoding of a training data set and the training data set could include confidence values for training images that are based on semantics of the training images.

[0091] The neural network can be trained to recognize specific features in the stereo camera image that cause errors when generating depth maps 35. The specific features can comprise areas with high element-wise homogeneity. The areas with high element-wise homogeneity can comprise sky regions.

[0092] The process continues to step 1102, where a portion of the knowledge map 80 is selected based on a confidence threshold. This selection uses the confidence map to identify areas that may require diagnostic attention because the confidence is not high enough.

[0093] Following the selection, the process moves to step 1103, where a self-diagnostic action is conducted for the stereo camera system. The self-diagnostic action evaluates the selected portion of the knowledge map 80 to determine if any issues require attention.

[0094] The process then reaches a decision point labeled "Is user alert needed?" If a user alert is needed (Yes branch), the process proceeds to step 1104, where a user alert with semantic information is generated. The semantic information provides context about the identified issue with the portion of the stereo camera image. The semantic information can be generated using the knowledge map 80 that was generated in step 1101. The user alert can include recommended actions to resolve an error source identified in the knowledge map 80.

[0095] If no user alert is needed (No branch), or after generating the user alert, the process moves to step 1105, where the process ends.

[0096] A number of implementations have been described. The maps disclosed herein can be used as filters to filter out portions of other data structures with associated portions. For example, the semantic map, knowledge map, and confidence map can all be used to filter out portions of the stereo camera images based on the content of the maps. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the disclosure. Accordingly, other implementations are within the scope of the following claims.

Claims

CLAIMS1. A method comprising: generating (901) a confidence map (55) for a stereo camera image (10), wherein the stereo camera image (10) is captured using a stereo camera system (20) that captures at least two camera images (30, 40) to form the stereo camera image (10); generating (902) a semantic map (50) using a semantic detector (60) and at least one of the at least two camera images (30, 40); adjusting (903) the confidence map (55) using the semantic map (50) to produce an adjusted confidence map (70); generating (904) a knowledge map (80) using the semantic map (50) and the confidence map (55); determining (905) if a system failure alert (90) should be generated using the adjusted confidence map (70); and determining (906) if an image area alert (95) should be generated using the knowledge map (80), wherein the image area alert (95) identifies a portion (15) of the stereo camera image (10) and a root cause of an error with the portion (15) of the stereo camera image (10).

2. The method of claim 1 , wherein the semantic detector is an artificial neural network.

3. The method of claim 1, further comprising generating (907) a depth map (35) from the stereo camera image (10) using a focal length of the stereo camera system (20).

4. The method of claim 1, wherein adjusting (903) the confidence map (55) comprises modifying confidence values for portions of the stereo camera image (10) identified in the semantic map (50) as having high element-wise homogeneity.

5. The method of claim 4, wherein the portions of the stereo camera image (10) identified as having high element-wise homogeneity include sky regions (85).

6. The method of claim 1, wherein the root cause of the error includes at least one of: a blocked line of sight, water or condensation on a lens, or dirt on a lens.

7. The method of claim 6, further comprising initiating (908) an automated corrective action based on the identified root cause of the error.

8. A method comprising: generating (1001) a confidence map (55) for a stereo camera image (10), wherein the stereo camera image (10) is captured using a stereo camera system (20) that captures at least two camera images (30, 40) to form the stereo camera image (10);generating (1002) a filter using a semantic detector and at least one of the at least two camera images (30, 40); generating (1003) a knowledge map (80) using the filter and the confidence map (55); selecting (1004) a portion (15) of the stereo camera image (10) using the knowledge map (80) and based on an associated portion of the confidence map (55) having a confidence value below a threshold; and conducting (1005) a self-diagnostic action for the stereo camera system (20) with respect to the portion of the stereo camera image (10).

9. The method of claim 8, wherein the self-diagnostic action comprises generating (1006) a user alert with semantic information related to the portion of the stereo camera image (10).

10. The method of claim 9, wherein the user alert includes recommended actions to resolve an error source (205) identified in the knowledge map (80).

11. The method of claim 8, wherein generating (1001) the confidence map (55) comprises performing (1006) a pixel matching disparity analysis of the stereo camera image (10) to generate a disparity map (225).

12. The method of claim 11, further comprising generating (1007) a depth map (35) from the disparity map (225) using a focal length of the stereo camera system (20).

13. The method of claim 8, wherein the semantic detector comprises an artificial neural network.

14. The method of claim 13, wherein the artificial neural network is trained to recognize specific features in the at least two camera images (30, 40) that can cause failures when matching points in the at least two camera images (30, 40).

15. A method comprising: generating (1101) a confidence map (55) of a stereo camera image (10) and a knowledge map (80) of the stereo camera image (10) using an artificial neural network (700); selecting (1102) a portion of the knowledge map (80) based on an associated portion of the confidence map (55) to the portion of the knowledge map having a confidence value below a threshold; and conducting (1103) a self-diagnostic action for the stereo camera system (20) with respect to an associated portion of the stereo camera image to the portion of the knowledge map.

16. The method of claim 15, wherein the self-diagnostic action comprises generating (1104) a user alert with semantic information related to the associated portion of the stereo camera image.

17. The method of claim 16, wherein the user alert includes recommended actions to resolve an error source identified in the knowledge map (80).

18. The method of claim 15, wherein the artificial neural network (700) is trained to recognize specific features in the stereo camera image (10) that cause errors when generating depth maps.

19. The method of claim 18, wherein the specific features comprise areas with high element-wise homogeneity.

20. The method of claim 19, wherein the areas with high element-wise homogeneity comprise sky regions(85).

21. The method of claim 15, wherein the artificial neural network (65) includes an encoding of a training data set and the training data set includes confidence values for training images that are based on semantics of the training images

Citation Information

Patent Citations

  • Method for monitoring of stereo camera arrangement used for detecting environment of e.g. vehicle during manufacturing vehicle, involves determining confidence measure indicating efficiency and / or reliability of arrangement

    DE102011108995A1

  • Method for using a stereovision camera arrangement

    EP2293588A1

  • Method and device for checking the visibility of a camera for surroundings of an automobile

    US20130070966A1

  • Method, apparatus and computer program product for disparity map estimation of stereo images

    US20150248769A1

  • Method for determining a state of obstruction of at least one camera installed in a stereoscopic system

    US20150279018A1