Method and device for identifying occlusion of vehicle-mounted camera lenses
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-30
- Publication Date
- 2026-08-11
AI Technical Summary
这几种镜头遮挡检测方案通常计算复杂度较高,不适用于在车载场景对车载镜头进行实时的遮挡检测
Smart Images

Figure CN114841910B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology (including assisted driving and driverless driving), and in particular to a method and device for identifying occlusion of vehicle-mounted cameras. Background Technology
[0002] Autonomous driving (including assisted driving and driverless driving) is a crucial direction in the development of intelligent vehicles, and an increasing number of vehicles are adopting autonomous driving systems to achieve autonomous driving functions. Typically, as a component of an autonomous driving system, onboard cameras are used to capture images of the vehicle's surrounding environment, allowing the system to formulate appropriate driving strategies based on these images and thus control the vehicle to achieve autonomous driving. Therefore, the performance of onboard cameras is critical to autonomous driving.
[0003] However, in some scenarios, the vehicle's onboard camera may be obstructed by foreign objects such as leaves or plastic bags, preventing it from functioning properly. This could potentially affect the implementation of autonomous driving functions. For example, in... Figure 1A or Figure 1B or Figure 1C When the vehicle's camera is obstructed, the image captured by the camera becomes blurry, affecting normal driving. For safety reasons, such as... Figure 2A As shown, it can detect whether the vehicle camera is obstructed. Once the vehicle camera is obstructed, an alarm mechanism is activated. At this time, the autonomous driving function can be stopped (such as stopping the advanced driver assistance system (ADAS) function), and the driver can take control of the driving (i.e., manual intervention) to reduce driving safety hazards.
[0004] Existing technologies offer several lens occlusion detection schemes based on depth information, motion information, and image statistical information. These schemes typically have high computational complexity and are not suitable for real-time occlusion detection of automotive lenses in automotive scenarios. Summary of the Invention
[0005] This application provides a method and apparatus for identifying occlusion of vehicle-mounted cameras, which improves the accuracy of determining road geometry, reduces the influence of non-road factors, and better assists the vehicle in determining driving strategies.
[0006] To achieve the above objectives, this application adopts the following technical solution:
[0007] In a first aspect, embodiments of this application provide a method for identifying occlusion of a vehicle-mounted camera. This method is applied to a device with autonomous driving (including assisted driving) functions, such as a vehicle, a chip system in the vehicle, and an operating system and driver running on the processor. The method includes: acquiring a frame difference image and determining whether the vehicle-mounted camera is occluded based on the frame difference image.
[0008] The frame difference image is an image obtained by performing a difference operation on the first image and the second image captured by the vehicle-mounted camera; the frame interval between the first image and the second image is N, where N is a positive integer; neither the first image nor the second image is a background image.
[0009] In one possible design, the frame difference image is one or more frame difference images, including a first frame difference image and a second frame difference image; the first frame difference image is an image obtained by performing a difference operation on a first image captured by the vehicle-mounted lens at a first moment and a second image preceding the first image; the second frame difference image is an image obtained by performing a difference operation on a first image captured by the vehicle-mounted lens at a second moment and a second image preceding the first image, wherein the first moment and the second moment are different.
[0010] In one possible design, the method further includes: obtaining a first confidence level and a first area corresponding to one or more frame difference images;
[0011] Wherein, the confidence score corresponding to a frame difference image is used to represent the degree of confidence that the first region in the first image corresponding to the frame difference image is identified as an occluded region, and the confidence score is obtained based on the LBP response score, Laplacian variance, and Laplacian mean of the first region; the first area corresponding to the frame difference image is the area of the first region in the first image corresponding to the frame difference image.
[0012] The first region is the region obtained by removing the preset region from the second region; the second region is included in the first image, and the position of the second region in the first image is the same as the position of the third region in the frame difference image corresponding to the first image; the third region is the region composed of pixels in the frame difference image whose pixel value is less than or equal to a first threshold; the preset region includes any one or more of the following: the imaging region of the road surface, the imaging region of the snow, the imaging region of the sky, and the imaging region of the grass.
[0013] In one possible design, there are M frame difference images; determining whether the vehicle-mounted camera is obstructed based on the frame difference images includes:
[0014] If P frame difference images out of the M frame difference images satisfy the first condition, then it is determined that the vehicle-mounted camera is blocked.
[0015] Wherein, the first condition is that the confidence level is greater than or equal to the first confidence threshold, and the first area is greater than or equal to the first area threshold; M and P are positive integers, and M is greater than or equal to P.
[0016] In one possible design, there are M frame difference images; determining whether the vehicle-mounted camera is obstructed based on the frame difference images includes:
[0017] If Q consecutive frame difference images among the M frame difference images satisfy the second condition, then it is determined that the vehicle-mounted camera is blocked;
[0018] The second condition is that the confidence level is greater than or equal to the second confidence threshold, and the first area is greater than or equal to the second area threshold; M and Q are positive integers, and M is greater than or equal to Q.
[0019] In one possible design, acquiring the frame difference image includes:
[0020] Obtain vehicle speed and / or a first identifier bit; the first identifier bit is used to indicate whether the image processing result of the image signal processing ISP is normal.
[0021] If the first identifier indicates that the image processing result of the ISP is normal, and / or the vehicle speed is greater than or equal to the vehicle speed threshold, then the frame difference image is acquired.
[0022] Secondly, this application provides a vehicle-mounted lens occlusion recognition device, the device comprising:
[0023] The acquisition module is used to acquire a frame difference image; the frame difference image is an image obtained by performing a difference operation on a first image and a second image captured by the vehicle-mounted camera; the frame interval between the first image and the second image is N, where N is a positive integer; neither the first image nor the second image is a background image.
[0024] The judgment module is used to determine whether the vehicle-mounted camera is obstructed based on the frame difference image.
[0025] In one possible design, the frame difference image is one or more frame difference images, including a first frame difference image and a second frame difference image; the first frame difference image is an image obtained by performing a difference operation on a first image captured by the vehicle-mounted lens at a first moment and a second image preceding the first image; the second frame difference image is an image obtained by performing a difference operation on a first image captured by the vehicle-mounted lens at a second moment and a second image preceding the first image, wherein the first moment and the second moment are different.
[0026] In one possible design, the acquisition module is further configured to acquire a first confidence level and a first area corresponding to one or more frame difference images;
[0027] Wherein, the confidence score corresponding to a frame difference image is used to represent the degree of confidence that the first region in the first image corresponding to the frame difference image is identified as an occluded region, and the confidence score is obtained based on the LBP response score, Laplacian variance, and Laplacian mean of the first region; the first area corresponding to the frame difference image is the area of the first region in the first image corresponding to the frame difference image.
[0028] The first region is the region obtained by removing the preset region from the second region; the second region is included in the first image, and the position of the second region in the first image is the same as the position of the third region in the frame difference image corresponding to the first image; the third region is the region composed of pixels in the frame difference image whose pixel value is less than or equal to a first threshold; the preset region includes any one or more of the following: the imaging region of the road surface, the imaging region of the snow, the imaging region of the sky, and the imaging region of the grass.
[0029] In one possible design, there are M frame difference images; the judgment module is used to determine whether the vehicle lens is blocked based on the frame difference images, including: if P frame difference images among the M frame difference images satisfy a first condition, then determine that the vehicle lens is blocked;
[0030] Wherein, the first condition is that the confidence level is greater than or equal to the first confidence threshold, and the first area is greater than or equal to the first area threshold; M and P are positive integers, and M is greater than or equal to P.
[0031] In one possible design, there are M frame difference images; the judgment module is used to determine whether the vehicle lens is blocked based on the frame difference images, including: if Q consecutive frame difference images in the M frame difference images satisfy a second condition, then determine that the vehicle lens is blocked.
[0032] The second condition is that the confidence level is greater than or equal to the second confidence threshold, and the first area is greater than or equal to the second area threshold; M and Q are positive integers, and M is greater than or equal to Q.
[0033] In one possible design, the acquisition module is used to acquire a frame difference image, including: acquiring vehicle speed and / or a first identifier bit; the first identifier bit is used to identify whether the image processing result of the image signal processing ISP is normal; if the first identifier bit indicates that the image processing result of the ISP is normal, and / or the vehicle speed is greater than or equal to a vehicle speed threshold, then the frame difference image is acquired.
[0034] Thirdly, a vehicle-mounted lens occlusion recognition device is provided, comprising: a processor and a memory; the memory is used to store computer execution instructions, and when the vehicle-mounted lens occlusion recognition device is running, the processor executes the computer execution instructions stored in the memory to cause the vehicle-mounted lens occlusion recognition device to perform the vehicle-mounted lens occlusion recognition method as described in the first aspect and any one of the first aspects.
[0035] Fourthly, a vehicle-mounted lens occlusion recognition device is provided, comprising: a processor; the processor is configured to be coupled to a memory, and after reading instructions from the memory, execute the vehicle-mounted lens occlusion recognition method as described in the first aspect and any one of the first aspects.
[0036] Fifthly, this application also provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the vehicle lens occlusion recognition method as described in the first aspect and any one of the first aspects.
[0037] Sixthly, this application also provides a computer program product, including instructions that, when run on a computer, cause the computer to execute the vehicle lens occlusion recognition method as described in the first aspect and any one of the first aspects.
[0038] Seventhly, embodiments of this application provide a vehicle-mounted lens occlusion recognition device. This device can be a chip system, including a processor and potentially a memory, for implementing the functions of the aforementioned method. The chip system can be composed of chips or may include chips and other discrete components.
[0039] Eighthly, a vehicle-mounted lens occlusion recognition device is provided. The device can be a circuit system, the circuit system including a processing circuit, the processing circuit being configured to perform the vehicle-mounted lens occlusion recognition method as described in the first aspect and any one of the first aspects. Attached Figure Description
[0040] Figures 1A-1C Example image of a vehicle-mounted camera being obstructed;
[0041] Figure 2A , Figure 2B This is a flowchart illustrating a conventional lens recognition method.
[0042] Figure 3 A schematic diagram of the structure of an autonomous vehicle provided in this application embodiment;
[0043] Figure 4 A second schematic diagram of the structure of an autonomous vehicle provided in this application embodiment;
[0044] Figure 5A schematic diagram of the structure of a computer system provided in an embodiment of this application;
[0045] Figure 6 This application provides an schematic diagram of a cloud-based command-driven autonomous vehicle according to an embodiment of the present application.
[0046] Figure 7 This application provides a schematic diagram of the structure of a computer program product according to an embodiment of the present application.
[0047] Figure 8 , Figure 9-1 , Figure 9-2 This is a schematic flowchart of the vehicle lens occlusion recognition method provided in an embodiment of this application;
[0048] Figure 9-3 A schematic diagram of the rule control module provided in an embodiment of this application;
[0049] Figures 10-21 A schematic diagram of a scenario for the vehicle-mounted camera occlusion recognition method provided in this application embodiment;
[0050] Figure 22 , Figure 23 This is a schematic diagram of the structure of the vehicle-mounted lens occlusion recognition device provided in the embodiments of this application. Detailed Implementation
[0051] For ease of understanding, the relevant terms used in the embodiments of this application are explained as follows:
[0052] Autonomous driving: Autonomous driving technology relies on the collaborative efforts of artificial intelligence, computer vision, radar, monitoring devices, and global positioning systems to enable computers to operate motor vehicles automatically and safely without any active human intervention. According to the classification standards of the Society of Automotive Engineers (SAE), autonomous driving technology is divided into: no automation (L0), driver assistance (L1), partial automation (L2), conditional automation (L3), high automation (L4), and full automation (L5).
[0053] Existing technologies include several lens occlusion detection schemes based on depth information, motion information, and image statistical information. Among these, depth information-based schemes are mostly used in security scenarios and require stereo lenses or other sensors besides a single lens to acquire depth information, placing high demands on lens type or sensor type. Motion information-based lens occlusion recognition schemes require calculating the motion information of objects in the image, resulting in high computational complexity. Schemes based on image statistical information are currently the mainstream approach, and are widely applicable to security, video surveillance, and similar scenarios. See also Figure 2BSpecifically, this can be implemented through the following steps: First, extract suspected occlusion areas from the image through background modeling. Then, calculate statistical information about the suspected occlusion areas, such as brightness changes, edge information, and grayscale histograms, and determine whether the lens is occluded based on the statistical information and corresponding thresholds.
[0054] Background modeling refers to modeling the background and extracting it from the image based on the constructed background model. Typically, relatively fixed portions of an image are considered the background. One possible approach is to divide the current frame into blocks, calculate the features of each block (e.g., using the center-symmetric local binary pattern (CSLBP) operator), and perform clustering based on the calculation results. For example, blocks with the same features are classified as background, while blocks with different features are classified as foreground. Blocks with CSLBP textures that meet certain conditions are considered background. In this way, the background can be extracted from the current frame.
[0055] Then, inter-frame difference (IF) is performed on the current frame image and the extracted background to obtain an IF image, and the suspected occlusion region is obtained based on the IF image. The occlusion region refers to the image formed on the image sensor by the obstruction through the lens when the lens is blocked. The suspected occlusion region (also called the candidate occlusion region) is an image area that is initially identified as potentially occluded. Further judgment is needed to determine whether the image is indeed caused by an obstruction of the vehicle-mounted lens. Optionally, after obtaining the suspected occlusion region, boundary retrieval can be performed to obtain a more accurate suspected occlusion region.
[0056] It's understandable that in some scenarios, the background may remain unchanged or change only slightly. Therefore, by comparing the current frame with the background frame, we can determine whether there are any objects obscuring the background. For example, if the frame difference indicates a significant difference between the background and the current frame, it suggests that there may be objects obscuring the background.
[0057] After extracting suspected occlusion areas, to determine whether they are actually occluded areas, specific information about these areas needs to be analyzed. This includes analyzing brightness variations, edge information, gradient histograms, and grayscale histograms. Edge information can be obtained using edge detection algorithms. When the statistical information meets predetermined conditions, the suspected occlusion area is determined to be a true occlusion area.
[0058] The aforementioned occlusion detection method also requires background extraction and background modeling, which results in a large computational load, complex calculations, and poor real-time performance.
[0059] Furthermore, in moving scenes, such as in-vehicle scenes, the background captured by the camera moves with the vehicle and is constantly changing. The frame difference between this constantly changing background and the current frame makes it difficult to detect occlusions. For example, if the background is extracted from a previous historical frame and the frame difference is calculated between that background and the current frame, the result may indicate a significant difference. In this case, it's impossible to determine whether the difference is due to an occlusion in the current frame or a significant difference between the background in a previous frame and the current frame. Therefore, even using background modeling to detect occlusion in moving scenes can lead to misjudgments or failure to detect occlusions; that is, the accuracy of identifying suspected occlusion areas through background modeling is low.
[0060] Furthermore, the detection results rely excessively on the segmentation and clustering strategies. The clustering results are strongly correlated with both the segmentation and clustering strategies; that is, the accuracy of the results largely depends on the appropriateness of these strategies. Inappropriate selection of these strategies can significantly lead to inaccurate results. Therefore, target occlusion detection methods suffer from high computational complexity and poor real-time performance, making them difficult to apply to occlusion detection in motion scenarios with high real-time requirements.
[0061] To address this, this application provides a method for identifying occlusion of vehicle-mounted cameras. Typically, when a vehicle-mounted camera is occluded, large connected regions with minimal changes in brightness and other properties will appear at the same location across several frames of images captured by the camera over a period of time. For example, ... Figure 1C As shown, the imaging area of the foreign object may exist in multiple frames of images captured by the vehicle-mounted camera. Based on this characteristic, in this embodiment, the suspected occlusion area is determined directly based on the frame difference between the current frame and the historical frames. This eliminates the need for background modeling and other operations, resulting in low computational load and complexity, reduced computation time, and improved real-time performance. Therefore, it can meet the occlusion detection requirements in motion scenarios.
[0062] The vehicle-mounted lens occlusion recognition method provided in this application is applied in motion scenarios. For example, it can be used in vehicle-mounted scenarios, outdoor robots, automated guided vehicles (AVGs) in indoor and outdoor factories, and mobile devices such as smartphones. When the portion of the image sensor obscured by a foreign object in a motion scenario exceeds a certain proportion of the total imaging area and persists for a certain period, the occlusion recognition method of this application can be used to promptly determine the obstruction and issue an alarm for subsequent response.
[0063] The technical methods of this application embodiment can be applied to vehicles with autonomous driving or assisted driving functions, or to other devices (such as cloud servers) with autonomous driving control functions. Vehicles can implement the vehicle lens occlusion recognition method provided in this application embodiment through their included components (including hardware and software) to identify whether the vehicle lens is occluded. Alternatively, other devices (such as servers) can be used to implement the vehicle lens occlusion recognition method of this application embodiment to identify whether the vehicle lens is occluded, in order to formulate corresponding driving strategies.
[0064] Figure 3 This is a functional block diagram of a vehicle 100 provided in an embodiment of this application. In one embodiment, the vehicle 100 is configured in either assisted driving or fully autonomous driving mode. For example, while in assisted driving or fully autonomous driving mode, the vehicle 100 can identify whether the onboard camera is obstructed, and formulate a driving strategy based on the identification result, thereby controlling the vehicle 100 to perform automated driving. When the vehicle 100 is in autonomous driving mode, the vehicle 100 does not interact with the driver and autonomously completes actions such as obstacle avoidance, following other vehicles, lane keeping, and automatic parking. When the vehicle 100 is in assisted driving mode, the vehicle 100 provides prompts to the driver according to the driving strategy, and the driver completes actions such as obstacle avoidance, following other vehicles, lane keeping, and automatic parking according to the prompts.
[0065] Vehicle 100 may include various subsystems, such as a mobility system 110, a sensor system 120, a control system 130, one or more peripheral devices 140, a power supply 150, a computer system 160, and a user interface 170. Optionally, vehicle 100 may include more or fewer subsystems, and each subsystem may include multiple components. Furthermore, each subsystem and component of vehicle 100 may be interconnected via wired or wireless means.
[0066] The propulsion system 110 may include components that provide power to the vehicle 110, such as an engine, transmission, etc.
[0067] Sensor system 120 may include several sensors for sensing information about the environment surrounding vehicle 100. For example, sensor system 120 may include at least one of a positioning system 121 (which may be a GPS system, a BeiDou system, or another positioning system), an inertial measurement unit (IMU) 122, a radar sensor 123, a lidar sensor 124, a vision sensor 125, an ultrasonic sensor 126, and a sonar sensor 127. Optionally, sensor system 120 may also include sensors for the internal systems of the monitored vehicle 100 (e.g., an in-vehicle air quality monitor, a fuel gauge, an oil temperature gauge, etc.). Sensor data from one or more of these sensors can be used to detect objects and their corresponding characteristics (position, shape, orientation, speed, etc.). This detection and identification is a key function for the safe operation of the autonomous vehicle 100.
[0068] The positioning system 121 can be used to estimate the geographical location of the vehicle 100. The IMU 122 is used to sense changes in the position and orientation of the vehicle 100 based on inertial acceleration. In one embodiment, the IMU 122 can be a combination of an accelerometer and a gyroscope.
[0069] The radar sensor 123 can use electromagnetic wave signals to sense objects in the surrounding environment of the vehicle 100. In some embodiments, in addition to sensing the position of an object, the radar sensor 123 can also be used to sense the radial velocity of the object and / or the radar cross section (RCS) of the object.
[0070] The ultrasonic sensor 126 can use ultrasonic waves to sense objects in the surrounding environment of the vehicle 100. In some embodiments, in addition to sensing the position of an object, the ultrasonic sensor 126 can also be used to sense the radial velocity of the object and / or the echo amplitude of the object.
[0071] Sonar sensor 127 can use sound waves to sense objects in the surrounding environment of vehicle 100. In some embodiments, in addition to sensing the position of an object, sonar sensor 127 can also be used to sense the radial velocity of the object and / or the sonar target intensity (sonar TS) of the object.
[0072] The lidar 124 can use lasers to sense objects in the environment in which the vehicle 100 is located. In some embodiments, the lidar 124 may include one or more laser sources, a laser scanner, and one or more detectors, as well as other system components.
[0073] The vision sensor 125 can be used to capture multiple images of the surrounding environment of the vehicle 100. The vision sensor 125 can be a still camera or a video camera.
[0074] In this embodiment, the vision sensor 125 can be used in a vehicle-mounted lens, or vehicle-mounted camera. It may be obstructed by foreign objects such as leaves or plastic bags. Therefore, the images it captures may possess certain characteristics, based on which it can be determined that the vision sensor 125 is obstructed.
[0075] The control system 130 can control the operation of the vehicle 100 and its components. The control system 130 may include at least one of various elements, such as a computer vision system 131, a route control system 132, and an obstacle avoidance system 133.
[0076] Computer vision system 131 is operable to process and analyze images captured by vision sensor 125 and measurement data obtained by radar sensor 123 to identify objects and / or features in the environment surrounding vehicle 100. The objects and / or features may include traffic signals, road boundaries, and obstacles. Computer vision system 131 may use object recognition algorithms, structure from motion (SFM) algorithms, video tracking, and other computer vision techniques. In some embodiments, computer vision system 131 may be used to map the environment, track objects, estimate object speeds, etc. In embodiments of this application, computer vision system 131 is operable to process and analyze images captured by an onboard camera to identify features of the image.
[0077] The route control system 132 is used to determine the driving route of the vehicle 100. In some embodiments, the route control system 132 may combine data from radar sensor 123, positioning system 121 and one or more predetermined maps to determine the driving route of the vehicle 100.
[0078] The obstacle avoidance system 133 is used to identify, assess and avoid or otherwise traverse potential obstacles in the environment of the vehicle 100.
[0079] Of course, in one instance, the control system 130 may include additional or alternative components besides those shown and described. Alternatively, some of the components shown above may be reduced.
[0080] Vehicle 100 can obtain necessary information using wireless communication system 140, which can communicate wirelessly with one or more devices directly or via a communication network. For example, wireless communication system 140 can use 3G cellular communication, such as CDMA, EVDO, GSM / GPRS, or 4G cellular communication, such as LTE, or 5G cellular communication. Wireless communication system 140 can communicate using WiFi and wireless local area network (WLAN). In some embodiments, wireless communication system 140 can communicate directly with devices using infrared links, Bluetooth, or ZigBee. Other wireless protocols, such as various vehicle communication systems, may also be used. For example, wireless communication system 140 may include one or more dedicated short-range communications (DSRC) devices.
[0081] Some or all of the functions of vehicle 100 are controlled by computer system 160. Computer system 160 may include at least one processor 161, which executes instructions 1621 stored in a non-transitory computer-readable medium such as data storage device 162. Computer system 160 may also be multiple computing devices that control individual components or subsystems of vehicle 100 in a distributed manner. In this embodiment, processor 161 may acquire multiple frames captured by onboard cameras and perform frame difference analysis between historical frames and the current frame to determine candidate occlusion areas based on the frame difference result. Optionally, areas such as sky, road surface, and snow may be excluded from the candidate occlusion areas.
[0082] Processor 161 can be any conventional processor, such as a commercially available central processing unit (CPU). Alternatively, the processor can be a special-purpose device such as an application-specific integrated circuit (ASIC) or other hardware-based processor. Although Figure 3The processor, memory, and other components within the same physical housing are functionally illustrated; however, those skilled in the art will understand that the processor, computer system, or memory may actually include multiple processors, computer systems, or memories that may be stored within the same physical housing, or multiple processors, computer systems, or memories that may not be stored within the same physical housing. For example, memory may be a hard disk drive, or other storage media located in a different physical housing. Therefore, references to processors or computer systems will be understood to include references to collections of processors or computer systems or memories that may operate in parallel, or collections of processors or computer systems or memories that may not operate in parallel. Unlike using a single processor to perform the steps described herein, some components, such as steering and deceleration components, may each have their own processor that performs calculations only related to the component's specific function.
[0083] In this embodiment of the application, the processor includes an image signal processor (ISP) for processing the signal output by the image sensor.
[0084] In the various aspects described herein, the processor may be located remotely from the vehicle and communicate wirelessly with the vehicle. In other aspects, some of the processes described herein are executed on a processor located within the vehicle, while others are executed by a remote processor, including taking the necessary steps to perform a single operation.
[0085] Optionally, the components described above are merely examples. In actual applications, components in each of the above modules may be added or removed as needed. Figure 3 This should not be construed as a limitation on the embodiments of this application.
[0086] Optionally, the autonomous vehicle 100 or the computing device associated with the autonomous vehicle 100 (such as...) Figure 3 The computer system 160, computer vision system 131, and data storage device 162 can identify whether the vehicle-mounted camera is obstructed by a foreign object based on the image captured by the camera. Optionally, the vehicle 100 can adjust its driving strategy based on whether the vehicle-mounted camera is obstructed. In some examples, if the vehicle-mounted camera is obstructed by a foreign object, then, for driving safety reasons, autonomous driving can be stopped and the vehicle can be switched to assisted driving.
[0087] The aforementioned vehicle 100 can be a car, truck, motorcycle, bus, ship, airplane, helicopter, lawnmower, recreational vehicle, amusement park vehicle, construction equipment, tram, golf cart, train, and handcart, etc., and this application embodiment does not impose any special limitations.
[0088] In other embodiments of this application, the autonomous vehicle may further include hardware structures and / or software modules to implement the above functions in the form of hardware structures, software modules, or a combination of hardware structures and software modules. Whether a particular function is implemented in the form of hardware structures, software modules, or a combination of hardware structures and software modules depends on the specific application and design constraints of the technical solution.
[0089] See Figure 4 For example, a vehicle may include the following modules:
[0090] The environmental perception module 201 is used to acquire measurement data of target objects detected by sensors. Sensors can be of various types, such as cameras, lidar, millimeter-wave radar, ultrasonic sensors, sonar sensors, etc. The environmental perception module transmits this data to the rule control module 202 so that the rule control module 202 can generate action commands.
[0091] Rule control module 202: This module is a traditional control module in autonomous vehicles. It receives image data collected from the environment perception module 201, and determines whether there are foreign objects obstructing the vehicle camera based on the image data, as well as the specific location of the obstruction, so as to generate corresponding driving strategies and corresponding action commands. It then sends the action commands to the vehicle control module 203, which instructs the vehicle control module 203 to perform driving control on the vehicle.
[0092] Vehicle control module 203: Used to receive action commands from rule control module 202 to control the driving operation of the vehicle.
[0093] Vehicle communication module 204 ( Figure 4 (Not shown in the image): Used for information exchange between the vehicle and other vehicles.
[0094] Storage component 205 ( Figure 4 (Not shown in the image), used to store the executable code of each of the above modules. Running this executable code can implement some or all of the method flows of the embodiments of this application.
[0095] In one possible implementation of the embodiments of this application, such as Figure 5 As shown, Figure 3The computer system 160 shown includes a processor 301 coupled to a system bus 302. The processor 301 can be one or more processors, each of which can include one or more processor cores. A video adapter 303 drives a display 309, which is coupled to the system bus 302. The system bus 302 is coupled to an input / output (I / O) bus 305 via a bus bridge 304. An I / O interface 306 is coupled to the I / O bus 305. The I / O interface 306 communicates with various I / O devices, such as input devices 307 (e.g., keyboard, mouse, touchscreen), a media tray 308 (e.g., CD-ROM, multimedia interface), a transceiver 315 (capable of sending and / or receiving radio communication signals), a camera 310 (capable of capturing still and moving digital video images), and an external universal serial bus (USB) interface 311. Optionally, the interface connected to the I / O interface 306 can be a USB interface.
[0096] The processor 301 can be any conventional processor, including a reduced instruction set computer (RISC) processor, a complex instruction set computer (CISC) processor, or a combination thereof. Optionally, the processor can be a special-purpose device such as an application-specific integrated circuit (ASIC). Optionally, the processor 301 can be a neural network processor or a combination of a neural network processor and the aforementioned conventional processors.
[0097] Alternatively, in the various embodiments described herein, the computer system 160 may be located remotely from the autonomous vehicle and may communicate wirelessly with the autonomous vehicle 100. In other aspects, some of the processes described herein may be set to execute on a processor within the autonomous vehicle, while others may be executed by a remote processor, including taking actions necessary to perform a single manipulation.
[0098] Computer system 160 can communicate with software deployment server 313 via network interface 312. Network interface 312 is a hardware network interface, such as a network interface card (NIC). Network 314 can be an external network, such as the Internet, or an internal network, such as Ethernet or a Virtual Private Network (VPN). Optionally, network 314 can also be a wireless network, such as a WiFi network or a cellular network.
[0099] In other embodiments of this application, the vehicle-mounted lens occlusion recognition method of this application can also be executed by a chip system. This application provides a chip system. Through the joint operation of a main CPU and a neural processing unit (NPU), it can achieve... Figure 3 The corresponding algorithms for the functions required by vehicle 100 can also be implemented. Figure 4 The corresponding algorithms for the functions required by the vehicle shown can also be implemented. Figure 3 The corresponding algorithms for the functions required by the computer system 160 shown.
[0100] In other embodiments of this application, computer system 160 may also receive information from or transfer information to other computer systems. Alternatively, sensor data collected from the sensor system 120 of vehicle 100 may be transferred to another computer for processing. Data from computer system 160 may be transmitted via a network to a cloud-based computer system for further processing. The network and intermediate nodes may include various configurations and protocols, including the Internet, World Wide Web, intranet, virtual private network, wide area network, local area network, private network using proprietary communication protocols of one or more companies, Ethernet, WiFi, and HTTP, as well as various combinations thereof. Such communication may be performed by any device capable of transmitting data to and from other computers, such as modems and wireless interfaces.
[0101] See Figure 6 This is an example of interaction between an autonomous vehicle and a cloud service center (cloud server). The cloud service center can receive information (such as data collected by vehicle sensors or other information) from autonomous vehicles 413, 412 within its environment 400 via a network 411, such as a wireless communication network.
[0102] Based on the received data, the cloud service center 420 runs its stored programs for identifying occlusion of vehicle cameras to determine whether the cameras of autonomous vehicles 413 and 412 are occluded. These programs can include: programs that calculate frame differences between historical and current frames; programs that perform color segmentation, histogram backprojection, or template matching; or programs that calculate occlusion confidence.
[0103] For example, cloud service center 420 can provide portions of a map to vehicles 413 and 412 via network 411. In other examples, operations can be divided among different locations. For instance, multiple cloud service centers can receive, verify, combine, and / or send information reports. In some examples, information reports and / or sensor data can also be sent between vehicles. Other configurations are also possible.
[0104] like Figure 7 As shown, in some examples, the signal-bearing medium 501 may comprise a computer-readable medium 503, such as, but not limited to, a hard disk drive, a compact disc (CD), a digital video optical disc (DVD), a digital magnetic tape, a memory, read-only memory (ROM), or random access memory (RAM), etc. In some embodiments, the signal-bearing medium 501 may comprise a computer-recordable medium 504, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, etc. In some embodiments, the signal-bearing medium 501 may comprise a communication medium 505, such as, but not limited to, digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, etc.). Therefore, for example, the signal-bearing medium 501 may be conveyed by a wireless communication medium 505 (e.g., a wireless communication medium conforming to the IEEE 802.11 standard or other transmission protocols). One or more program instructions 502 may be, for example, computer-executable instructions or logical implementation instructions. In some examples, such as for... Figures 3 to 7 The described computing device can be configured to provide various operations, functions, or actions in response to program instructions 502 transmitted to the computing device via one or more of computer-readable media 503, and / or computer-recordable media 504, and / or communication media 505. It should be understood that the arrangements described herein are merely illustrative. Therefore, those skilled in the art will understand that other arrangements and other elements (e.g., machines, interfaces, functions, sequences, and functional groups, etc.) can be used instead, and some elements can be omitted depending on the desired result. Furthermore, many of the described elements are functional entities that can be implemented as discrete or distributed components, or in any suitable combination and location in conjunction with other components.
[0105] The vehicle-mounted camera occlusion recognition method provided in this application embodiment can be applied to mobile scenarios. Optionally, mobile scenarios include vehicle-mounted scenarios. Vehicle-mounted scenarios include autonomous / semi-autonomous driving scenarios. This method can identify whether the vehicle-mounted camera is occluded, for example, by identifying... Figure 1A The occlusion is shown. This method can be derived from... Figures 3-7 The method is executed by a device (such as a vehicle, a vehicle processor or other components, or a server). The following describes in detail, with reference to the accompanying drawings, the vehicle-mounted camera occlusion recognition method according to embodiments of this application.
[0106] This application provides a method for identifying occlusion of vehicle-mounted cameras, such as... Figure 8 As shown, the method includes the following steps:
[0107] S101. Obtain the frame difference image.
[0108] The frame difference image is an image obtained by performing a difference operation on the first image and the second image captured by the vehicle-mounted camera; the frame interval between the first image and the second image is N, where N is a positive integer; neither the first image nor the second image is a background image.
[0109] Neither the first image nor the second image is a background image. This means that, in this embodiment of the application, there is no need to perform background modeling operations on the images captured by the vehicle-mounted camera.
[0110] In some embodiments, performing inter-frame difference operations on the current frame and historical frames refers to performing difference operations on the grayscale images of the current frame and the grayscale images of historical frames. Of course, the embodiments of this application are not limited to this. If the inter-frame difference refers to the difference operation between grayscale images, then it is necessary to first convert the current frame and historical frames into their respective grayscale images, and then perform the difference operation on the grayscale images of the current frame and the grayscale images of historical frames. The operations of converting the current frame from a color image to a grayscale image, and converting the historical frames from color images to grayscale images, can be found in the prior art.
[0111] Inter-frame differencing can be performed between the grayscale images of the current frame and the grayscale images of previous frames. This can be achieved by subtracting the pixels at corresponding positions (i.e., those with the same coordinates) in the grayscale images of the current and previous frames. This yields a frame difference image. Of course, other inter-frame differencing methods can also be used, and the specific method of inter-frame differencing in this embodiment is not limited.
[0112] As can be seen, the frame difference image can represent the difference between two frames.
[0113] S102. Determine whether the vehicle-mounted camera is obstructed based on the frame difference image.
[0114] It is understandable that when a vehicle camera lens is obstructed by a foreign object, the object is likely to remain obstructed for a considerable period. This means that the positional relationship between the foreign object and the camera lens is likely to remain unchanged or change only slightly. Consequently, the image formed by the foreign object through the camera lens may show minimal changes over a longer period. In this embodiment, the frame difference image can represent the difference between the first and second images. Therefore, based on the frame difference image, it is possible to detect which pixels in the first or second image are likely to be obstructed by a foreign object, i.e., images suspected of being obstructed by a foreign object. Specifically, the area formed by pixels with pixel values less than or equal to a pixel threshold in the frame difference image corresponds to images obstructed by a foreign object. Thus, based on images suspected of being obstructed by a foreign object, it is possible to further determine whether the vehicle camera lens is indeed obstructed.
[0115] The vehicle-mounted camera occlusion recognition method provided in this application directly performs frame difference statistics on the first and second images to identify regions where features such as brightness remain unchanged or change only slightly. Thus, since background modeling and clustering calculations are unnecessary, it can quickly and efficiently obtain suspected occlusion regions. Furthermore, compared to extracting candidate occlusion regions through background modeling, the technical solution in this application extracts candidate occlusion regions with higher accuracy.
[0116] The methods of the embodiments of this application are described in detail below.
[0117] In some cases, multiple frame difference images can be acquired, and the occlusion of the vehicle-mounted camera can be determined based on these multiple frame difference images. These multiple frame difference images include a first frame difference image and a second frame difference image; the first frame difference image is obtained by performing a difference operation on a first image and a second image (with an interval of X frames between the second and first images) captured by the vehicle-mounted camera at a first moment; the second frame difference image is obtained by performing a difference operation on the first image captured by the vehicle-mounted camera at the first moment and a second image preceding the first image (with an interval of Y frames between the second and first images). For example... Figure 9-1 As shown, the method provided in this application embodiment includes:
[0118] S201, Obtain vehicle speed information Vcur and ISP flag Flag_ISP.
[0119] One possible approach is to use data from one or more sensors to obtain vehicle speed information. For example, an onboard camera can collect road information and obtain the vehicle's real-time location. Based on the vehicle's location and time at different times, the vehicle speed can be calculated.
[0120] The ISP flag is used to indicate whether the ISP processing is normal. If the ISP flag is true or set (e.g., set to 1), it means the ISP processing is normal; conversely, if the ISP flag is false or not set (e.g., set to 0), it means the ISP processing is abnormal.
[0121] Typically, after the ISP acquires the raw image signal from the sensor, it processes the signal. If the processing is normal, the ISP flag is set to 1; if the processing is abnormal, the ISP flag is set to 0. The ISP flag can be stored in a processor such as processor 160. When the ISP flag needs to be retrieved, it can be obtained from processor 160 to determine whether the ISP processing is normal. The ISP flag can be stored in processor 160.
[0122] ISP processing errors include, but are not limited to, overexposure, underexposure, and color distortion. Normal ISP processing includes proper exposure and no color distortion.
[0123] S202. Determine whether the vehicle speed has reached the first threshold and whether the ISP flag is true. If the vehicle speed has reached the first threshold and the ISP flag is true, proceed to step S203. Otherwise, end this process.
[0124] The first threshold refers to a vehicle speed that is greater than or equal to the first threshold. The first threshold can be flexibly set according to the application scenario.
[0125] Typically, at lower vehicle speeds, below the first threshold, driving safety hazards may not be significant, or the performance requirements for the vehicle-mounted camera are not high. In such scenarios, it's not considered a moving scene, and detecting whether the vehicle-mounted camera is obstructed is unnecessary, reducing the workload. Furthermore, a false ISP flag often indicates that the image captured by the vehicle-mounted camera may be overexposed, underexposed, or have color distortion. In other words, the captured image cannot accurately reflect the state of the object being photographed. Consequently, due to the inability to capture an accurate image, it's impossible to accurately determine whether the vehicle-mounted camera is actually obstructed, or even if an obstruction can be determined, the result will be inaccurate. In this case, since detecting whether the vehicle-mounted camera is obstructed is difficult based on inaccurate data, it's not performed. Optionally, in this situation, an ISP or image acquisition and display anomaly alarm should be reported.
[0126] In other words, in this embodiment of the application, subsequent steps are only executed when the vehicle speed reaches the first threshold (indicating that the safety hazard is relatively prominent, the performance requirements of the vehicle-mounted camera are high, and it is in line with the motion scenario) and the ISP flag is true (meaning that the image captured by the vehicle-mounted camera is relatively accurate), so as to determine whether the vehicle-mounted camera is blocked.
[0127] By judging whether the ISP is abnormal in steps S201 and S202, abnormal scenarios such as overexposure, underexposure, and color distortion are eliminated, ensuring that the quality of the acquired image meets the requirements.
[0128] S203. Obtain the current frame FrameCur and L-1 historical frames FrameHistory.
[0129] Where L is a positive integer, greater than or equal to 2. The current frame is the image captured by the vehicle-mounted camera at the current moment. The history frames are images captured by the vehicle-mounted camera at previous moments.
[0130] As one possible implementation, the vehicle-mounted camera captures images in real time, the processor acquires the current frame image in real time, and simultaneously saves L-1 historical image frame information, resulting in a total of L frames. Each L frame includes L-1 historical frames and the current frame.
[0131] For the current frame and each of the L-1 historical frames (taking the i-th historical frame as an example), perform the following steps S204-S209:
[0132] S204. Perform inter-frame difference operation on the i-th historical frame FrameHistory and the current frame FrameCur to obtain the frame difference image FrameSub corresponding to the i-th historical frame (the frame difference image corresponding to the i-th historical frame can be simply referred to as the i-th frame difference image).
[0133] Where i is a positive integer, i is greater than or equal to 1, and less than or equal to L-1.
[0134] For details on the specific method of inter-frame differential, please refer to step S101 above.
[0135] Optional, such as Figure 9-3 As shown, the rule control module 202 includes a frame difference submodule. The current frame and the i-th historical frame are input into the frame difference submodule, and the i-th frame difference image can be output.
[0136] S205. Threshold the i-th frame difference image to obtain the FrameBinary image corresponding to the i-th historical frame (the FrameBinary image corresponding to the i-th historical frame can be simply referred to as the i-th binary image).
[0137] It can be understood that the frame difference image represents the difference between the current frame and the historical frames. Therefore, the pixels in the frame difference image with pixel values greater than a certain threshold are the pixels with larger changes between the current frame and the historical frames.
[0138] One possible implementation is to threshold the frame difference image. This can be achieved by setting a binarization threshold. For pixels in the frame difference image whose pixel value (or grayscale value) is greater than this threshold, their pixel value (or grayscale value) is set to the first pixel value, such as 0 (i.e., black). For pixels in the frame difference image whose pixel value (or grayscale value) is less than or equal to this threshold, their pixel value (or grayscale value) is set to the second pixel value, such as 255 (i.e., white). This means that pixels with larger changes in the frame difference image are set to black, and vice versa, pixels with smaller changes are set to white, resulting in a binary image. For example, this can be applied to historical frames captured by a vehicle-mounted camera. Figure 1A The current frame is subjected to frame difference operation to obtain a frame difference image, and then the frame difference image is binarized to obtain a binary image as shown. Figure 10 As shown.
[0139] The binarization threshold can be flexibly set according to the situation.
[0140] Of course, there are other ways to binarize an image, such as setting smaller pixel values to 0 and larger pixel values to 255, or other possible settings. This application does not limit the specific method of image binarization.
[0141] Optional, such as Figure 9-3 As shown, the rule control module 202 includes a binarization submodule. The i-th frame difference image is input into the binarization submodule, and the i-th binary image, that is, the binary image corresponding to the i-th frame difference image, can be output.
[0142] S206. Detect the target connected component in the i-th binary image.
[0143] First, we define a connected region. A connected region is a region consisting of pixels with the same pixel value that are adjacent (or contiguous).
[0144] Typically, the connected components may differ depending on the adjacency relationship. The following mainly uses four-adjacency and eight-adjacency as examples to illustrate the methods of dividing connected components.
[0145] like Figure 11 As shown in (1), the positional relationship between pixel p and the four pixels marked 1-4 (i.e., directly above, below, to the left, and to the right of pixel p) is called a four-adjacency relationship. In a four-adjacency relationship, some pixels are directly adjacent, such as pixel p and pixel 4. Some pixels are adjacent to each other through other intermediate pixels, such as pixel 2 and pixel 4 being adjacent through pixel p.
[0146] like Figure 11 As shown in (2), the positional relationship between pixel p and the eight pixels 1-8 directly below, to the left, to the right, and diagonally is called an eight-adjacency relationship. Similar to the four-adjacency relationship, in the eight-adjacency relationship, some pixels can be adjacent to each other through other intermediate pixels.
[0147] like Figure 11 As shown in (3), in the case of four-adjacent relationships, for the black pixel marked 2, since none of the surrounding black pixels can form a four-adjacent relationship with itself, this pixel is divided into a separate connected component. Similarly, for the pixel marked 1 indicated by the arrow and the pixel marked 3 indicated by the arrow, these two pixels cannot directly form a four-adjacent relationship, nor can they form a four-adjacent relationship with other intermediate pixels. Therefore, these two pixels are divided into two different connected components.
[0148] like Figure 11As shown in (4), in the case of eight adjacency, black pixels can form an eight adjacency relationship directly or with the help of other intermediate pixels. Therefore, all black pixels are divided into the same connected domain.
[0149] In this embodiment of the application, the connected component partitioned in the case of four adjacent nodes can be called a four-connected connected component. That is... Figure 11 The diagram (3) shows three 4-connected domains. A connected domain partitioned in the case of eight adjacencies can be called an 8-connected domain, i.e. Figure 11 The diagram shown in (4) includes one 8-connected domain.
[0150] from Figure 11 (3) or Figure 11 As can be seen from (4), the number of pixels included in an 8-connected region is usually greater than or equal to the number of pixels included in a 4-connected region. In some embodiments, 8-connected regions in a binary image can be detected, and of course, 4-connected regions in a binary image can also be detected. This application does not limit this.
[0151] The location of a connected component can be represented using coordinates. For example, the location of the connected component can be represented by the coordinates (x, y) of its top-left pixel. The size of a connected component can be represented by its length h, width w, or the number of pixels.
[0152] The target connected component refers to a connected component with an area greater than or equal to the connected component threshold and a pixel value of the second pixel value. Pixels with the second pixel value are those pixels that show the smallest change between the current frame and the i-th historical frame, i.e., pixels in the corresponding frame difference image whose pixel value is less than or equal to the first threshold.
[0153] It is understandable that when a vehicle-mounted camera is obstructed by a foreign object, the object may remain obstructed for a considerable period. This means that the positional relationship between the object and the camera is likely to remain unchanged or change only slightly. Consequently, the image formed by the object through the camera may show minimal changes over a longer period. Therefore, in this embodiment, pixels with minimal changes (including brightness changes) between the current frame and historical frames (such as...) can be detected. Figure 10 As shown, the pixel values of these pixels are set as the second pixel value in the binary image. If these pixels are detected, it indicates that the small changes in brightness and other parameters of these pixels may be due to foreign object occlusion. To further rule out false detections, it is possible to check whether the area of the connected component formed by these pixels with small changes in brightness and other parameters is large enough (i.e., whether it constitutes the target connected component). When the area of the connected component reaches a certain threshold, it indicates that the small changes in brightness and other parameters of these pixels are likely due to foreign object occlusion.
[0154] For example, suppose in a binary image, pixels that change little between the current frame and historical frames are set to white, such as... Figure 12As shown in (1), the area of the connected region formed by the white pixels is detected. If the area of connected region 1 is greater than or equal to a certain threshold, then connected region 1 can be regarded as the target connected region. The area of region 2 is smaller, so region 2 is not regarded as the target connected region.
[0155] It should be noted that the number of target connected components in the binary image corresponding to the i-th historical frame can be zero, one, or more. If there are no target connected components in the binary image corresponding to the i-th historical frame, then the operation is performed on the next historical frame after the current i-th historical frame, i.e., the (i+1)-th historical frame, i.e., jump to step S204. If there are target connected components in the binary image corresponding to the i-th historical frame, it indicates that the target connected component is likely due to an obstruction blocking the image corresponding to the vehicle-mounted lens. Therefore, in order to further confirm whether it is caused by an obstruction blocking the lens, subsequent steps need to be executed.
[0156] Optional, such as Figure 9-3 As shown, the rule control module 202 includes a connected component extraction submodule. The i-th binary image is input into the connected component extraction submodule, which can output the target connected component in the i-th binary image.
[0157] It should be noted that the positional relationship between the binary image and the target connected component can be mapped to the positional relationship between the frame difference image and the third region. In other words, if the coordinates of the target connected component in the binary image are (x, y), then the coordinates of the third region in the frame difference image are also (x, y), and the shape and size of the third region are the same as those of the target connected component.
[0158] S207. Determine the region of interest (ROI) corresponding to the target connected component, and remove the imaging areas such as sky, snow, and road surface from the ROI to obtain candidate occluded regions.
[0159] First, we introduce the method for determining the region of interest (i.e., the second region). The region of interest is an image region composed of pixels in the current frame or the i-th historical frame that share the same coordinates as the target connected component. That is, each frame corresponds to one or more regions of interest.
[0160] For example, the location of the target connected component, i.e., connected component 1, in the binary image is as follows: Figure 12 As shown in (1), the corresponding position of the region of interest in the current frame is as follows: Figure 12As shown in (2), it can be seen that the position of the target connected component in the binary image is the same as the position of the region of interest (i.e., the second region) in the current frame (i.e., the first image). In other words, the position of the third region in the frame difference image is the same as the position of the region of interest (i.e., the second region) in the current frame (i.e., the first image). The region of interest corresponds to the third region, which corresponds to the pixel region with small changes between the current frame and the i-th historical frame. This means that the region of interest is likely to include the occlusion region (i.e., the image formed by the occlusion through the vehicle lens).
[0161] Of course, even if the vehicle-mounted camera is not obstructed, a target connectivity region may still exist, and such as Figure 13 As shown, the region of interest corresponding to the connected component of the target can also be determined.
[0162] Furthermore, as mentioned above, obstructions to the vehicle's camera lens can cause minimal or no change in pixel values between images of the obstructed object for a certain period. For example, ... Figure 1C As shown, for a period of time, the foreign object remained obstructed at the same position on the vehicle-mounted lens, and the multiple images formed by the lens remained essentially unchanged. In some cases, image areas with minimal pixel value changes between the current frame and previous frames may also be due to the imaging of interfering objects. For example, imaging large areas of natural landscapes such as sky, snow, ground, and grass can also lead to image areas with little pixel value change between different frames. For example, see... Figure 14 (1) and Figure 14 (2) Due to the presence of a large area of sky, the pixel values of some image regions in the multi-frame images captured by the vehicle-mounted camera remain relatively small. For example, the pixel values in the two elliptical regions are basically close.
[0163] To further avoid misclassifying interfering objects such as the sky as occluded areas within the region of interest, the influence of these interfering objects can be eliminated. This involves removing the imaging area of a preset region (sky, snow, road surface, grass, lake, or one or more other objects) from the region of interest (i.e., the second region). In other words, the candidate occluded region (i.e., the first region) is the region obtained by removing the preset region from the region of interest (i.e., the second region).
[0164] One possible implementation is to detect specified features in the region of interest and select pixels within the region of interest that do not have the specified features as candidate occlusion regions.
[0165] Set the pixel value of the pixel with the specified feature within the region of interest to the specified pixel value.
[0166] The specified features are those of the aforementioned interfering objects. Interfering objects include, but are not limited to, natural landscapes with large imaging areas and where the pixel values of the images tend to be consistent, such as the sky, roads, snow, grasslands, and lakes.
[0167] This means that the imaging areas representing the sky, road surface, etc. in the region of interest are removed, and the remaining areas constitute the candidate occlusion areas.
[0168] The following describes a method for removing images of the sky, road surface, and other areas from the region of interest.
[0169] Taking the extraction of Region of Interest (ROI) from the current frame as an example, for instance, in the case of occlusion, the ROI extracted from the current frame would be as follows: Figure 12 As shown in (2), in the absence of obstructions, the ROI extracted in the current frame is as follows: Figure 13 The middle border is shown. As one possible implementation, removing the road surface imaging area from the ROI can be achieved by using template matching to extract areas such as... Figure 15 The shown road template is matched to the ROI by drawing windows. The template matching process can be found in [link to documentation]. Figure 16 The road surface template is slid across the entire Region of Interest (ROI) from left to right and from top to bottom. Each time the template reaches a pixel region (or sliding window), each pixel within that window is compared to the corresponding pixel in the road surface template. If the metric value between the area to be detected within the sliding window and the template meets a threshold condition, it is considered a match, meaning that the pixel region is likely an image of the road surface. Then, the pixel value of the pixels in that region is set to 0 (i.e., set to black), meaning the pixel region that is likely the road surface is set to black. For example... Figure 16 In the pixel area circled by the dashed box shown, if each pixel is only slightly different from the corresponding pixel in the road surface template, then that pixel area is set to black.
[0170] The metrics between the sliding window and the template include, but are not limited to, the difference of squares, the standard difference of squares, and correlation values.
[0171] Figure 16 The pixel region shown in the dashed box is the pixel region with the specified characteristics. This pixel region is the pixel region that is eliminated in this occlusion detection process, that is, the pixel region that is not considered as a candidate occlusion region.
[0172] In the template matching process, the sliding window size depends on the road surface template size. It should be noted that methods for detecting road surface imaging regions from images are not limited to the template matching method described above.
[0173] One possible implementation is to remove the sky image region from the ROI by setting a sky color segmentation range [r1~r2, g1~g2, b1~b2], and then performing color segmentation on all pixels in the ROI according to this range. Specifically, for a pixel in the ROI, if its r component is within the range of r1~r2, its g component is within the range of g1~g2, and its b component is within the range of b1~b2, then this pixel is considered to be an image pixel of the sky. Furthermore, this pixel is set to 0 (i.e., set to black).
[0174] like Figure 17 (2) shows from Figure 17 The ROI shown in (1) is also known as Figure 13 The result is obtained by removing the sky imaging region from the ROI shown. The black pixels represent the removed pixels, which may have been part of the sky imaging region.
[0175] As one possible approach, removing the snow-covered image area from the ROI can be achieved by: setting a color direction projection histogram of the snow (referred to as a color histogram, Hist), back-projecting the color histogram onto the ROI, finding the region in the ROI that best matches the snow image, and removing that region.
[0176] Among them, the color histogram is used to characterize the distribution of an image in the color space.
[0177] Back projection is typically used to find the best-matching point or region of a specific image (i.e., a snow image) within an input image (i.e., a Region of Interest, or ROI), thus locating the snow within the ROI. Specifically, the ROI can be divided into multiple image blocks of the same size as the snow image, and the following operation is performed on each block: the color histogram of the image block is compared with the color histogram of the snow image. If the two color histograms are similar, for example, the similarity is greater than or equal to a threshold, it indicates that the image block in the ROI is likely a snow image block, and the top-left pixel of that image block is set to 0 (i.e., set to black). In other words, the portion of the ROI suspected of containing snow is set to black. In this embodiment, the method for detecting snow images from images is not limited to back projection; if back projection is used, the method for comparing the similarity of color histograms is not limited to comparing the similarity of color histograms.
[0178] Obtaining the color histogram of an image patch within a Region of Interest (ROI) can be achieved as follows: First, quantize the colors, i.e., divide the color space (e.g., RGB (red, green, blue) color space) into multiple subspaces, each called a bin. Then, count the number of pixels in the image patch whose colors fall within each bin to obtain the color histogram of that image patch.
[0179] It should be noted that for a given historical frame or the current frame, it may correspond to one or more candidate occlusion regions.
[0180] Because rain, fog, snow, and large areas of sky or road surface can cause blurring and misjudgment, this application's embodiments use color segmentation, color histogram backprojection, and template matching to eliminate potentially occluded areas such as the sky, road surface, and snow, resulting in candidate occlusion areas and improving the accuracy of occlusion detection. This ensures that the occlusion area is caused by a real obstruction rather than a misjudgment of the natural scene. It can accurately determine the presence or absence of occlusion in scenes with complex weather changes throughout the day. Furthermore, the frame difference results also eliminate the influence of rain and fog, because in rain and fog, if the vehicle-mounted lens is not obstructed, a large area with little brightness variation will not consistently exist in the same location.
[0181] Optional, such as Figure 9-3 As shown, the rule control module 202 includes a candidate occlusion region extraction submodule. Taking the acquisition of candidate occlusion regions in the current frame as an example, the occlusion region extraction submodule can obtain the region of interest in the current frame based on the current frame and the target connected component corresponding to the current frame, and remove regions such as sky and snow from the region of interest to obtain candidate occlusion regions.
[0182] S208. Calculate the occlusion confidence of the candidate occlusion region and the area of the candidate occlusion region.
[0183] For example, the area of the candidate occlusion region (i.e., the first area corresponding to the frame difference image) can be represented by the number of pixels included in the candidate occlusion region. Other representations are also possible.
[0184] The occlusion confidence score (i.e., the first confidence score corresponding to the frame difference image) is used to represent the blurriness of the candidate occluded region (i.e., the first region) in the current frame (i.e., the first image) corresponding to the frame difference image. Typically, when an object obstructs the lens, the object is usually close to the lens, resulting in a blurry image of the occluded region. In this embodiment, a higher occlusion confidence score indicates a more blurred candidate occluded region, suggesting a higher probability that the image is formed due to an object obstructing the lens. The first confidence score is obtained based on the local binary patterns (LBP) response score, Laplacian variance, and Laplacian mean of the first region. The calculation method for the first confidence score is detailed below.
[0185] Optional, such as Figure 9-3 As shown, the rule control module 202 includes a first calculation submodule. The candidate occlusion region is input into the first calculation submodule, which can output the confidence level and area of the candidate occlusion region.
[0186] The candidate occluded regions are converted into grayscale images, and the LBP-like response score, Laplace variance, and Laplace mean of the candidate occluded regions (grayscale images) are calculated respectively. Then, the LBP-like response score, Laplace variance, and Laplace mean are weighted and summed to obtain the occlusion confidence.
[0187] Among them, the LBP-like response score can be used to characterize the texture richness of an image.
[0188] First, we introduce how the LBP response score is calculated. One possible implementation is... Figure 18 For each pixel in the candidate occlusion region (I(x, y)) (taking pixel P as an example), the following operation is performed: The pixel value of pixel P is compared with the pixel values of its 8 neighboring pixels. If the absolute difference between the pixel value of pixel P and the pixel value of a neighboring pixel is greater than or equal to a preset threshold (e.g., 3), the pixel value is set to 1; otherwise, if the absolute difference is less than the preset threshold, the pixel value is set to 0. For example, Figure 18 The pixel value of the top-left corner pixel is 5. The absolute value of the difference between its value and that of pixel P is 1, which is less than the threshold of 3. Therefore, the corresponding value of that pixel is set to 0. This process is repeated to obtain the binary string corresponding to the 8 pixels in the neighborhood of pixel P, as shown below. Figure 18 The binary string shown is 01100101. Then, compare this binary string (for example...). Figure 18 The number of 1s in the given string (01100101) is counted. If the number is greater than or equal to a preset threshold (e.g., 4), then pixel P is marked as a strong response point. Alternatively, the number of 0s in the giant binary string can be counted; if the number is less than or equal to a certain threshold, then pixel P is marked as a strong response point. A strong response point is a pixel with a large difference in pixel value from many neighboring pixels, or a pixel with rich surrounding texture. For example, a pixel with a large difference in pixel value from more than 4 neighboring pixels.
[0189] It should be noted that, as Figure 18 As shown above, the example above mainly uses the binary value of each pixel to expand in a clockwise direction to obtain a binary string. In other scenarios, there may be other expansion methods, such as expanding in a counterclockwise direction to obtain something like 11010100.
[0190] As another possible implementation, taking the labeling of pixel P as a strong or weak response point as an example, such as... Figure 18As shown, after obtaining the neighborhood binary string corresponding to pixel P, i.e., 01100101, we can determine whether pixel P is a strong response point by counting the number of numerical transitions in this binary string. Specifically, we count the number of transitions from 0 to 1 and from 1 to 0 in the binary string. When the number of transitions is greater than or equal to a preset threshold, pixel P is marked as a strong response point. Taking a preset threshold of 4 as an example, in 01100101, there are 3 transitions from 0 to 1 and 2 transitions from 1 to 0, for a total of 5 transitions (greater than the threshold of 4), so pixel P is marked as a strong response point.
[0191] It should be noted that for pixel P, either of the two methods described above can be used to determine if it is a strong response point. This can be done by counting the number of 1s in the binary string or by counting the number of transitions in the binary string. If either method determines that pixel P is a strong response point, then pixel P is marked as such. If neither method determines that pixel P is a strong response point, then it is marked as such. A weak response point refers to a pixel with sparse surrounding texture.
[0192] Similarly, referring to the method of marking pixel P as a strong or weak response point, a similar judgment is made on the remaining pixels in the candidate occluded region so that all pixels in the candidate occluded region are marked as strong or weak response points.
[0193] Next, the number of strong response points (Num_Strong) and the total number of pixels (Num_Sum) within the candidate occlusion region are calculated. Optionally, the class LBP response score (D) is calculated based on the number of strong response points (Num_Strong) and the total number of pixels (Num_Sum) according to the following formula.
[0194]
[0195] The LBP-like response score is used to characterize the blurriness and texture richness of candidate occluded regions. As can be seen from Equation 1, the fewer strong response points in the image, the lower the LBP-like response score, indicating that the pixel values of the image are more consistent, that is, the smaller the difference in pixel values between pixels, the more blurred the image and the sparser the texture.
[0196] The Laplace operator is a second-order differential algorithm. It's typically used to process images to detect varying degrees of blur or occlusion. Specifically, processing a sharp image with the Laplace operator enhances contrast, sharpens boundaries, and makes textures clearer, resulting in greater pixel value deviation. A blurry image processed with the Laplace operator produces a more uniform distribution of pixel values. This characteristic allows us to use statistical variance or mean to evaluate image blur. Specifically, sharper images produce larger variances and mean values after Laplace processing, while blurrier images produce smaller variances and mean values.
[0197] To obtain the Laplace variance and mean, firstly, the candidate occlusion region is processed using the Laplace operator to obtain the Laplace response. Then, the variance of the Laplace response is calculated, and the mean of the Laplace response is calculated, which is the Laplace mean.
[0198] First, we introduce a method for obtaining the Laplace response. One possible implementation involves convolving the candidate occlusion region I(x, y) with a Laplace kernel H(x, y) (i.e., a Laplace operator) to obtain the Laplace response G(x, y). Here, G(x, y) = I(x, y) * H(x, y). * represents the convolution operator.
[0199] The convolution kernel H(x,y) can be a two-dimensional matrix, and its size can be selected according to actual needs, for example, it can be... Figure 19 The 3x3 (i.e., 3 by 3) matrix shown
[0200] The candidate occlusion region I(x,y) can also be viewed as a two-dimensional matrix, such as Figure 19 Candidate occlusion areas shown
[0201] The process of convolving I(x,y) with H(x,y) involves, for each pixel in the I(x,y) image, calculating the product of each of its eight neighboring pixels and the corresponding element in the Laplace convolution kernel, and then summing these products to obtain the pixel value. For example... Figure 19As shown, taking the updating of pixel value Q as an example, the initial pixel value of pixel Q in the candidate occlusion region is 1, and the updated pixel value of pixel Q is 0*4+0*0+0*0+0*0+1*0+1*0+0*0+1*0+2*(-4)=-8. Referring to the method of updating the pixel value of pixel Q, H(x,y) is used to update the pixel values of the remaining pixels in the candidate occlusion region I(x,y), thus obtaining the convolution result of I(x,y) and H(x,y), which yields the Laplace response G(x,y) of the candidate occlusion region I(x,y).
[0202] Of course, there are also some boundary pixels that do not have an eight-neighborhood, such as Figure 19 The pixel R shown is an example. For these boundary pixels, boundary padding methods can be used to update their pixel values. For example, zero-padding, boundary copy padding, mirror padding, and block padding can be used to update the boundary pixel values.
[0203] Next, the variance of the Laplace response G(x,y) is calculated to obtain the Laplace variance Var_Laplace. The mean of G(x,y) is then calculated to obtain the Laplace mean Mean_Laplace.
[0204] As mentioned above, the blurrier the image, the smaller the Var_Laplace and Mean_Laplace values, while the clearer the image, the larger the Var_Laplace and Mean_Laplace values.
[0205] After calculating the Laplace variance, Laplace mean, and LBP-like response score, the weighted sum of these three factors yields the occlusion confidence score. Optionally, the occlusion confidence score can be calculated using the following formula 2:
[0206]
[0207] Here, SCORE is the occlusion confidence score, used to characterize features such as the degree of blurring and texture richness of an image. It is also used to characterize the confidence level of a candidate occluded region (the first region) being identified as an occluded region. D is the LBP-like response score, and a is the weighting coefficient of D; Var_Laplace is the Laplace variance, and b is the weighting coefficient of the Laplace variance; Mean_Laplace is the Laplace mean, and c is the weighting coefficient of the Laplace mean; N is a constant value, and N is related to the size of the candidate occluded region.
[0208] As mentioned above, the smaller the LBP-like response score, the greater the image blurriness; the smaller the Laplace mean, the greater the image blurriness; and the smaller the Laplace variance, the greater the image blurriness. Therefore, when the LBP-like response score, Laplace mean, and Laplace variance are all small, the occlusion confidence is usually high, indicating a high degree of image blurriness. In other words, by judging the level of occlusion confidence, the degree of blurriness of the candidate occluded region can be determined.
[0209] In this embodiment, the occlusion confidence score is obtained based on the Laplace mean, variance, and LBP-like response score of the candidate occlusion region as the occlusion judgment criterion. Multi-dimensional features are used to determine whether the suspected occlusion region is blurred, improving the accuracy and robustness of the blur judgment. This, in turn, enhances the robustness and accuracy of the occlusion region judgment results.
[0210] S209. Determine whether the vehicle-mounted camera is obstructed based on the occlusion confidence and area of the candidate occlusion region.
[0211] Optional, such as Figure 9-3 As shown, the rule control module 202 includes a judgment submodule. The confidence level and area are input into the judgment submodule, and the judgment result can be output, namely whether the vehicle lens is blocked and the specific blocking location.
[0212] like Figure 9-1 As shown, this application provides two judgment mechanisms applicable to occlusion judgment in different scenarios. When the candidate occlusion area corresponding to the i-th historical frame satisfies the first branch, the first judgment mechanism is used to determine whether the vehicle-mounted camera is occluded. When the candidate occlusion area corresponding to the i-th historical frame satisfies the second branch, the second judgment mechanism is used to determine whether the vehicle-mounted camera is occluded. This allows for faster identification of scenarios with large-area occlusion by foreign objects and more accurate detection of small-area occlusion. The two mechanisms for determining whether the vehicle-mounted camera is occluded are described below.
[0213] In the judgment mechanism corresponding to the first branch, a queue of length L is used to store the judgment results corresponding to the current frame and the previous L-1 frames. For a candidate occlusion region corresponding to a certain frame, it is determined whether both the occlusion confidence score and the area are greater than or equal to the first confidence threshold and the area is greater than or equal to the first area threshold. If both are true, a 1 is stored in the queue; otherwise, a 0 is stored in the queue if either condition is not met. After the candidate occlusion regions corresponding to L-1 historical frames and the current frame have been judged, the number of 1s in the L values of the queue is counted. If the number of 1s is greater than or equal to a certain threshold, the vehicle camera is determined to be occluded. Otherwise, the vehicle camera is determined not to be occluded.
[0214] It is understandable that if a candidate occlusion region in a given frame has an occlusion confidence score greater than or equal to a first confidence threshold, and its area is greater than or equal to a first area threshold, then this candidate occlusion region is considered an occlusion region, meaning the image is formed due to an object obstructing the lens. By judging candidate occlusion regions across multiple frames, the influence of misjudging a candidate occlusion region in a single frame can be avoided, thus improving the accuracy of occlusion detection results.
[0215] For example, such as Figure 20 As shown, assuming the threshold is 4, the number of 1s in the queue is 6 (greater than the threshold 4), indicating that the candidate occlusion area corresponds to most frames, thus confirming that the vehicle camera is occluded.
[0216] In other embodiments, the judgment mechanism corresponding to the second branch is typically used to determine large-area occlusion. In some scenarios, due to the large size of the occluding object or other factors, the area of the image captured by the occluding lens (i.e., the occluded area) is large, and the occluded area usually appears in consecutive frames, making it less prone to fluctuation. Based on this characteristic, in this embodiment, a frame counter is set. When the area of the candidate occluded area is greater than or equal to a second area threshold, and the occlusion confidence of the candidate occluded area is greater than or equal to the second confidence threshold, the frame counter is incremented by one; otherwise, the counter is set to 0. If the counter is greater than or equal to the threshold, it indicates that an occluded area exists in consecutive frames. In this case, it is determined that the vehicle-mounted lens is occluded.
[0217] In the second judgment mechanism, when the vehicle-mounted camera is largely obstructed by a foreign object, for example, when... Figure 21 The occlusion shown can be detected and identified more quickly through continuous cumulative judgment.
[0218] In this embodiment of the application, the queue-based cumulative decision method can smooth out abnormal situations such as intermittent brightness changes and improve the robustness of the judgment results.
[0219] Optionally, an alarm mechanism can be triggered after it is determined that the vehicle-mounted camera is wholly or partially obstructed. For example, autonomous driving can be stopped, and the driver can take over driving.
[0220] In other cases, multiple frame difference images can be acquired, and the occlusion of the vehicle-mounted camera can be determined based on these multiple frame difference images. These multiple frame difference images include a first frame difference image and a second frame difference image; the first frame difference image is obtained by performing a difference operation on a first image captured by the vehicle-mounted camera at a first moment and a second image preceding the first image; the second frame difference image is obtained by performing a difference operation on a first image captured by the vehicle-mounted camera at a second moment and a second image preceding the first image, wherein the first moment and the second moment are different.
[0221] like Figure 9-2 As shown, the method provided in this application embodiment includes:
[0222] S301, obtain vehicle speed information and ISP identifier.
[0223] This step can be found in step S201 above.
[0224] S302. Determine whether the vehicle speed has reached the first threshold and whether the ISP flag is true. If the vehicle speed has reached the first threshold and the ISP flag is true, proceed to step S303. Otherwise, end this process.
[0225] This step can be found in step S202 above.
[0226] S303, Obtain the current frame and historical frames.
[0227] The current frame is the first image, and the historical frame is the second image.
[0228] It should be noted that in this embodiment of the application, multiple frame difference images can be acquired, and the vehicle lens can be determined as to be blocked based on the multiple frame difference images. Figure 9-2 This example uses one frame difference image from multiple frame difference images for illustration. In one example, the current frame at a first moment and a historical frame X frames away from the current frame can be obtained, and the frame difference image between the current frame and the historical frame can be calculated. Similarly, the current frame at a second moment and a historical frame X frames away from the current frame can be obtained, and the frame difference image between the current frame and the historical frame can be calculated. X is a positive integer. This process continues, obtaining M (M is a positive integer) frame difference images, and then using these M frame difference images to perform subsequent processes to determine whether the vehicle-mounted camera is obstructed.
[0229] In some embodiments, a queue of length X can be used to store X historical frames. For the current frame at a given moment, a frame difference image is calculated using historical frames that are X frames away from the current frame and the current frame itself. After calculating the frame difference image, the current frame is inserted at the end of the queue, and the corresponding calculated historical frames are dequeued. In this way, the current frame becomes a historical frame for subsequent moments. The same operation can be performed on current frames at other moments, and historical frames that are X frames away from the current frames at other moments. In this way, the current frame at a given moment that has completed calculation can be inserted into the queue in real time to become a historical frame for subsequent moments, and the calculated historical frames can be deleted from the queue to update the elements in the queue. This ensures that subsequent calculations are completed using a shorter queue, effectively saving storage space.
[0230] S304. Perform inter-frame difference operation on the historical frame and the current frame to obtain the frame difference image.
[0231] S305. Threshold the frame difference image to obtain a binary image.
[0232] S306. Detect the target connected component in a binary image.
[0233] S307. Determine the region of interest corresponding to the target connected component, and remove the imaging areas such as sky, snow, and road surface from the region of interest to obtain candidate occlusion areas.
[0234] S308. Calculate the occlusion confidence of the candidate occlusion region and the area of the candidate occlusion region.
[0235] S309. Determine whether the vehicle-mounted camera is obstructed based on the occlusion confidence and area of the candidate occlusion region.
[0236] As one possible implementation, if a total of M frame difference images are acquired, and P of these M frame difference images satisfy a first condition, then it is determined that the vehicle-mounted camera is occluded. The first condition is that the confidence level is greater than or equal to a first confidence threshold, and the first area is greater than or equal to a first area threshold; M and P are positive integers, with M greater than or equal to P. Alternatively, if Q consecutive frame difference images among the M frame difference images satisfy a second condition, then it is determined that the vehicle-mounted camera is occluded; the second condition is that the confidence level is greater than or equal to a second confidence threshold, and the first area is greater than or equal to a second area threshold; M and Q are positive integers, with M greater than or equal to Q.
[0237] For example, see Figure 9-2 If the number of 1s in the queue is greater than P, it means that there are occlusion areas in all P frames. In this case, it is determined that the vehicle camera is occluded. For example, in another example, if the frame counter reaches Q, it means that there are occlusion areas in all Q consecutive frames. In this case, it is determined that the vehicle camera is occluded.
[0238] For the specific implementation of steps S303-S309, please refer to steps S203-S209 above.
[0239] This application embodiment can divide the vehicle lens occlusion recognition device into functional modules according to the above method example. When each functional module is divided according to its corresponding function, Figure 22 This diagram illustrates a possible structure of the vehicle-mounted lens occlusion recognition device described in the above embodiments. Figure 22 As shown, the vehicle-mounted lens occlusion recognition device includes an acquisition module 401 and a judgment module 402. Of course, the vehicle-mounted lens occlusion recognition device may also include other functional modules, or it may include even fewer functional modules.
[0240] The acquisition module 401 is used to acquire a frame difference image; the frame difference image is an image obtained by performing a difference operation on a first image and a second image captured by the vehicle-mounted camera; the frame interval between the first image and the second image is N, where N is a positive integer; neither the first image nor the second image is a background image.
[0241] The judgment module 402 is used to determine whether the vehicle-mounted camera is blocked based on the frame difference image.
[0242] In one possible design, the frame difference image is one or more frame difference images, including a first frame difference image and a second frame difference image; the first frame difference image is an image obtained by performing a difference operation on a first image captured by the vehicle-mounted lens at a first moment and a second image preceding the first image; the second frame difference image is an image obtained by performing a difference operation on a first image captured by the vehicle-mounted lens at a second moment and a second image preceding the first image, wherein the first moment and the second moment are different.
[0243] In one possible design, the acquisition module 401 is further configured to acquire a first confidence level and a first area corresponding to one or more frame difference images;
[0244] Wherein, the confidence score corresponding to a frame difference image is used to represent the degree of confidence that the first region in the first image corresponding to the frame difference image is identified as an occluded region, and the confidence score is obtained based on the LBP response score, Laplacian variance, and Laplacian mean of the first region; the first area corresponding to the frame difference image is the area of the first region in the first image corresponding to the frame difference image.
[0245] The first region is the region obtained by removing the preset region from the second region; the second region is included in the first image, and the position of the second region in the first image is the same as the position of the third region in the frame difference image corresponding to the first image; the third region is the region composed of pixels in the frame difference image whose pixel value is less than or equal to a first threshold; the preset region includes any one or more of the following: the imaging region of the road surface, the imaging region of the snow, the imaging region of the sky, and the imaging region of the grass.
[0246] In one possible design, there are M frame difference images; the judgment module 402 is used to determine whether the vehicle lens is blocked based on the frame difference images, including: if there are P frame difference images among the M frame difference images that satisfy a first condition, then determine that the vehicle lens is blocked.
[0247] Wherein, the first condition is that the confidence level is greater than or equal to the first confidence threshold, and the first area is greater than or equal to the first area threshold; M and P are positive integers, and M is greater than or equal to P.
[0248] In one possible design, there are M frame difference images; the judgment module 402 is used to determine whether the vehicle lens is blocked based on the frame difference images, including: if Q consecutive frame difference images in the M frame difference images satisfy a second condition, then determine that the vehicle lens is blocked.
[0249] The second condition is that the confidence level is greater than or equal to the second confidence threshold, and the first area is greater than or equal to the second area threshold; M and Q are positive integers, and M is greater than or equal to Q.
[0250] In one possible design, the acquisition module 401 is used to acquire a frame difference image, including: acquiring vehicle speed and / or a first identifier bit; the first identifier bit is used to identify whether the image processing result of the image signal processing ISP is normal; if the first identifier bit indicates that the image processing result of the ISP is normal, and / or the vehicle speed is greater than or equal to a vehicle speed threshold, then the frame difference image is acquired.
[0251] See Figure 23 This application also provides a vehicle-mounted camera lens occlusion recognition device, including a processor 510 and a memory 520. The processor 510 and the memory 520 are connected (e.g., interconnected via a bus 540).
[0252] Optionally, the vehicle-mounted camera occlusion detection device may also include a transceiver 530, which is connected to the processor 510 and the memory 520. The transceiver is used to receive / send data, such as data from other vehicles.
[0253] Processor 510 can execute Figure 8 , Figure 9-1 , Figure 9-2 The operation of any corresponding implementation scheme and its various feasible implementation methods. For example, the operation of the acquisition module 401, the judgment module 402, and / or other operations described in the embodiments of this application.
[0254] The processor 510 described above (or described as a controller) can implement or execute various exemplary logic blocks, unit modules, and circuits described in conjunction with the disclosure of this application. The processor or controller can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, unit modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computing functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0255] Bus 540 can be an extended industry standard architecture (EISA) bus, etc. Bus 540 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 23 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0256] For details on the specific working processes of the processor, memory, bus, and transceiver, please refer to the above text; they will not be repeated here.
[0257] This application also provides a vehicle-mounted camera lens occlusion recognition device, including a non-volatile storage medium and a central processing unit (CPU). The non-volatile storage medium stores an executable program, and the CPU is connected to the non-volatile storage medium and executes the executable program to implement the present application, for example... Figures 8-9-1 , Figure 9-2 The method for recognizing occlusion of vehicle-mounted cameras is shown.
[0258] Another embodiment of this application provides a computer-readable storage medium including one or more program codes, the one or more programs including instructions, wherein when a processor executes the program code, the vehicle lens occlusion recognition device performs the following... Figure 8 , Figure 9-1 , Figure 9-2 The method for recognizing occlusion of vehicle-mounted cameras is shown.
[0259] In another embodiment of this application, a computer program product is also provided, comprising computer-executable instructions stored in a computer-readable storage medium. At least one processor of the vehicle-mounted lens occlusion recognition device can read the computer-executable instructions from the computer-readable storage medium, and the at least one processor executes the computer-executable instructions to cause the vehicle-mounted lens occlusion recognition device to perform its functions. Figure 8 , Figure 9-1 , Figure 9-2 The corresponding steps in the vehicle-mounted camera occlusion recognition method shown.
[0260] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0261] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, it can appear, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated.
[0262] A computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL0)) or wireless (e.g., infrared, wireless, microwave, etc.) means. A computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. This available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks, SSDs), etc.
[0263] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0264] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0265] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0266] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0267] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0268] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for identifying occlusion of a vehicle-mounted camera lens, characterized in that, include: Obtain a frame difference image; the frame difference image is an image obtained by performing a difference operation on a first image and a second image captured by the vehicle-mounted camera, and the frame difference image is one or more frame difference images; the frame interval between the first image and the second image is N, where N is a positive integer; neither the first image nor the second image is a background image. Obtain the first confidence level and the first area corresponding to one or more frame difference images; Wherein, the confidence score corresponding to a frame difference image is used to represent the degree of confidence that the first region in the first image corresponding to the frame difference image is identified as an occluded region, and the confidence score is obtained based on the LBP response score, Laplacian variance, and Laplacian mean of the first region; the first area corresponding to the frame difference image is the area of the first region in the first image corresponding to the frame difference image. The first region is the region obtained by removing a preset region from the second region; the second region is included in the first image, and the position of the second region in the first image is the same as the position of the third region in the frame difference image corresponding to the first image; the third region is the region composed of pixels in the frame difference image whose pixel value is less than or equal to a first threshold; the preset region includes any one or more of the following: the imaging region of the road surface, the imaging region of the snow, the imaging region of the sky, and the imaging region of the grass. The vehicle-mounted camera is determined to be obstructed based on the first confidence level and the first area.
2. The vehicle-mounted lens occlusion recognition method according to claim 1, characterized in that, The plurality of frame difference images include a first frame difference image and a second frame difference image; the first frame difference image is an image obtained by performing a difference operation on a first image captured by the vehicle-mounted lens at a first moment and a second image preceding the first image; the second frame difference image is an image obtained by performing a difference operation on a first image captured by the vehicle-mounted lens at a second moment and a second image preceding the first image, wherein the first moment and the second moment are different.
3. The vehicle-mounted lens occlusion recognition method according to claim 1 or 2, characterized in that, The frame difference images are M; Determining whether the vehicle-mounted camera is obstructed based on the frame difference image includes: If P frame difference images out of the M frame difference images satisfy the first condition, then it is determined that the vehicle-mounted camera is blocked. Wherein, the first condition is that the confidence level is greater than or equal to the first confidence threshold, and the first area is greater than or equal to the first area threshold; M and P are positive integers, and M is greater than or equal to P.
4. The vehicle-mounted lens occlusion recognition method according to claim 1 or 2, characterized in that, The frame difference images are M; Determining whether the vehicle-mounted camera is obstructed based on the frame difference image includes: If Q consecutive frame difference images among the M frame difference images satisfy the second condition, then it is determined that the vehicle-mounted camera is blocked; The second condition is that the confidence level is greater than or equal to the second confidence threshold, and the first area is greater than or equal to the second area threshold; M and Q are positive integers, and M is greater than or equal to Q.
5. The vehicle-mounted lens occlusion recognition method according to claim 1 or 2, characterized in that, The acquisition of the frame difference image includes: Obtain vehicle speed and / or a first identifier bit; the first identifier bit is used to indicate whether the image processing result of the image signal processing ISP is normal. If the first identifier indicates that the image processing result of the ISP is normal, and / or the vehicle speed is greater than or equal to the vehicle speed threshold, then the frame difference image is acquired.
6. A vehicle-mounted lens occlusion recognition device, characterized in that, include: The acquisition module is used to acquire frame difference images; The frame difference image is an image obtained by performing a difference operation on the first image and the second image captured by the vehicle-mounted camera. The frame difference image is one or more frame difference images. The frame interval between the first image and the second image is N, where N is a positive integer. Neither the first image nor the second image is a background image. The judgment module is used to determine whether the vehicle-mounted camera is obstructed based on the frame difference image; The acquisition module is further configured to acquire a first confidence level and a first area corresponding to one or more frame difference images; Wherein, the confidence score corresponding to a frame difference image is used to represent the degree of confidence that the first region in the first image corresponding to the frame difference image is identified as an occluded region, and the confidence score is obtained based on the LBP response score, Laplacian variance, and Laplacian mean of the first region; the first area corresponding to the frame difference image is the area of the first region in the first image corresponding to the frame difference image. The first region is the region obtained by removing the preset region from the second region; the second region is included in the first image, and the position of the second region in the first image is the same as the position of the third region in the frame difference image corresponding to the first image; the third region is the region composed of pixels in the frame difference image whose pixel value is less than or equal to a first threshold; the preset region includes any one or more of the following: the imaging region of the road surface, the imaging region of the snow, the imaging region of the sky, and the imaging region of the grass.
7. The vehicle-mounted lens occlusion recognition device according to claim 6, characterized in that, The plurality of frame difference images include a first frame difference image and a second frame difference image; the first frame difference image is an image obtained by performing a difference operation on a first image captured by the vehicle-mounted lens at a first moment and a second image preceding the first image; the second frame difference image is an image obtained by performing a difference operation on a first image captured by the vehicle-mounted lens at a second moment and a second image preceding the first image, wherein the first moment and the second moment are different.
8. The vehicle-mounted lens occlusion recognition device according to claim 6 or 7, characterized in that, The frame difference images are M; The judgment module is used to determine whether the vehicle lens is blocked based on the frame difference image, including: if there are P frame difference images among the M frame difference images that satisfy the first condition, then determine that the vehicle lens is blocked. Wherein, the first condition is that the confidence level is greater than or equal to the first confidence threshold, and the first area is greater than or equal to the first area threshold; M and P are positive integers, and M is greater than or equal to P.
9. The vehicle-mounted lens occlusion recognition device according to claim 6 or 7, characterized in that, The frame difference images are M; the judgment module is used to determine whether the vehicle lens is blocked based on the frame difference images, including: if Q consecutive frame difference images in the M frame difference images satisfy a second condition, then determine that the vehicle lens is blocked. The second condition is that the confidence level is greater than or equal to the second confidence threshold, and the first area is greater than or equal to the second area threshold; M and Q are positive integers, and M is greater than or equal to Q.
10. The vehicle-mounted lens occlusion recognition device according to claim 6 or 7, characterized in that, The acquisition module is used to acquire a frame difference image, including: acquiring vehicle speed and / or a first identifier bit; the first identifier bit is used to indicate whether the image processing result of the image signal processing ISP is normal; if the first identifier bit indicates that the image processing result of the ISP is normal, and / or the vehicle speed is greater than or equal to a vehicle speed threshold, then the frame difference image is acquired.
11. A vehicle-mounted lens occlusion recognition device, characterized in that, include: The device includes a processor, a memory, and a communication interface; wherein the communication interface is used to communicate with other devices or communication networks, the memory is used to store one or more programs, the one or more programs including computer-executable instructions, and when the device is running, the processor executes the computer-executable instructions stored in the memory to cause the device to perform the vehicle lens occlusion recognition method as described in any one of claims 1-5.
12. A computer-readable storage medium, characterized in that, The system includes programs and instructions, which, when run on a computer, implement the vehicle lens occlusion recognition method as described in any one of claims 1-5.
13. A computer program product containing instructions, characterized in that, When the computer program product is run on a computer, the computer performs the vehicle lens occlusion recognition method as described in any one of claims 1-5.
14. A chip system, characterized in that, The system includes a processor coupled to a memory, the memory storing program instructions, which, when executed by the processor, implement the vehicle lens occlusion recognition method according to any one of claims 1-5.
Citation Information
Patent Citations
Camera blocking area detection method and device, equipment and storage medium
CN111932596A