An image processing method and related device
By combining neural networks with region of interest analysis, the problem of vehicle image acquisition quality under severe weather conditions has been solved, the accuracy of weather condition determination has been improved, and vehicle safety has been ensured.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2022-09-30
- Publication Date
- 2026-07-31
AI Technical Summary
The quality of images captured by vehicles in adverse weather conditions is affected, leading to a decrease in target detection accuracy and potentially causing safety issues.
The first neural network generates predicted weather conditions and regions of interest for each image. The comprehensive analysis determines the weather conditions of the surrounding environment. Multiple cameras are used to reduce the influence of lens obstructions and improve accuracy.
It improves the accuracy of determining the surrounding weather conditions, reduces errors caused by camera obstructions, and ensures safe vehicle operation.
Smart Images

Figure CN115661770B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to an image processing method and related equipment. Background Technology
[0002] During vehicle operation, images of the surrounding environment can be captured by the first camera, and then target detection can be performed on the captured images. However, the quality of the images captured by the first camera can be affected by severe weather (such as rain, snow, or other weather conditions), which will affect the accuracy of the target detection process. For example, the target detection results obtained based on images captured in severe weather may have problems such as false detection and missed detection.
[0003] If the vehicle's driving path is still planned according to the above target detection results, it may cause safety problems. Therefore, a solution to determine the weather conditions of the surrounding environment is urgently needed. Summary of the Invention
[0004] This application provides an image processing method and related equipment, which integrates the predicted weather state of each first image generated by a first neural network and at least one region of interest in the first image to finally determine the weather state of the surrounding environment, thereby improving the accuracy of the final determined weather state.
[0005] To address the aforementioned technical problems, the embodiments of this application provide the following technical solutions:
[0006] In a first aspect, embodiments of this application provide an image processing method applicable to the image processing field within the field of artificial intelligence. The method includes: an execution device acquiring at least one first image, wherein the at least one first image is captured by at least one first camera of the surrounding environment; the execution device generating first information through a first neural network, the first information indicating the predicted weather state of each first image, and determining the weather state of the surrounding environment based on the first information and second information, wherein the second information indicates at least one region of interest in each first image, and the first region of interest within the at least one region of interest includes a region associated with the predicted weather state of the first image. Exemplarily, the regions in the first image associated with multiple weather states may include any one or more of the following: blurred regions, water droplet regions, muddy water regions, icy regions, snowflake regions, or other regions associated with weather states.
[0007] This application provides a scheme for determining the weather state of the surrounding environment. After acquiring a first image of the surrounding environment, not only is a predicted weather state for each first image generated by a first neural network, but also at least one region of interest in each first image is acquired. The first region of interest in the at least one region of interest includes a region caused by an obstruction of the first camera and which is related to the predicted weather state of the first image. By combining the predicted weather state of each first image generated by the first neural network and the at least one region of interest in the first image, the weather state of the surrounding environment is finally determined, which helps to improve the accuracy of the finally determined weather state.
[0008] Optionally, the number of first images is at least two, and the at least two first images contain images from different first cameras; the aforementioned different first cameras can be first cameras located at different positions, and the orientation of the first cameras at different positions can be the same or different. In this application, since the area in a first image that is related to the predicted weather state may be caused by the surrounding environment, for example, if the surrounding environment is rainy, it will result in the presence of raindrop areas in the first image; however, the area in a first image that is related to the predicted weather state may also be caused by obstructions on the lens of the first camera, for example, if there are water droplets on the lens of a first camera, it will also result in the presence of raindrop areas in the first image. Obstructions on the lens of the first camera may reduce the accuracy of the prediction result of the weather state of the surrounding environment. Therefore, the multiple first images obtained are taken by different first cameras. Since the probability of the same obstruction on the lenses of different first cameras is low, determining the weather state of the surrounding environment based on multiple first images taken by different first cameras is beneficial to improving the accuracy of the finally determined weather state.
[0009] Optionally, the execution device determines the weather state of the surrounding environment based on the first information and the second information, which may include: the execution device determining parameter information corresponding to each first region of interest based on the second information, and determining the weather state of the surrounding environment based on the first information and the parameter information corresponding to the first region of interest; wherein, the parameter information includes any one or more of the following: the coverage of the first region of interest in the first image, the area of the first region of interest, the brightness of the first region of interest, or the blur of the first region of interest. Optionally, the greater the coverage of each first region of interest in the first image, the greater the probability that the weather state of the surrounding environment is determined by the predicted weather state of the first image. When at least one first image includes at least two first images taken by the same first camera, the area of the same type of first region of interest (hereinafter referred to as "target region of interest") in each of the aforementioned at least two first images can be obtained. If the target region of interest varies more in the aforementioned at least two first images, the greater the probability that the weather state of the surrounding environment is determined by the predicted weather state of the first image. In the case where a certain first region of interest is a raindrop, the greater the brightness of the first region of interest, the greater the rainfall in the surrounding environment. For example, when the first region of interest is a raindrop, the greater the ambiguity of the first region of interest, the greater the thickness of the raindrop, which also indicates a greater amount of rainfall in the surrounding environment.
[0010] In this application, several parameter information corresponding to the first region of interest are determined based on the second information, which reduces the feasibility of this solution. It can not only utilize the coverage of the first region of interest in the first image, but also the area, brightness, and blur of the first region of interest, so as to utilize each first region of interest in the first image from more dimensions, which is conducive to obtaining a more accurate weather condition of the surrounding environment.
[0011] Optionally, each first image may include an image region of the surrounding environment of the first camera and an image region of the obstruction of the first camera; the region in each first image that is related to the weather state may include a weather-related region in the image region of the surrounding environment of the first camera, and a weather-related region in the image region of the obstruction of the first camera.
[0012] Optionally, the execution device generates first information via a first neural network, which may include: the execution device predicts the weather state of the second image via the first neural network to obtain the first information, wherein the second image includes an image region in the first image other than a second region of interest in at least one region of interest, and the second region of interest includes a region caused by an obstruction of the first camera and which is related to the weather state. "An obstruction of the first camera" may include an obstruction directly located on the lens of the first camera, and an obstruction suspended outside the first camera, such as an obstruction located between the first camera and the surrounding environment.
[0013] In this application, since the obstructions of the first camera can include obstructions on the lens of the first camera and obstructions suspended outside the first camera, the area caused by the obstructions of the first camera and related to the weather conditions may be caused by obstructions on the lens of the first camera, such as raindrops, mud, ice, or other coverings on the lens of the first camera; it may also be caused by obstructions suspended outside the first camera, for example, if the surrounding environment of the first camera is rainy, then the obstructions suspended outside the first camera may also have raindrops. Therefore, the area in the first image caused by the obstructions of the first camera and related to the weather conditions may be caused by the weather conditions of the surrounding environment or by the first camera itself. By removing the area in the first image caused by the obstructions of the first camera and related to the weather conditions, a second image is obtained. The weather conditions of the second image are then predicted using a first neural network. The second image can more realistically reflect the weather conditions of the surrounding environment, thereby improving the accuracy of the first information output by the first neural network.
[0014] Optionally, when the executing device is a vehicle, at least one second region of interest in at least one region of interest in the first image may further include the region in the first image caused by the vehicle. For example, the second region of interest in at least one region of interest may include any one or more of the following: blurred region, water droplet region, muddy region, icy region, snowflake region, or glare region. This application provides specific examples of the types of regions that the second region of interest in at least one region of interest may represent, which reduces the feasibility of this solution.
[0015] Optionally, the method may further include: the execution device determining a third region of interest in a first image obtained by a second camera based on second information, wherein the second camera belongs to at least one first camera, the first image obtained by the second camera is included in at least one first image, and the third region of interest includes the area in the first image obtained by the second camera caused by an occlusion of the second camera; and determining the degree of failure of the second camera based on the area in the first image obtained by the second camera caused by an occlusion of the second camera. Optionally, the execution device determining the degree of failure of the second camera based on the area in the first image obtained by the second camera caused by an occlusion of the second camera may include: the execution device determining the coverage of all areas caused by occlusions of the second camera lens in each first image based on all areas caused by occlusions of the second camera lens in each first image obtained by the second camera; and determining the degree of failure of the second camera based on the coverage of all areas caused by occlusions of the second camera lens in each first image.
[0016] In this application, the area caused by the occlusion of the second camera in the first image obtained by the second camera (i.e., any one of the first cameras) can also be determined based on the second information. Then, the degree of failure of the second camera can be determined based on the area caused by the occlusion of the second camera in the first image obtained by the second camera. Since the image obtained by the camera with a high degree of failure has low reliability, timely detection of the camera with a high degree of failure is beneficial to improving the quality of the first image obtained, and thus beneficial to improving the accuracy of the determined weather conditions of the surrounding environment.
[0017] Secondly, embodiments of this application provide an image processing method applicable to the image processing field within the field of artificial intelligence. The method may include: an execution device acquiring at least one first image, each first image being captured by a first camera of the surrounding environment; the execution device acquiring second information, wherein the second information indicates at least one region of interest in each first image, the at least one region of interest including a second region of interest, the second region of interest including an area caused by an obstruction of the first camera and an area correlated with weather conditions; and predicting the weather conditions of the second image using a first neural network to obtain first information, the second image including image areas in the first image excluding the second region of interest, the first information indicating the predicted weather conditions of each first image, and the first information being used to determine the weather conditions of the surrounding environment.
[0018] In this application, since the obstructions of the first camera can include obstructions on the lens of the first camera and obstructions suspended outside the first camera, the area caused by the obstructions of the first camera and related to the weather conditions may be caused by obstructions on the lens of the first camera, such as raindrops, mud, ice, or other coverings on the lens of the first camera; it may also be caused by obstructions suspended outside the first camera, for example, if the surrounding environment of the first camera is rainy, then the obstructions suspended outside the first camera may also have raindrops. Therefore, the area in the first image caused by the obstructions of the first camera and related to the weather conditions may be caused by the weather conditions of the surrounding environment or by the first camera itself. By removing the area in the first image caused by the obstructions of the first camera and related to the weather conditions, a second image is obtained. The weather conditions of the second image are then predicted using a first neural network. The second image can more realistically reflect the weather conditions of the surrounding environment, thereby improving the accuracy of the first information output by the first neural network.
[0019] Optionally, the method may further include: the execution device determining the weather state of the surrounding environment based on first information and a first region of interest in at least one region of interest, wherein the first region of interest in at least one region of interest includes a region that is associated with the predicted weather state of the first image.
[0020] Optionally, the execution device determines the weather state of the surrounding environment based on the first information and the first region of interest in at least one region of interest, which may include: the execution device determining parameter information corresponding to the first region of interest based on the second information, wherein the parameter information includes any one or more of the following: the coverage of the first region of interest in the first image, the area of the first region of interest, the brightness of the first region of interest, or the blur of the first region of interest; the execution device determines the weather state of the surrounding environment based on the first information and the parameter information corresponding to the first region of interest.
[0021] In the second aspect of this application, the execution device can also be used to execute the steps executed by the execution device in the first aspect and various possible implementations of the first aspect. The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in the various possible implementations of the second aspect can all be found in the first aspect, and will not be repeated here.
[0022] Thirdly, embodiments of this application provide a method for training a neural network, applicable to image processing within the field of artificial intelligence. The method may include: a training device acquiring a training image, which is obtained by capturing the surrounding environment using a camera; and generating second information through a first neural network, wherein the second information indicates the predicted location of at least one region of interest in the training image, the at least one region of interest including a region caused by camera occlusion, or, alternatively, a region caused by a vehicle. The training device trains the first neural network according to a first loss function term, the first loss function term indicating the similarity between the second information and first expected information, the first expected information indicating the correct location of at least one region of interest in the training image.
[0023] In this application, since camera obstructions can include obstructions on the camera lens and obstructions suspended outside the camera, such as raindrops, mud, ice, or other coverings on the camera lens, or raindrops on obstructions suspended outside the camera if the surrounding environment is rainy, the areas in the image caused by camera obstructions will affect the observation of the actual surrounding environment. The first neural network trained using this scheme can identify the areas in the image caused by camera obstructions, which is beneficial for the execution device of the trained first neural network to understand the surrounding environment more accurately.
[0024] Optionally, after the training device generates the second information through the first neural network, the method may further include: the training device predicting the weather state of the second image through the first neural network to obtain the first information, wherein the second image includes image regions in the training image other than at least one region of interest, and the first information indicates the predicted weather state of each training image. The training device training the first neural network based on a first loss function term may include: the training device training the first neural network based on a first loss function term and a second loss function term, wherein the second loss function term indicates the similarity between the first information and the second expected information, and the second expected information indicates the correct weather state of the training image.
[0025] In the third aspect of this application, the training device can also be used to perform the steps of the first aspect and the various possible implementations of the first aspect. The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in the various possible implementations of the third aspect can be found in the first aspect, and will not be repeated here.
[0026] Fourthly, embodiments of this application provide an image processing apparatus that can be used in the field of image processing within the field of artificial intelligence. The apparatus includes: an acquisition module for acquiring at least one first image, wherein the at least one first image is obtained by capturing the surrounding environment through at least one first camera; a generation module for generating first information through a first neural network, wherein the first information indicates the predicted weather state of each first image; and a determination module for determining the weather state of the surrounding environment based on the first information and second information, wherein the second information indicates at least one region of interest in each first image, and the first region of interest in the at least one region of interest includes a region that is associated with the predicted weather state of the first image.
[0027] In the fourth aspect of this application, the image processing apparatus can also be used to perform the steps of the execution device in the first aspect and various possible implementations of the first aspect. The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in the various possible implementations of the fourth aspect can be found in the first aspect, and will not be repeated here.
[0028] Fifthly, embodiments of this application provide an image processing apparatus that can be used in the field of image processing within the field of artificial intelligence. The apparatus includes: an acquisition module for acquiring at least one first image, each first image being captured by a first camera of the surrounding environment; an acquisition module for acquiring second information, wherein the second information indicates at least one region of interest in each first image, the at least one region of interest including a second region of interest, the second region of interest including a region caused by an obstruction of the first camera and a region related to weather conditions; and a processing module for predicting the weather conditions of the second image using a first neural network to obtain first information, wherein the second image includes image regions in the first image other than the second region of interest, the first information indicating the predicted weather conditions of each first image, and the first information being used to determine the weather conditions of the surrounding environment.
[0029] In the fifth aspect of this application, the image processing apparatus can also be used to perform the steps of the execution device in the second aspect and various possible implementations of the second aspect. The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in the various possible implementations of the fifth aspect can be found in the second aspect, and will not be repeated here.
[0030] Sixthly, embodiments of this application provide a neural network training apparatus that can be used in the field of image processing within the field of artificial intelligence. The apparatus includes: an acquisition module for acquiring a training image, wherein the training image is obtained by capturing the surrounding environment through a camera; a processing module for generating second information through a first neural network, wherein the second information indicates the predicted position of at least one region of interest in the training image, and the at least one region of interest includes a region caused by an occlusion of the camera; and a training module for training the first neural network according to a first loss function term, wherein the first loss function term indicates the similarity between the second information and first expected information, and the first expected information indicates the correct position of at least one region of interest in the training image.
[0031] According to the sixth aspect of this application, the training device for the neural network can also be used to perform the steps of the execution device in the third aspect and various possible implementations of the third aspect. The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in the various possible implementations of the sixth aspect can be found in the third aspect, and will not be repeated here.
[0032] In a seventh aspect, embodiments of this application provide a computer program product, which includes a program that, when run on a computer, causes the computer to perform the methods described in the first to third aspects above.
[0033] Eighthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods described in the first to third aspects above.
[0034] Ninthly, embodiments of this application provide an execution device, including a processor and a memory, the processor being coupled to the memory, the memory being used to store a program; the processor being used to execute the program in the memory, causing the execution device to perform the image processing method described in the first or second aspect above.
[0035] In a tenth aspect, embodiments of this application provide a training device, including a processor and a memory, wherein the processor is coupled to the memory, the memory is used to store a program, and the processor is used to execute the program in the memory, such that the training device performs the neural network training method described in the third aspect above.
[0036] Eleventhly, this application provides a chip system including a processor for supporting an execution device or training device in implementing the functions involved in the foregoing aspects, such as transmitting or processing data and / or information involved in the foregoing methods. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the terminal device or communication device. This chip system may be composed of chips or may include chips and other discrete devices. Attached Figure Description
[0037] Figure 1 A schematic diagram of the main framework of artificial intelligence provided in the embodiments of this application;
[0038] Figure 2a A system architecture diagram of an image processing system provided in an embodiment of this application;
[0039] Figure 2b This is a schematic flowchart of an image processing method provided in an embodiment of this application;
[0040] Figure 3 Another schematic flowchart of the image processing method provided in the embodiments of this application;
[0041] Figure 4 A schematic diagram illustrating the position of at least one first camera provided in an embodiment of this application;
[0042] Figure 5 A schematic diagram showing the relationship between the weather state and the first image provided in the embodiments of this application;
[0043] Figure 6 A schematic diagram of an area caused by an obstruction of the first camera and related to weather conditions, provided in an embodiment of this application;
[0044] Figure 7 A comparative schematic diagram of the first and second images provided in the embodiments of this application;
[0045] Figure 8 Two schematic diagrams of the first and second images provided for embodiments of this application;
[0046] Figure 9 A schematic flowchart illustrating a neural network training method provided in an embodiment of this application;
[0047] Figure 10 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this application;
[0048] Figure 11 This is another schematic diagram of the image processing apparatus provided in the embodiments of this application;
[0049] Figure 12 A schematic diagram of a neural network training device provided in an embodiment of this application;
[0050] Figure 13 A schematic diagram of the structure of the execution device provided in the embodiments of this application;
[0051] Figure 14 A schematic diagram of the structure of a training device provided in an embodiment of this application;
[0052] Figure 15 This is a schematic diagram of a chip structure provided in an embodiment of this application. Detailed Implementation
[0053] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0054] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0055] First, the overall workflow of the artificial intelligence system is described; please refer to [link / reference]. Figure 1 , Figure 1 The diagram illustrates a structural framework for artificial intelligence (AI). The framework is further elaborated below along two dimensions: the "Intelligent Information Chain" (horizontal axis) and the "IT Value Chain" (vertical axis). The "Intelligent Information Chain" reflects a series of processes from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a condensation process of "data—information—knowledge—wisdom." The "IT Value Chain" reflects the value that AI brings to the information technology industry, from the underlying infrastructure of human intelligence and information (provided and processed through technological means) to the industrial ecosystem of the system.
[0056] (1) Infrastructure
[0057] The infrastructure provides computing power to support artificial intelligence systems, enabling communication with the external world and providing support through a basic platform. Communication with the outside world is achieved through sensors; computing power is provided by intelligent chips, which can specifically employ hardware acceleration chips such as central processing units (CPUs), embedded neural network processing units (NPUs), graphics processing units (GPUs), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). The basic platform includes distributed computing frameworks and related platform guarantees and support, which may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to acquire data, and this data is provided to intelligent chips in the distributed computing system provided by the basic platform for computation.
[0058] (2) Data
[0059] The data at the next layer of infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, and IoT data from traditional devices, including business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0060] (3) Data processing
[0061] Data processing typically includes methods such as data training, machine learning, deep learning, search, reasoning, and decision-making.
[0062] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, and training on data, including symbolization and formalization.
[0063] Reasoning refers to the process in which, in a computer or intelligent system, the machine thinks and solves problems by simulating human intelligent reasoning, based on reasoning control strategies and using formalized information. Typical functions include search and matching.
[0064] Decision-making refers to the process of making decisions based on intelligent information after reasoning, and it typically provides functions such as classification, sorting, and prediction.
[0065] (4) General ability
[0066] After the data processing mentioned above, the results of the data processing can be used to form some general capabilities, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0067] (5) Smart Products and Industry Applications
[0068] Intelligent products and industry applications refer to products and applications of artificial intelligence systems in various fields. They encapsulate overall artificial intelligence solutions, productize intelligent information decision-making, and realize practical applications. Their application areas mainly include: intelligent terminals, intelligent manufacturing, intelligent transportation, smart homes, intelligent healthcare, intelligent security, autonomous driving, and smart cities.
[0069] This application can be used in various scenarios for obtaining surrounding weather conditions. Specifically, it can be used to determine the weather conditions of the surrounding environment based on images of the surrounding environment. For example, in the field of autonomous driving, images of the surrounding environment can be collected by cameras configured on the vehicle, and the weather conditions can be determined based on the collected images. As another example, in the field of intelligent security, cameras with monitoring functions are installed to collect images of the surrounding environment, and the weather conditions can be determined based on the collected images. As yet another example, in the field of intelligent transportation, multiple cameras can be installed on roads to collect images of the surrounding environment, and the weather conditions can be determined based on the collected images. As yet another example, in the field of intelligent manufacturing, cameras can also be configured on base stations to collect images of the surrounding environment, and the weather conditions can be determined based on the collected images, and so on. The application scenarios of this application are not exhaustively listed here.
[0070] Before providing a detailed description of the image processing method provided in the embodiments of this application, please refer to [the relevant documentation]. Figure 2a , Figure 2a This is a system architecture diagram of an image processing system provided in an embodiment of this application. Figure 2a In the image processing system 200, there are training devices 210, database 220, execution devices 230 and data storage system 240, and the execution devices 230 include computing modules 231.
[0071] The database 220 stores a training dataset. During the training phase, the training device 210 generates a first model / rule 201 and iteratively trains the first model / rule 201 using the training dataset to obtain the trained first model / rule 201. The first model / rule 201 can be specifically represented as a neural network or as a non-neural network model. In this embodiment, the example of the first model / rule 201 being represented as a neural network is used for illustration.
[0072] The first model / rule 201 obtained by the training device 210 can be applied to the execution device 230. The execution device 230 can call data, code, etc., from the data storage system 240, and can also store data, instructions, etc., in the data storage system 240. The data storage system 240 can be located within the execution device 230, or it can be an external memory relative to the execution device 230. In the application phase, the execution device 230 can use the first model / rule 201 to execute the image processing method provided in the embodiments of this application. For details, please refer to... Figure 2b , Figure 2b This is a schematic flowchart of an image processing method provided in an embodiment of this application. A1. The execution device 230 acquires at least one first image, which is obtained by capturing the surrounding environment using at least one first camera. A2. The execution device 230 generates first information using a first model / rule 201, the first information indicating the predicted weather state for each first image. A3. The execution device 230 determines the weather state of the surrounding environment based on the first information and second information, the second information indicating at least one region of interest in each first image. The first region of interest in the at least one region of interest includes a region related to the predicted weather state of the first image. For example, the region in the first image related to the weather state may include any one or more of the following: blurred regions, water droplet regions, muddy water regions, icy regions, snowflake regions, or other regions related to the weather state, etc., and is not exhaustively listed here. It should be understood that... Figure 2b The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0073] In some embodiments of this application, please refer to Figure 2a The execution device 230 can be configured in the client device, and the "user" can interact directly with the execution device 230. For example, when the client device is a vehicle, the execution device 230 can be a module in the vehicle's host CPU used for image processing. The execution device 230 can also be a graphics processing unit (GPU) or neural network processor (NPU) in the vehicle. The GPU or NPU is mounted on the host processor as a coprocessor, and the host processor allocates tasks.
[0074] It is worth noting that Figure 2a This is merely a schematic diagram of the architecture of an image processing system provided by an embodiment of the present invention. The positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in some other embodiments of this application, the execution device 230 and the client device can be separate independent devices. The execution device 230 is configured with an input / output (I / O) interface to interact with the client device. The "user" can input a first image through the client device, and the client device sends the first image to the execution device 230 through the I / O interface. After the execution device 230 determines the weather state of the surrounding environment through the first machine learning model / rule 201 in the calculation module 231, it can return the aforementioned weather state of the surrounding environment to the client device through the I / O interface and provide it to the user.
[0075] Based on the above description, the specific implementation process of the application stage and training stage of the first neural network provided in the embodiments of this application will be described in detail below.
[0076] I. Application Phase
[0077] In the embodiments of this application, please refer to Figure 3 , Figure 3 This is another schematic flowchart illustrating the image processing method provided in this application embodiment. The image processing method provided in this application embodiment may include:
[0078] 301. Acquire at least one first image, wherein the at least one first image is obtained by taking a picture of the surrounding environment with at least one first camera.
[0079] In this embodiment, the execution device can acquire one or more first images, which are captured by at least one first camera of the surrounding environment. The first camera can be configured on the execution device; for example, if the execution device is a vehicle, the first camera can be a camera mounted on the vehicle. Alternatively, the first camera can be communicatively connected to the execution device; for example, if the first camera is a camera in a monitoring system, and the execution device is a data processing device that processes images acquired by multiple cameras in the monitoring system. This embodiment does not limit the positional relationship between the first camera and the execution device. Any first camera can be of any type, such as a standard camera, a fisheye camera, or other types of cameras, etc., without exhaustive list. Any first camera can be a camera that includes a cleaning device or a camera that does not include a cleaning device; for example, the cleaning device can be any combination of one or more of the following: a windshield wiper, a washer, a dryer, or other cleaning devices, etc., without exhaustive list.
[0080] If the executing device acquires multiple first images, in one scenario, these multiple first images can be captured by a single first camera at different times, observing the surrounding environment. The orientation of this single first camera at these different times can be the same or different. In another scenario, the multiple first images can be captured by at least two first cameras. For example, the number of at least two first cameras can be 2, 3, 5, 6, 7, or other numbers, etc., and is not limited here. The at least two first cameras are located at different positions, and the orientations of the first cameras at different positions can be the same or different. For a more intuitive understanding of this solution, please refer to [link to relevant documentation]. Figure 4 , Figure 4 This is a schematic diagram showing the position of at least one first camera provided in an embodiment of this application. Figure 4 Taking vehicles as an example of the execution equipment in China and Israel, such as Figure 4 As shown, at least two first cameras may include any combination of one or more of the following cameras: a front-view pinhole camera, a front-view fisheye camera, a left-view pinhole camera, a left-view fisheye camera, a right-view pinhole camera, a right-view fisheye camera, a rear-view pinhole camera, and a rear-view fisheye camera. It should be understood that... Figure 4 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0081] In this embodiment, since the area in a first image that is related to the predicted weather state may be caused by the surrounding environment, for example, if the surrounding environment is rainy, it will result in the presence of raindrop areas in the first image; however, the area in a first image that is related to the predicted weather state may also be caused by obstructions on the lens of the first camera, for example, if there are water droplets on the lens of a first camera, it will also result in the presence of raindrop areas in the first image. Obstructions on the lens of the first camera may reduce the accuracy of the prediction result of the weather state of the surrounding environment. Therefore, the multiple first images are taken by different first cameras. Since the probability of the same obstruction on the lenses of different first cameras is low, determining the weather state of the surrounding environment based on multiple first images taken by different first cameras is beneficial to improving the accuracy of the final determined weather state.
[0082] The at least two first cameras capture images of the surrounding environment at one or more identical moments, resulting in at least two first images. For example, the at least two first cameras include camera 1 and camera 2. Both camera 1 and camera 2 capture images of the surrounding environment at moment A, resulting in image 1 and image 2. Both camera 1 and camera 2 then capture images of the surrounding environment at moment B, resulting in image 3 and image 4. The at least two first images include image 1, image 2, image 3, and image 4. Therefore, the at least two first cameras (i.e., camera 1 and camera 2) capture images of the surrounding environment at the same time. It should be understood that this example is only for the convenience of understanding the solution and is not intended to limit the solution.
[0083] When the execution device needs to monitor the weather conditions of the surrounding environment, each first camera can capture a first image of the surrounding environment at a first frequency and transmit the image to the processor of the execution device at the same first frequency. That is, the processor of the execution device can acquire the first image sent by each first camera once within each first time period. Optionally, the execution device can perform the weather condition determination operation of the surrounding environment at a second frequency, meaning the execution device can perform the weather condition determination operation of the surrounding environment once within each second time period. Further optionally, the second frequency can be an integer multiple of the first frequency. For example, if the length of the first time period is 1 second, then the length of the second time period can be 1 second, 2 seconds, or other lengths, etc.; or, for example, if the length of the first time period is 30 milliseconds, then the length of the second time period can be 30 milliseconds, 1 second, 2 seconds, or other lengths, etc. It should be noted that the frequency at which each first camera sends the first image to the processor of the execution device and the frequency at which the execution device determines the weather condition of the surrounding environment can be flexibly set according to the actual application scenario. The examples here are only for the convenience of understanding this solution and are not intended to limit this solution.
[0084] 302. Obtain second information, the second information indicating at least one region of interest in each first image, the at least one region of interest including a region that is related to the weather state.
[0085] In this embodiment of the application, after the execution device acquires at least one first image of the surrounding environment, it can acquire second information. The second information indicates at least one region of interest in each first image, that is, the second information indicates the position of each region of interest in each first image.
[0086] Wherein, at least one region of interest may include at least one fourth region of interest, each fourth region of interest referring to a region in the first image that is associated with S weather states, where S is an integer greater than or equal to 1. For example, the regions in the first image associated with weather states may include any one or more of the following: blurred regions, water droplet regions, muddy regions, icy regions, snowflake regions, or other regions associated with weather states, etc., which are not exhaustively listed here. Further, each first image may include an image region of the surrounding environment of the first camera and an image region of the first camera's obstructions; the regions in each first image associated with weather states may include weather-related regions in the image region of the first camera's surrounding environment, and weather-related regions in the image region of the first camera's obstructions. For example, the image region of the surrounding environment in the first image may include green leaves, and the water droplet region in the image region of the surrounding environment included in the first image may include the aforementioned water droplets on the green leaves; as another example, the image region of the surrounding environment in the first image may include a road, and the muddy region in the image region of the surrounding environment included in the first image may include the aforementioned muddy water in the road, etc. These examples are only for the convenience of understanding this scheme and are not intended to limit this scheme.
[0087] For a more intuitive understanding of this solution, please refer to [link / reference]. Figure 5 , Figure 5 This is a schematic diagram illustrating the relationship between the first image provided in this application and the weather conditions. Figure 5 This diagram includes two sub-diagrams, C1 and C2. C1 is an example of the first image. C1 shows the image area of the surrounding environment in the first image, including the image areas containing objects such as roads, tall buildings, and trees. C2 refers to the area in the image area of the first camera obstructing the view in the first image shown in C1 that is related to the weather conditions. Specifically, it refers to the raindrop area on the lens in the first image shown in C1. The raindrop area on the lens is the obstruction directly located on the lens of the first camera. Figure 5 The examples in the text are only for the convenience of understanding the concepts of "the area in the image region of the first camera that is related to the weather state" and "the area in the image region of the surrounding environment in the first image that is related to the weather state" and are not intended to limit this scheme.
[0088] "Obstructions to the first camera" can include obstructions directly on the lens of the first camera, such as water droplets, ice, mud, or other obstructions on the first camera, etc., which are not exhaustive examples here. "Obstructions to the first camera" can also include obstructions suspended outside the first camera, such as obstructions located between the first camera and the surrounding environment; for example, if the first camera is configured on a vehicle and the surrounding environment of the first camera is the environment around the vehicle, when the weather condition of the surrounding environment of the first camera is light rain, the obstruction to the first camera can be raindrops between the lens of the first camera and the surrounding environment of the vehicle; as another example, if the first camera is a camera in a monitoring system, when the weather condition of the surrounding environment of the first camera is light snow, the obstruction to the first camera can be snowflakes between the lens of the first camera and the surrounding environment of the first camera, etc. It should be understood that the examples here are only for the convenience of understanding this solution and are not intended to limit this solution.
[0089] Optionally, each fourth region of interest may include an area obstructed by the first camera and correlated with weather conditions. For a more intuitive understanding of this scheme, please refer to [link to relevant documentation]. Figure 6 , Figure 6 This is a schematic diagram of an area caused by an obstruction of the first camera and which is related to the weather conditions, as provided in an embodiment of this application. Figure 6 This includes two sub-illustrative diagrams, D1 and D2. D1 shows an example of the first image, while D2 refers to the water droplet area in the first image shown in D1 caused by obstructions from the first camera (i.e., an example of an area related to weather conditions). The obstructions shown in both D1 and D2 are located between the first camera and the vehicle's surroundings, specifically the water droplet area on the vehicle's windshield. It should be understood that... Figure 6 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0090] Optionally, at least one region of interest in the first image may include at least one second region of interest, and the at least one second region of interest may include regions caused by M types of occlusions from the first camera, where M is an integer greater than or equal to 1. The aforementioned M types of occlusions may only include occlusions related to weather conditions, such as raindrops, mud, snowflakes, frost, or other occlusions related to weather conditions, etc., and are not exhaustively listed here. Optionally, the aforementioned M types of occlusions may also include occlusions not related to weather conditions, such as paper occlusions, plastic occlusions, or other occlusions not related to weather conditions, etc., and are not exhaustively listed here.
[0091] It should be noted that if each fourth region of interest includes a region caused by an obstruction of the first camera and is related to the weather conditions, and all M types of obstructions are obstructions related to the weather, then at least one second region of interest and at least one fourth region of interest refer to the same region in the first image, which is a region caused by an obstruction of the first camera and is related to the weather conditions.
[0092] If each fourth region of interest includes a region caused by obstructions from the first camera and is related to the weather conditions; and if the M types of obstructions include not only obstructions related to the weather conditions but also obstructions not related to the weather conditions, then all types of fourth regions of interest are subsets of at least one second region of interest.
[0093] If each fourth region of interest is not required to be caused by an obstruction of the first camera, the second information may include a first sub-information and a second sub-information, wherein the first sub-information is used to indicate the position of each fourth region of interest in the first image within at least one fourth region of interest, and the second sub-information is used to indicate the position of each second region of interest in the first image within at least one second region of interest.
[0094] Optionally, when the executing device is a vehicle, at least one second region of interest in at least one region of interest in the first image may also include regions in the first image caused by the vehicle. For example, regions in the first image caused by the vehicle may include glare regions in the first image caused by the vehicle's lights, or other regions caused by the vehicle, etc., which are not exhaustively listed here. Exemplarily, the second region of interest in at least one region of interest includes any one or more of the following: blurred regions, water droplet regions, muddy regions, icy regions, snowflake regions, glare regions, or other types of regions, etc. The specific regions set as the second region of interest can be flexibly determined according to the actual situation, and are not limited here. In the embodiments of this application, the specific types of regions represented by the second region of interest in at least one region of interest are provided, which reduces the feasibility of this solution.
[0095] Specifically, step 302 may include: the execution device inputting the first image into the second neural network to obtain second information output by the second neural network. The second information indicates the position of each region of interest in each first image within at least one region of interest. The second neural network is a neural network that has already undergone training. The training process of the second neural network will be described in subsequent embodiments and will not be elaborated here. The second neural network can be a neural network used to perform object detection tasks. For example, the second neural network can be a convolutional neural network, a residual neural network, or other neural networks used to perform object detection tasks, etc., and will not be exhaustively listed here.
[0096] 303. Generate first information through a first neural network, the first information indicating the predicted weather state of each first image.
[0097] In this embodiment, the execution device generates first information through a first neural network, indicating the predicted weather state of each first image. The first information may include the probability that the first image generated by the first neural network belongs to each of N weather states, where N is an integer greater than or equal to S. For example, the N weather states may include normal, light rain, heavy rain, and others; or, for example, normal, rainy, snowy, sleet, and others. It should be understood that the specific weather states represented by the N weather states can be flexibly determined according to the actual application scenario. The examples given here are only for the convenience of understanding this solution and are not intended to limit this solution. The first neural network may specifically be a convolutional neural network, a residual neural network, or other types of neural networks, etc., and is not limited here.
[0098] In one implementation, step 303 may include: the execution device may input each of the at least one first image into a first neural network to obtain first information output by the first neural network; optionally, the at least one first image comes from different first cameras. Optionally, before inputting each first image into the first neural network, the execution device may also perform scaling processing on the first image so that the different first images have the same size; the execution device may also perform other processing on the first image before inputting each first image into the first neural network, which will not be elaborated here.
[0099] Optionally, the second neural network and the first neural network are the same neural network. The execution device inputs the first image into the first neural network (i.e., into the second neural network), and the first neural network generates second information. The execution device uses the first neural network to predict the weather state of the second image to obtain the first information. The second image includes the image region in the first image excluding the second region of interest (ROI) within at least one ROI. A detailed explanation of the second ROI within at least one ROI can be found in step 302 above, and will not be repeated here.
[0100] To understand this solution more intuitively, please refer to [link / reference]. Figure 7 , Figure 7 This is a comparative schematic diagram of a first image and a second image provided for an embodiment of this application. Figure 7 Includes two sub-diagrams, left and right. Figure 7 The left sub-image in the diagram represents the first image. Figure 7 The right sub-image in the diagram represents the second image. Figure 7Taking at least one region of interest as an example, the second region of interest includes an area caused by obstructions from the first camera and that is related to weather conditions (i.e.) Figure 7 The left sub-illustration shows the raindrop area caused by the obstruction of the first camera. The second image includes the image area in the first image excluding the raindrop area caused by the obstruction of the first camera. Figure 7 As can be seen from the left and right sub-illustrations, after removing the image area outside the raindrop area caused by the obstruction of the first camera in the first image, it is clear that there is no rain in the surrounding environment. This should be understood. Figure 7 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0101] In another implementation, the second neural network and the first neural network are different neural networks. Step 303 may include: the execution device generating at least one second image corresponding one-to-one with each of the at least one first image based on the second information and each of the at least one first image. Each second image includes an image region in each first image other than a second region of interest in at least one region of interest. The second region of interest includes a region caused by occlusion of the first camera and a region related to the weather state. Optionally, at least one first image may contain images from different first cameras. The execution device may input the second image into the first neural network to obtain the first information output by the first neural network.
[0102] For a more intuitive understanding of this solution, please refer to [link / reference]. Figure 8 , Figure 8 Two schematic diagrams of the first and second images provided for embodiments of this application. Figure 8The diagram includes four sub-illustrations: F1, F2, F3, and F4. F1 shows the first image captured by the vehicle's forward-facing camera. F2 represents the second image corresponding to the first image shown in F1. F2 includes the image area in F1 excluding the area caused by obstructions to the first camera and related to weather conditions (i.e., the blurred and icy areas caused by obstructions to the first camera in F1). F3 shows the first image captured by the vehicle's left-facing camera. F4 represents the second image corresponding to the first image shown in F3. F4 includes the image area in F3 excluding the area caused by obstructions to the first camera and related to weather conditions (i.e., the icy areas caused by obstructions to the first camera in F3). Since the second image shown in F2 carries relatively little useful information, the first neural network has difficulty accurately predicting the weather conditions of the surrounding environment based on the second image shown in F2. Removing the icy area caused by the obstruction of the first camera in F3 yields the second image shown in F4. Based on the second image shown in F4, the first neural network can predict the weather conditions of the surrounding environment as sunny. Figure 8 As can be seen from the images shown, the steps of "removing areas in the first image caused by obstructions from the first camera and related to weather conditions" and "acquiring multiple first images from different first cameras" help improve the accuracy of the obtained weather conditions of the surrounding environment. It should be understood that... Figure 8 The examples in this document are for illustrative purposes only and are not intended to limit the scope of this solution.
[0103] In this embodiment, since the obstruction of the first camera may include obstructions on the lens of the first camera and obstructions suspended outside the first camera, the area caused by the obstruction of the first camera and related to the weather condition may be caused by obstructions on the lens of the first camera, such as raindrops, mud, ice, or other coverings on the lens of the first camera; it may also be caused by obstructions suspended outside the first camera, for example, if the surrounding environment of the first camera is rainy, then the obstructions suspended outside the first camera may also have raindrops; therefore, the area in the first image caused by the obstruction of the first camera and related to the weather condition may be caused by the weather condition of the surrounding environment or it may be caused by the first camera itself. By removing the area in the first image caused by the obstruction of the first camera and related to the weather condition, a second image is obtained, and the weather condition of the second image is predicted using a first neural network. The second image can more realistically reflect the weather condition of the surrounding environment, thereby improving the accuracy of the first information output by the first neural network.
[0104] 304. Based on the first and second information, determine the weather conditions of the surrounding environment.
[0105] In this embodiment, the execution device can determine the weather state of the surrounding environment within a first time period based on the first information and the second information. Specifically, step 304 may include: after acquiring the predicted weather state of the first image, the execution device can determine at least one first region of interest (ROI) related to the predicted weather state of the first image from at least one ROI indicated by the second information. That is, it can determine at least one first ROI related to the predicted weather state of the first image from at least one fourth ROI. The concept of a first ROI can be found in the above description of the fourth ROI, and will not be repeated here. The execution device determines parameter information corresponding to the first ROI based on the second information. The parameter information includes any one or more of the following: the coverage of the first ROI in the first image, the area of the first ROI, the brightness of the first ROI, or the blurriness of the first ROI. Based on the first information and the parameter information corresponding to the first ROI, the execution device determines the weather state of the surrounding environment within the first time period.
[0106] For example, if the predicted weather condition of the first image is rainy, such as light rain, moderate rain, or heavy rain, then at least one first region of interest (ROI) associated with light rain, moderate rain, or heavy rain may include blurred regions, raindrop regions, muddy regions, or other regions associated with rainy weather. As another example, if the predicted weather condition of the first image is snowy, then at least one ROI associated with snowy weather may include snowflake regions, icy regions, muddy regions, blurred regions, or other regions associated with snowy weather. It should be noted that when the predicted weather condition of the first image is other weather conditions, the at least one ROI associated with the predicted weather condition of the first image can be other types of ROI. These examples are only for the convenience of understanding this scheme and are not intended to limit this scheme.
[0107] The first information may include the probability that each first image generated by the first neural network belongs to each of the N weather states. Optionally, the execution device can determine a first weather state and a first probability value based on the probability that each first image in at least one of the first images belongs to each of the N weather states. The first probability value represents the probability that the surrounding weather obtained based on the first neural network is the first weather state. If the first probability value is greater than or equal to a preset probability value, the execution device can determine the first weather state as the weather state of the surrounding environment; if the first probability value is less than the preset probability value, the execution device can determine the weather state of the surrounding environment based on the parameter information corresponding to the first region of interest. For example, the preset probability value can be 0.7, 0.75, 0.8, or other values, etc., and can be flexibly set according to the actual situation, which is not limited here.
[0108] The greater the coverage of each first region of interest in the first image, the greater the probability that the weather state of the surrounding environment is based on the predicted weather state of the first image. When at least one first image includes at least two first images taken by the same first camera, the area of the same type of first region of interest (hereinafter referred to as "target region of interest") in each of the aforementioned at least two first images can be obtained. If the target region of interest varies more in the aforementioned at least two first images, the probability that the weather state of the surrounding environment is based on the predicted weather state of the first image is greater. For example, the at least one first image acquired by the execution device includes image 1, image 2, and image 3 from camera 1. The predicted weather state of the first image is light rain. At least one first region of interest associated with light rain may include raindrop regions. If the raindrop regions vary more in image 1, image 2, and image 3, the probability that the weather state of the surrounding environment is light rain is greater. It should be understood that this example is only for the convenience of understanding this solution and is not intended to limit this solution.
[0109] For example, when the first region of interest is a raindrop, the greater the brightness of the first region of interest, the greater the rainfall in the surrounding environment. Similarly, when the first region of interest is a raindrop, the greater the blurriness of the first region of interest, the greater the thickness of the raindrop, which also represents a greater rainfall in the surrounding environment.
[0110] For example, this example uses N weather conditions, including light rain, heavy rain, normal, and others, and the parameter information corresponding to the first region of interest is based on the coverage rate of each region of interest. Normal weather conditions refer to weather conditions that do not affect the normal operation of the vehicle's intelligent driving system. For example, normal weather conditions can include sunny days, cloudy days, or other types of weather, etc., which are not exhaustively listed here. The execution device can pre-set S coverage rate thresholds corresponding to S weather conditions. Here, a first coverage rate threshold corresponding to light rain and a second coverage rate threshold corresponding to heavy rain are pre-set. For example, the value of the first coverage rate threshold can be 0.05, and the value of the second coverage rate threshold can be 0.35. The specific implementation method of "determining the weather conditions of the surrounding environment based on the first information and the parameter information corresponding to the first region of interest" is described below with reference to Table 1.
[0111]
[0112] Table 1
[0113] As shown in Table 1, the execution device can determine whether a first probability value is greater than or equal to a preset probability value. If the first probability value is greater than or equal to the preset probability value, the execution device can determine the first weather state as the weather state of the surrounding environment. If the first probability value is less than the preset probability value, the execution device acquires at least one first region of interest (ROI) that is related to the first weather state. It can determine whether the coverage rate of each ROI is greater than or equal to the coverage rate threshold corresponding to the first weather state. If the coverage rate of any ROI is greater than or equal to the coverage rate threshold corresponding to the first weather state, the weather state of the surrounding environment can be determined as the first weather state. If the coverage rate of each ROI is less than the coverage rate threshold corresponding to the first weather state, the weather state of the surrounding environment can be determined as normal. For example, as shown in Table 1, when the first weather condition is light rain, the coverage rate of each first region of interest (ROI) associated with light rain (e.g., raindrop region, blurred region, and muddy region in Table 1) is determined to be greater than or equal to 0.05 (i.e., the first coverage rate threshold). If the coverage rate of any first ROI is greater than or equal to 0.05, the weather condition of the surrounding environment is determined to be light rain; if the coverage rate of each first ROI is less than 0.05, the weather condition of the surrounding environment is determined to be normal. As another example, as shown in Table 1, when the first weather condition is heavy rain, the coverage rate of each first ROI associated with heavy rain (e.g., raindrop region, blurred region, and muddy region in Table 1) is determined to be greater than or equal to 0.35 (i.e., the second coverage rate threshold). If the coverage rate of any first ROI is greater than or equal to 0.35, the weather condition of the surrounding environment is determined to be heavy rain; if the coverage rate of each first ROI is less than 0.35, the weather condition of the surrounding environment is determined to be normal.
[0114] It should be noted that the description above, in conjunction with Table 1, is merely an example for ease of understanding this solution. In the process of finally determining the weather conditions of the surrounding environment, the device can compare the coverage rate of each first region of interest with a coverage threshold, or it can compare the coverage rate of all first regions of interest within at least one region of interest with a coverage threshold. Both N and S weather conditions can include more or fewer weather conditions, and the specific settings can be flexibly configured according to the actual application scenario; no limitations are imposed here.
[0115] In this embodiment, several parameter information corresponding to the first region of interest are determined based on the second information, which reduces the feasibility of this solution. It can not only utilize the coverage of the first region of interest in the first image, but also the area, brightness, and blur of the first region of interest, so as to utilize each first region of interest in the first image from more dimensions, which is conducive to obtaining a more accurate weather condition of the surrounding environment.
[0116] Regarding the implementation process of the execution device "determining a first weather state and a first probability value based on the probability that each first image belongs to each of N weather states", one implementation can use a voting method to determine the first weather state and the first probability value. Specifically, the execution device can first determine a predicted weather state for each first image based on the probability that each first image belongs to each of the N weather states, thereby obtaining T predicted weather states that correspond one-to-one with at least one first image (hereinafter referred to as "T first images" for convenience). The T predicted weather states may contain the same weather state. The execution device can determine the weather state that appears most frequently among the T predicted weather states as the first weather state. For example, T is 5, and the five first images are image 1, image 2, image 3, image 4 and image 5. Based on the probability that image 1 belongs to each of the N weather states, the predicted weather state of image 1 is determined to be normal. Based on the probability that image 2 belongs to each of the N weather states, the predicted weather state of image 2 is determined to be light rain. Based on the probability that image 3 belongs to each of the N weather states, the predicted weather state of image 3 is determined to be other. Based on the probability that image 4 belongs to each of the N weather states, the predicted weather state of image 4 is determined to be normal. Based on the probability that image 5 belongs to each of the N weather states, the predicted weather state of image 5 is determined to be normal. That is, the T predicted weather states are normal, light rain, other, normal and normal. The voting results are shown in Table 2 below.
[0117] normal Light rain other 3 1 1
[0118] Table 2
[0119] As shown in Table 2 above, the weather state that appears most frequently among the T predicted weather states is normal. Therefore, the first weather state can be determined as normal. It should be understood that the example here is only for the convenience of understanding this scheme and is not intended to limit this scheme.
[0120] The execution device can also obtain a second probability value for each first image belonging to a first weather state based on the probability of each first image belonging to each of the N weather states. That is, it obtains T second probability values corresponding one-to-one with T first images. The first probability value can be obtained by weighted summation of the aforementioned T second probability values. Optionally, the execution device can obtain the first probability value by averaging the aforementioned T probability values.
[0121] In another implementation, for any one of the N weather states (hereinafter referred to as the "target weather state" for convenience), the execution device obtains the probability that each first image belongs to the target weather state based on the probability that each first image belongs to each of the N weather states. This results in T third probability values corresponding to the T first images. A weighted sum of these T third probability values yields the probability value corresponding to the target weather state. The execution device repeats this process to obtain N probability values corresponding to the N weather states. The execution device can then determine the highest probability value among these N probability values as the first probability value and select a first weather state from the N weather states that corresponds to this first probability value.
[0122] It should be noted that the execution device can also use other schemes to achieve "determining the first weather state and the first probability value based on the probability that each first image in at least one first image belongs to each of the N weather states". The example here is only to demonstrate the feasibility of this scheme and is not intended to limit this scheme.
[0123] Regarding the implementation method of the execution device "determining the parameter information corresponding to the first region of interest based on the second information", this explanation will first take the coverage of the first region of interest as an example. For the specific methods of obtaining the "area of the first region of interest", "brightness of the first region of interest", and "blurriness of the first region of interest", please refer to the method of obtaining the "coverage of the first region of interest", which will not be elaborated here. Specifically, for any type of first region of interest (hereinafter referred to as the "target first region of interest") in the first image, if the first image is obtained in step 301, the coverage parameter in the parameter information corresponding to the target first region of interest can be directly determined based on the position of the target first region of interest in the first image.
[0124] In the case that at least two first images are obtained in step 301 (i.e., T is greater than or equal to 2), in one implementation, the execution device can obtain the first coverage rate of the target first region of interest in each first image, and obtain T first coverage rates. The execution device can perform a weighted summation of the T first coverage rates to obtain the coverage parameter in the parameter information corresponding to the target first region of interest.
[0125] In another implementation, if at least two first images come from different first cameras, the coverage rate of the target first region of interest in each first image captured by the first camera can be determined first. Then, the coverage rates of the target first region of interest in each first image captured by the first camera are weighted and summed to obtain the coverage parameter in the parameter information corresponding to the target first region of interest. The reference factors for the weighting of different first cameras can include the importance of each first camera. For example, the location of a first camera, the frequency of use of images captured by a first camera, or other factors can all serve as reference factors for the importance of that first camera.
[0126] For example, at least two first images are obtained from cameras 1, 2, and 3. Camera 1 is a forward-looking camera, camera 2 is a left-side camera, and camera 3 is a right-side camera. The weights of cameras 1, 2, and 3 are 0.4, 0.3, and 0.3, respectively. The coverage of the raindrop region (an example of the target's first region of interest) in the first image obtained by camera 1 is 0.06, the coverage of the raindrop region in the first image obtained by camera 2 is 0.04, and the coverage of the raindrop region in the first image obtained by camera 3 is 0.05. The weighted sum of 0.06, 0.04, and 0.05 yields the coverage parameter in the parameter information corresponding to the target's first region of interest. That is, the value of the coverage parameter in the parameter information corresponding to the raindrop region is 0.051. It should be understood that this example is only for the convenience of understanding this scheme and is not intended to limit this scheme.
[0127] Furthermore, regarding the specific method for obtaining the "coverage rate of the first region of interest in each first image acquired by the first camera," if at least two first images include at least two first images from the same first camera (hereinafter referred to as the "target first camera" for ease of description), the execution device can obtain the coverage rate of the first region of interest in each first image acquired by the target first camera. The execution device can perform a weighted average of the coverage rates of the first region of interest in each first image acquired by the target first camera to obtain the coverage rate of the first region of interest in the first image acquired by the target first camera. Alternatively, the execution device can also determine the coverage rate of the first region of interest in each first image acquired by the target first camera as the median of the coverage rates of the first region of interest in each first image acquired by the target first camera. Alternatively, the execution device can also determine the coverage rate of the first region of interest in each first image acquired by the target first camera using a moving average, exponential smoothing, or other methods. The specific method used can be flexibly determined based on the actual situation and is not limited here.
[0128] Optionally, the execution device can determine the final weather state of the surrounding environment within the first time period based on the weather state of the surrounding environment within at least one second time period prior to the first time period, and the weather state of the surrounding environment within the first time period. The length of each second time period can be the same as the length of the first time period. Since the weather state of the surrounding environment changes gradually, i.e., the weather state of the surrounding environment is continuous, in order to reduce the probability of the execution device incorrectly determining the weather state, one or more switching conditions can be set on the execution device. When the execution device determines that the weather state within the first time period needs to switch from the weather state before the first time period, the execution device can determine whether the weather state within at least one second time period meets the switching conditions. Only if the switching conditions are met will the final weather state within the first time period be determined as the weather state within the first time period; if the switching conditions are not met, the final weather state within the first time period will be determined as the weather state before the first time period.
[0129] To provide a more intuitive understanding of this scheme, examples are provided below using Tables 3 and 4, which both show the switching conditions when transitioning between different weather states. Table 3 uses a value of 4 for N, with the four weather states being normal, light rain, heavy rain, and others as examples.
[0130]
[0131] Table 3
[0132] As shown in Table 3, for example, when the weather condition needs to switch from normal to light rain, it requires three consecutive time periods of light rain; that is, not only the first time period, but also the two preceding second time periods. Similarly, when the weather condition needs to switch from normal to heavy rain, it requires five consecutive time periods of heavy rain; that is, not only the first time period, but also the four preceding second time periods. Furthermore, when the weather condition needs to switch from heavy rain to normal, it requires ten consecutive time periods of normal; that is, not only the first time period, but also the nine preceding second time periods. The switching conditions between other weather conditions in Table 3 can be found in the preceding descriptions and will not be elaborated upon here. It should be noted that the examples in Table 3 are only for facilitating understanding of this scheme and are not intended to limit this scheme.
[0133] Please refer to Table 4. In Table 4, N takes the value of 5, and the 5 weather conditions are normal, rainy, snowy, sleet and others.
[0134]
[0135] Table 4
[0136] Table 4 shows the switching conditions when switching between five weather states: normal, rainy, snowy, sleet, and others. For an understanding of Table 4, please refer to the description of Table 3 above. It will not be repeated here. It should be noted that the examples in Table 4 are only for the convenience of understanding this scheme and are not intended to limit this scheme.
[0137] 305. Based on the second information, determine a third region of interest in the first image obtained by the second camera, wherein the second camera belongs to at least one first camera, the first image obtained by the second camera is included in at least one first image, and the third region of interest includes the area in the first image obtained by the second camera caused by an occlusion of the second camera.
[0138] In some embodiments of this application, for any one of the at least one first camera (hereinafter referred to as the "second camera" for convenience), the execution device can determine the first image obtained by the first and second cameras from the at least one first image acquired in step 301, and determine a third region of interest in each first image obtained by the second camera based on the second information. The third region of interest includes the area in each first image obtained by the second camera caused by occlusions of the second camera. The concept of "the area in the first image caused by occlusions of the second camera" can be referred to the description in step 302 above, and will not be repeated here.
[0139] 306. Determine the degree of failure of the second camera based on the area caused by the occlusion of the second camera in the first image obtained by the second camera.
[0140] In some embodiments of this application, the execution device can determine the coverage rate of all areas caused by the occlusion of the second camera lens in each first image obtained by the second camera; and determine the degree of failure of the second camera based on the coverage rate of all areas caused by the occlusion of the second camera lens in each first image. The execution device can set H degrees of failure and H failure conditions corresponding one-to-one with the H degrees of failure, where H is an integer greater than or equal to 2. For example, the H degrees of failure may include failure and no failure; or, for example, the H degrees of failure may include mild failure, moderate failure, and severe failure, etc. These examples are only for the convenience of understanding this solution and are not intended to limit this solution.
[0141] Optionally, if the number of first images obtained by the first camera is at least two, after determining the coverage rate of all areas caused by the occlusion of the second camera in each first image, the executing device can perform a weighted summation of the coverage rates of all areas caused by the occlusion of the second camera in each first image to obtain the coverage rate of all areas caused by the occlusion of the second camera within the lens of the second camera. The coverage rate of all areas caused by the occlusion of the second camera within the lens of the second camera is then compared with H failure conditions to determine the degree of failure of the second camera.
[0142] For example, the H failure levels may include mild failure, moderate failure, and severe failure. The failure conditions for mild failure are coverage rates of 0 to 0.25, the failure conditions for moderate failure are coverage rates of 0.25 to 0.5, and the failure conditions for severe failure are coverage rates greater than or equal to 0.5.
[0143] Optionally, the execution device can also output a prompt message based on the degree of failure of the second camera. This prompt message informs the user of the current failure level of a particular second camera. For example, when the failure level of the second camera is moderate, an indicator light can be emitted and the location of the moderately failed second camera can be displayed to the user; when the failure level of the second camera is severe, an indicator light and warning sound can be emitted and the location of the severely failed second camera can be displayed to the user, and so on. It should be noted that the execution device can flexibly set how to emit the prompt message to the user after determining the failure level of the second camera, depending on the actual scenario, and there is no limitation here.
[0144] In this embodiment of the application, the area caused by the occlusion of the second camera in the first image obtained by the second camera (i.e., any first camera) can also be determined according to the second information. Then, the degree of failure of the second camera can be determined according to the area caused by the occlusion of the second camera in the first image obtained by the second camera. Since the image obtained by the camera with a high degree of failure has low reliability, timely detection of the camera with a high degree of failure is beneficial to improving the quality of the first image obtained, and thus beneficial to improving the accuracy of the determined weather conditions of the surrounding environment.
[0145] It should be noted that steps 305 and 306 are optional. If steps 305 and 306 are not executed, other steps will be executed after step 304.
[0146] In this embodiment of the application, a scheme for determining the weather state of the surrounding environment is provided. After acquiring a first image of the surrounding environment, not only is a predicted weather state for each first image generated by a first neural network, but also at least one region of interest in each first image is acquired. The first region of interest in the at least one region of interest includes a region caused by an obstruction of the first camera and which is related to the predicted weather state of the first image. By combining the predicted weather state of each first image generated by the first neural network and the at least one region of interest in the first image, the weather state of the surrounding environment is finally determined, which helps to improve the accuracy of the finally determined weather state.
[0147] II. Training Phase
[0148] In the embodiments of this application, please refer to Figure 9 , Figure 9 This is a flowchart illustrating a neural network training method provided in an embodiment of this application. The neural network training method provided in an embodiment of this application may include:
[0149] 901. Acquire training images. Training images are obtained by taking pictures of the surrounding environment with a camera.
[0150] 902. Generate second information through a first neural network, wherein the second information indicates the predicted location of at least one region of interest in a training image, the at least one region of interest including a region caused by an obstruction of the camera lens, or the at least one region of interest also including a region caused by a vehicle.
[0151] 903. The weather state of the second image is predicted by the first neural network to obtain first information, wherein the second image includes image regions in the training images other than at least one region of interest, and the first information indicates the predicted weather state of each training image.
[0152] In this embodiment, the training device can acquire training images, input the training images into a first neural network, and generate second information through the first neural network; optionally, the weather state of the second image can also be predicted through the first neural network to obtain the first information. The specific implementation of steps 902 and 903 by the training device can be found in the above description. Figure 3 The descriptions in the corresponding embodiments are not repeated here.
[0153] 904. The first neural network is trained according to a first loss function term, wherein the first loss function term indicates the similarity between second information and first expected information, and the first expected information indicates the correct location of at least one region of interest in the training image.
[0154] In this embodiment, step 903 is optional. If step 903 is executed, step 904 may include: the training device iteratively training the first neural network based on a first loss function term and a second loss function term. The first loss function term indicates the similarity between second information and first expected information. The first expected information indicates the correct location of at least one region of interest in the training image. Therefore, the goal of training the first neural network using the first loss function term includes improving the similarity between the second information and the first expected information output by the first neural network. The second loss function term indicates the similarity between the first information and the second expected information. The second expected information indicates the correct weather state of the training image. Therefore, the goal of training the first neural network using the second loss function term includes improving the similarity between the first information and the second expected information output by the first neural network.
[0155] Specifically, after executing steps 902 and 903, the training device can generate the values of the first loss function term and the second loss function term. Based on the values of the first and second loss function terms, the value of the total loss function is determined, and the gradient derivative of the total loss function value is calculated to update the weight parameters of the first neural network in reverse order, thereby completing one training of the first neural network. The training device repeats the aforementioned steps to achieve iterative training of the first neural network.
[0156] If step 903 is not executed, then step 904 may include: the training device iteratively training the first neural network based on the first loss function term. Specifically, after executing step 902, the training device can generate the value of the first loss function term and perform gradient differentiation on the value of the first loss function term to update the weight parameters of the first neural network in reverse order, thereby completing one training of the first neural network; the training device repeats the aforementioned steps to achieve iterative training of the first neural network.
[0157] In this embodiment, since the camera's obstruction may include obstructions on the camera lens and obstructions suspended outside the camera, for example, the camera lens may be covered with raindrops, mud, ice, or other coverings, or, for example, if the surrounding environment of the camera is rainy, the obstructions suspended outside the camera may also have raindrops; and the area in the image caused by the camera's obstruction will affect the observation of the actual surrounding environment, the first neural network trained using this scheme can identify the area in the image caused by the camera's obstruction, which is beneficial for the execution device of the trained first neural network to understand the surrounding environment more accurately.
[0158] exist Figures 1 to 9 Based on the corresponding embodiments, in order to better implement the above-described solutions of this application, related equipment for implementing the above solutions is also provided below. See details. Figure 10 , Figure 10 This is a schematic diagram of an image processing apparatus provided in an embodiment of this application. The image processing apparatus 1000 includes: an acquisition module 1001, configured to acquire at least one first image, wherein the at least one first image is obtained by capturing the surrounding environment through at least one first camera; a generation module 1002, configured to generate first information through a first neural network, wherein the first information indicates the predicted weather state of each first image; and a determination module 1003, configured to determine the weather state of the surrounding environment based on the first information and second information, wherein the second information indicates at least one region of interest in each first image, and the first region of interest in the at least one region of interest includes a region that is related to the predicted weather state of the first image.
[0159] Optionally, the number of first images is at least two, and the at least two first images include images from different first cameras.
[0160] Optionally, module 1003 is specifically used for:
[0161] Based on the second information, parameter information corresponding to each first region of interest is determined. The parameter information includes any one or more of the following: the coverage of the first region of interest in the first image, the area of the first region of interest, the brightness of the first region of interest, or the blur of the first region of interest.
[0162] Based on the first information and the parameter information corresponding to the first region of interest, the weather conditions of the surrounding environment are determined.
[0163] Optionally, the generation module 1002 is specifically used to predict the weather state of the second image through the first neural network to obtain first information, wherein the second image includes image regions in the first image other than a second region of interest in at least one region of interest, and the second region of interest includes a region caused by an obstruction of the first camera and which is related to the weather state.
[0164] Optionally, the second region of interest in at least one region of interest includes any one or more of the following: a blurred region, a water droplet region, a muddy region, an icy region, a snowflake region, or a glare region.
[0165] Optionally, the determining module 1003 is further configured to determine a third region of interest in the first image obtained by the second camera based on the second information, wherein the second camera belongs to at least one first camera, the first image obtained by the second camera is included in at least one first image, and the third region of interest includes the area in the first image obtained by the second camera caused by an occlusion of the second camera; the determining module is further configured to determine the degree of failure of the second camera based on the area in the first image obtained by the second camera caused by an occlusion of the second camera.
[0166] It should be noted that the information interaction and execution process between the modules / units in the image processing device 1000 are based on the same concept as the various method embodiments in this application. For details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0167] Please see Figure 11 , Figure 11 This is another schematic diagram of the image processing apparatus provided in an embodiment of this application. The image processing apparatus 1100 includes: an acquisition module 1101, configured to acquire at least one first image, each first image being captured by a first camera of the surrounding environment; the acquisition module 1101, configured to acquire second information, wherein the second information indicates at least one region of interest in each first image, the at least one region of interest including a second region of interest, the second region of interest including a region caused by an obstruction of the first camera and a region related to the weather state; and a processing module 1102, configured to predict the weather state of the second image through a first neural network to obtain first information, wherein the second image includes image regions in the first image other than the second region of interest, the first information indicating the predicted weather state of each first image, and the first information being used to determine the weather state of the surrounding environment.
[0168] Optionally, the processing module 1102 is further configured to determine the weather state of the surrounding environment based on the first information and a first region of interest in at least one region of interest, wherein the first region of interest in at least one region of interest includes a region that is associated with the predicted weather state of the first image.
[0169] Optionally, the processing module 1002 is specifically used for:
[0170] Based on the second information, determine the parameter information corresponding to the first region of interest. The parameter information includes any one or more of the following: the coverage of the first region of interest in the first image, the area of the first region of interest, the brightness of the first region of interest, or the blur of the first region of interest.
[0171] Based on the first information and the parameter information corresponding to the first region of interest, the weather conditions of the surrounding environment are determined.
[0172] It should be noted that the information interaction and execution process between the modules / units in the image processing device 1100 are based on the same concept as the various method embodiments in this application. For details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0173] Please see Figure 12 , Figure 12 This is a schematic diagram of a neural network training device provided in an embodiment of this application. The neural network training device 1200 includes: an acquisition module 1201 for acquiring training images, which are obtained by capturing images of the surrounding environment using a camera; a processing module 1202 for generating second information through a first neural network, wherein the second information indicates the predicted position of at least one region of interest in the training image, and the at least one region of interest includes a region caused by an obstruction of the camera; and a training module 1203 for training the first neural network according to a first loss function term, wherein the first loss function term indicates the similarity between the second information and first expected information, and the first expected information indicates the correct position of at least one region of interest in the training image.
[0174] Optionally, the processing module 1202 is further configured to predict the weather state of the second image through the first neural network to obtain first information, wherein the second image includes image regions in the training image other than at least one region of interest, and the first information indicates the predicted weather state of each training image; the training module 1203 is specifically configured to train the first neural network according to a first loss function term and a second loss function term, wherein the second loss function term indicates the similarity between the first information and the second expected information, and the second expected information indicates the correct weather state of the training image.
[0175] It should be noted that the information interaction and execution process between the modules / units in the neural network training device 1200 are based on the same concept as the various method embodiments in this application. For details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0176] The following describes an execution device provided in an embodiment of this application. Please refer to [link / reference]. Figure 13 , Figure 13 This is a schematic diagram of an execution device provided in an embodiment of this application. The execution device 1300 can specifically be a vehicle, mobile phone, laptop computer, or monitoring data processing device, etc., and is not limited thereto. Specifically, the execution device 1300 includes: a receiver 1301, a transmitter 1302, a processor 1303, and a memory 1304 (wherein the execution device 1300 may have one or more processors 1303). Figure 13 (Taking a processor as an example), processor 1303 may include application processor 13031 and communication processor 13032. In some embodiments of this application, receiver 1301, transmitter 1302, processor 1303 and memory 1304 may be connected via bus or other means.
[0177] Memory 1304 may include read-only memory and random access memory, and provides instructions and data to processor 1303. A portion of memory 1304 may also include non-volatile random access memory (NVRAM). Memory 1304 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.
[0178] Processor 1303 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together through a bus system, which may include not only the data bus, but also power buses, control buses, and status signal buses. However, for clarity, all buses are referred to as the bus system in the diagram.
[0179] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 1303. The processor 1303 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 1303 or by instructions in software form. The processor 1303 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1303 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 1304. Processor 1303 reads the information in memory 1304 and, in conjunction with its hardware, completes the steps of the above method.
[0180] Receiver 1301 can be used to receive input digital or character information, and to generate signal inputs related to the settings and function control of the execution device. Transmitter 1302 can be used to output digital or character information through the first interface; transmitter 1302 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; transmitter 1302 may also include a display device such as a display screen.
[0181] In this embodiment of the application, the processor 1303 is used to execute... Figures 2b to 8 The image processing method executed by the execution device in the corresponding embodiment. It should be noted that the specific manner in which the application processor 13031 in processor 1303 executes the above steps is different from that in this application. Figures 2b to 8 The various method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 2b to 8 The corresponding method embodiments are the same, and for details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0182] This application also provides a training device; please refer to [link / reference]. Figure 14 , Figure 14 This is a schematic diagram of a training device provided in an embodiment of this application. The training device 1400 is implemented by one or more servers. The training device 1400 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1422 (e.g., one or more processors) and memory 1432, and one or more storage media 1430 (e.g., one or more mass storage devices) for storing application programs 1442 or data 1444. The memory 1432 and storage media 1430 can be temporary or persistent storage. The program stored in the storage media 1430 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the training device. Furthermore, the CPU 1422 may be configured to communicate with the storage media 1430 and execute the series of instruction operations in the storage media 1430 on the training device 1400.
[0183] The training device 1400 may also include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input / output interfaces 1458, and / or one or more operating systems 1441, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0184] In this embodiment of the application, the central processing unit 1422 is used to execute... Figure 9 The training method for the neural network executed by the training device in the corresponding embodiment. It should be noted that the specific manner in which the central processing unit 1422 executes each step differs from that in this application. Figure 9 The various method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figure 9 The corresponding method embodiments are the same, and for details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0185] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned actions. Figures 2b to 8 The method described in the illustrated embodiment executes steps performed by the execution device, or causes the computer to perform the steps as described above. Figure 9 The steps performed by the training device in the method described in the illustrated embodiment.
[0186] This application embodiment also provides a computer-readable storage medium storing a program for performing signal processing, which, when run on a computer, causes the computer to perform the aforementioned actions. Figures 2b to 8 The method described in the illustrated embodiment executes steps performed by the execution device, or causes the computer to perform the steps as described above. Figure 9 The steps performed by the training device in the method described in the illustrated embodiment.
[0187] The execution device or training device provided in this application embodiment can specifically be a chip. The chip includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in the storage unit to cause the chip to perform the above-mentioned operations. Figures 2b to 8 The image processing method described in the illustrated embodiment, or, to cause a chip within a training device to perform the above-described... Figure 9 The image processing method described in the illustrated embodiment. Optionally, the storage unit is a storage unit within the chip, such as a register or cache. Alternatively, the storage unit can be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).
[0188] For details, please refer to Figure 15 , Figure 15 This is a schematic diagram of a chip provided in an embodiment of this application. The chip can be represented as a neural network processor (NPU) 150. The NPU 150 is mounted as a coprocessor on the host CPU, and tasks are assigned by the host CPU. The core part of the NPU is the arithmetic circuit 150, which is controlled by the controller 1504 to extract matrix data from the memory and perform multiplication operations.
[0189] In some implementations, the arithmetic circuit 1503 internally includes multiple processing engines (PEs). In some implementations, the arithmetic circuit 1503 is a two-dimensional pulsating array. The arithmetic circuit 1503 can also be a one-dimensional pulsating array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1503 is a general-purpose matrix processor.
[0190] For example, suppose we have an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from the weight memory 1502 and caches it in each PE of the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from the input memory 1501 and performs matrix operations with matrix B. The partial result or the final result of the obtained matrix is stored in the accumulator 1508.
[0191] Unified memory 1506 is used to store input and output data. Weight data is directly transferred to weight memory 1502 via Direct Memory Access Controller (DMAC) 1505. Input data is also transferred to unified memory 1506 via DMAC.
[0192] BIU stands for Bus Interface Unit, which is used for interaction between the AXI bus and the DMAC and the Instruction Fetch Buffer (IFB) 1509.
[0193] The Bus Interface Unit (BIU) 1510 is used by the instruction fetch memory 1509 to fetch instructions from external memory, and also by the memory access controller 1505 to fetch the original data of the input matrix A or the weight matrix B from external memory.
[0194] The DMAC is mainly used to move input data from external memory DDR to unified memory 1506, or to weight data to weight memory 1502, or to input data to input memory 1501.
[0195] The vector computation unit 1507 includes multiple arithmetic processing units that further process the output of the computation circuit as needed, such as vector multiplication, vector addition, exponential operations, logarithmic operations, size comparisons, etc. It is mainly used for computation in non-convolutional / fully connected layers of neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.
[0196] In some implementations, the vector computation unit 1507 can store the processed output vector in the unified memory 1506. For example, the vector computation unit 1507 can apply linear and / or nonlinear functions to the output of the computation circuit 1503, such as performing linear interpolation on feature planes extracted by convolutional layers, or accumulating a vector of values to generate activation values. In some implementations, the vector computation unit 1507 generates normalized values, pixel-level summed values, or both. In some implementations, the processed output vector can be used as activation input to the computation circuit 1503, for example, for use in subsequent layers of the neural network.
[0197] The instruction fetch buffer 1509 connected to the controller 1504 is used to store the instructions used by the controller 1504;
[0198] Unified memory 1506, input memory 1501, weighted memory 1502, and instruction fetch memory 1509 are all on-chip memories. External memory is proprietary to this NPU hardware architecture.
[0199] in, Figures 2b to 9 The operations of each layer in the neural network shown can be performed by the operation circuit 1503 or the vector calculation unit 1507.
[0200] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of a program in the first aspect of the method.
[0201] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0202] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0203] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0204] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. An image processing method, characterized in that, The method includes: At least one first image is acquired, the at least one first image being captured by at least one first camera of the surrounding environment; the number of the first images is at least two, and the at least two first images include images from different first cameras; First information is generated through a first neural network, the first information indicating the predicted weather state for each of the first images; Determining the weather state of the surrounding environment based on the first information and the second information includes: determining parameter information corresponding to each first region of interest based on the second information, the parameter information including any one or more of the following: the coverage of the first region of interest in the first image, the area of the first region of interest, the brightness of the first region of interest, or the blur of the first region of interest; determining the weather state of the surrounding environment based on the first information and the parameter information corresponding to the first region of interest; the second information indicates at least one region of interest in each first image, the first region of interest in the at least one region of interest includes a region associated with the predicted weather state of the first image; the region associated with the predicted weather state of the first image is caused by an obstruction on the lens of the first camera.
2. The method according to claim 1, characterized in that, The generation of first information through the first neural network includes: The weather conditions of the second image are predicted by the first neural network to obtain the first information. The second image includes an image region in the first image other than the second region of interest in the at least one region of interest. The second region of interest includes a region caused by the obstruction of the first camera and which is related to the weather conditions.
3. The method according to claim 2, characterized in that, The second region of interest in at least one region of interest includes any one or more of the following: a blurred region, a water droplet region, a muddy water region, an icy region, a snowflake region, or a glare region.
4. The method according to claim 1, characterized in that, The method further includes: Based on the second information, a third region of interest is determined in the first image obtained by the second camera, wherein the second camera belongs to the at least one first camera, the first image obtained by the second camera is included in the at least one first image, and the third region of interest includes the area in the first image obtained by the second camera caused by the occlusion of the second camera; The degree of failure of the second camera is determined based on the area in the first image obtained by the second camera caused by the obstruction of the second camera.
5. An image processing method, characterized in that, The method includes: At least one first image is acquired, each first image being captured by a first camera of the surrounding environment; the number of first images is at least two, and the at least two first images contain images from different first cameras; Acquire second information, wherein the second information indicates at least one region of interest in each of the first images, the at least one region of interest including a second region of interest, the second region of interest including an area caused by an obstruction on the lens of the first camera and an area that is related to the weather conditions; The weather conditions of the second image are predicted by the first neural network to obtain first information. The second image includes the image region in the first image other than the second region of interest. The first information indicates the predicted weather conditions of each first image and is used to determine the weather conditions of the surrounding environment. Determining the weather state of the surrounding environment based on the first information and the first region of interest (ROI) of the at least one region of interest includes: determining parameter information corresponding to the first ROI based on the second information, wherein the parameter information includes any one or more of the following: the coverage of the first ROI in the first image, the area of the first ROI, the brightness of the first ROI, or the blur of the first ROI; determining the weather state of the surrounding environment based on the first information and the parameter information corresponding to the first ROI, wherein the first ROI of the at least one region of interest includes a region that is associated with the predicted weather state of the first image.
6. A method for training a neural network, characterized in that, The method includes: Acquire training images, which are obtained by capturing the surrounding environment with a camera; generate second information through a first neural network, wherein the second information indicates the predicted location of at least one region of interest in the training images, the at least one region of interest including an area caused by an obstruction on the lens of the camera and an area that is related to the weather conditions; The first neural network is trained according to a first loss function term, wherein the first loss function term indicates the similarity between the second information and the first expected information, and the first expected information indicates the correct location of the at least one region of interest in the training image; The weather state of the second image is predicted using the first neural network to obtain first information, wherein the second image includes image regions in the training image other than the at least one region of interest, and the first information indicates the predicted weather state of each training image; the first information is used to determine the weather state of the surrounding environment with parameter information corresponding to the first region of interest in the at least one region of interest, and the parameter information includes any one or more of the following: the coverage of the first region of interest in the training image, the area of the first region of interest, the brightness of the first region of interest, or the blur of the first region of interest; the parameter information is determined according to the second information; the first region of interest in the at least one region of interest includes a region that is associated with the predicted weather state of the training image; training the first neural network according to the first loss function term includes: The first neural network is trained based on the first loss function term and the second loss function term, wherein the second loss function term indicates the similarity between the first information and the second expected information, and the second expected information indicates the correct weather state of the training image.
7. An image processing apparatus, characterized in that, The device includes: An acquisition module is configured to acquire at least one first image, wherein the at least one first image is obtained by capturing the surrounding environment through at least one first camera; the number of the first images is at least two, and the at least two first images contain images from different first cameras; A generation module is configured to generate first information through a first neural network, wherein the first information indicates the predicted weather state for each of the first images; A determining module is configured to determine the weather state of the surrounding environment based on the first information and the second information, wherein the second information indicates at least one region of interest in each of the first images, and the first region of interest in the at least one region of interest includes a region that is associated with the predicted weather state of the first image; the region that is associated with the predicted weather state of the first image is caused by an obstruction on the lens of the first camera. The determining module is specifically configured to: determine parameter information corresponding to each of the first regions of interest based on the second information, wherein the parameter information includes any one or more of the following: the coverage of the first region of interest in the first image, the area of the first region of interest, the brightness of the first region of interest, or the blur of the first region of interest; and determine the weather state of the surrounding environment based on the first information and the parameter information corresponding to the first region of interest.
8. The apparatus according to claim 7, characterized in that, The generation module is specifically used to predict the weather state of the second image through the first neural network to obtain the first information. The second image includes an image region in the first image other than the second region of interest in the at least one region of interest. The second region of interest includes a region caused by the obstruction of the first camera and which is related to the weather state.
9. The apparatus according to claim 8, characterized in that, The second region of interest in at least one region of interest includes any one or more of the following: a blurred region, a water droplet region, a muddy water region, an icy region, a snowflake region, or a glare region.
10. The apparatus according to claim 9, characterized in that, The determining module is further configured to determine a third region of interest in the first image obtained by the second camera based on the second information, wherein the second camera belongs to the at least one first camera, the first image obtained by the second camera is included in the at least one first image, and the third region of interest includes the area in the first image obtained by the second camera caused by the occlusion of the second camera; The determining module is further configured to determine the degree of failure of the second camera based on the area in the first image obtained by the second camera caused by the occlusion of the second camera.
11. An image processing apparatus, characterized in that, The device includes: An acquisition module is used to acquire at least one first image, each first image being captured by a first camera of the surrounding environment; the number of first images is at least two, and the at least two first images contain images from different first cameras; The acquisition module is used to acquire second information, wherein the second information indicates at least one region of interest in each of the first images, the at least one region of interest including a second region of interest, the second region of interest including an area caused by an obstruction on the lens of the first camera and an area that is related to the weather conditions; The processing module is used to predict the weather state of the second image through a first neural network to obtain first information. The second image includes the image region in the first image other than the second region of interest. The first information indicates the predicted weather state of each first image and is used to determine the weather state of the surrounding environment. The processing module is further configured to determine the weather state of the surrounding environment based on the first information and the first region of interest in the at least one region of interest, wherein the first region of interest in the at least one region of interest includes a region that is associated with the predicted weather state of the first image; The processing module is specifically used to: determine parameter information corresponding to the first region of interest based on the second information, wherein the parameter information includes any one or more of the following: the coverage of the first region of interest in the first image, the area of the first region of interest, the brightness of the first region of interest, or the blur of the first region of interest. Based on the first information and the parameter information corresponding to the first region of interest, the weather state of the surrounding environment is determined.
12. A training device for a neural network, characterized in that, The device includes: An acquisition module is used to acquire training images, which are obtained by taking pictures of the surrounding environment with a camera. A processing module is configured to generate second information through a first neural network, wherein the second information indicates the predicted location of at least one region of interest in the training image, the at least one region of interest including an area caused by an obstruction of the camera lens and an area related to weather conditions; A training module is configured to train the first neural network based on a first loss function term, wherein the first loss function term indicates the similarity between the second information and the first expected information, and the first expected information indicates the correct location of the at least one region of interest in the training image. The processing module is further configured to predict the weather state of the second image using the first neural network to obtain first information, wherein the second image includes image regions in the training image other than the at least one region of interest, and the first information indicates the predicted weather state of each training image; the first information is used to determine the weather state of the surrounding environment together with parameter information corresponding to a first region of interest in the at least one region of interest, and the parameter information includes any one or more of the following: the coverage of the first region of interest in the training image, the area of the first region of interest, the brightness of the first region of interest, or the blur of the first region of interest; the parameter information is determined according to the second information; the first region of interest in the at least one region of interest includes a region that is associated with the predicted weather state of the training image; The training module is specifically used to train the first neural network based on the first loss function term and the second loss function term, wherein the second loss function term indicates the similarity between the first information and the second expected information, and the second expected information indicates the correct weather state of the training image.
13. A computer program product, characterized in that, The computer program product includes a program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 6.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 6.
15. An execution device, characterized in that, It includes a processor and a memory, the processor being coupled to the memory, the memory being used to store programs; The processor is configured to execute a program in the memory, causing the execution device to perform the method as described in any one of claims 1 to 4, or to cause the execution device to perform the method as described in claim 5.
16. A training device, characterized in that, It includes a processor and a memory, the processor being coupled to the memory, the memory being used to store programs; The processor is configured to execute a program in the memory, causing the training device to perform the method as described in claim 6.