Machine Vision-Based Image Processing Method and Device

Through the combination of the preloaded environment detection model and the lightweight scene detection model, the problem of large amount of computing in different scenarios is solved, and efficient image processing and detection performance improvement on devices with limited computing power are achieved.

CN112204566BActive Publication Date: 2025-08-05SZ ZHUOYU TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980033604.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-08-15
Publication Date
2025-08-05
Estimated Expiration
2039-08-15

AI Technical Summary

Technical Problem

The existing object detection algorithm has a large amount of computation in different scenarios, making it difficult to maintain efficient recognition performance on devices with limited computing power.

Method used

The preloaded environment detection model is used to determine the current scene, and the lightweight scene detection model matching the current scene is loaded, and the lightweight scene detection model is selected to process the environment image.

Benefits of technology

When computing power is limited, the efficiency of image processing and detection performance in different scenarios are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112204566B_ABST
    Figure CN112204566B_ABST
Patent Text Reader

Abstract

A machine vision-based image processing method and device, applied to a mobile platform equipped with an image acquisition device, comprises: acquiring an environmental image (101); using a preloaded environmental detection model to determine a current scene based on the environmental image (102); loading a scene detection model that matches the current scene (103); and processing the environmental image based on the scene detection model (104). When computing power is constrained, a lightweight scene detection model corresponding to the current scene is selected, thereby improving processing efficiency and performance in different scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of intelligent control and perception technology, and in particular to an image processing method and device based on machine vision. Background Art

[0002] Target detection algorithm is one of the key technologies for autonomous driving and intelligent drones. It can detect and identify the location, category and confidence of objects of interest in visual images, providing the necessary observation information for subsequent intelligent functions.

[0003] In related technologies, object detection algorithms typically use a single universal model for all scenarios, such as a trained neural network model or a perceptual algorithm based on feature point recognition. To ensure highly reliable recognition results across different scenarios, neural network models must be trained on a large amount of data from diverse scenarios. While achieving high-performance detection results across diverse scenarios often requires complex model design, significantly increasing computational effort. Summary of the Invention

[0004] The present disclosure provides an image processing method and device based on machine vision, which improves image processing efficiency.

[0005] In a first aspect, the present disclosure provides an image processing method based on machine vision, which is applied to a mobile platform equipped with an image acquisition device, the method comprising:

[0006] Get the environment image;

[0007] Using a preloaded environment detection model, determining a current scene based on the environment image;

[0008] Loading a scene detection model that matches the current scene;

[0009] An environment image is processed based on the scene detection model.

[0010] In a second aspect, the present disclosure provides a vehicle equipped with a camera device, a memory, and a processor, wherein the memory is used to store instructions, and the instructions are executed by the processor to implement any one of the methods in the first aspect.

[0011] In a third aspect, the present disclosure provides a drone equipped with a camera, a memory, and a processor, wherein the memory is used to store instructions, and the instructions are executed by the processor to implement any one of the methods in the first aspect.

[0012] In a fourth aspect, the present disclosure provides an electronic device that can be communicatively connected to a camera device, the electronic device comprising a memory and a processor, the memory being used to store instructions, the instructions being executed by the processor to implement any one of the methods described in the first aspect.

[0013] In a fifth aspect, the present disclosure provides a handheld gimbal, comprising: a camera device, a memory, and a processor, wherein the memory is used to store instructions, and the instructions are executed by the processor to implement any one of the methods in the first aspect.

[0014] In a sixth aspect, the present disclosure provides a mobile terminal comprising: a camera device, a memory, and a processor, wherein the memory is used to store instructions, and the instructions are executed by the processor to implement any one of the methods in the first aspect.

[0015] The present disclosure provides a machine vision-based image processing method and device, which acquire an environmental image; use a preloaded environmental detection model to determine the current scene based on the environmental image; load a scene detection model that matches the current scene; process the environmental image based on the scene detection model, and when computing power is constrained, select a lightweight scene detection model corresponding to the current scene, thereby improving the efficiency of image processing and the respective performance in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0017] Figure 1 A schematic diagram of a drone provided in one embodiment of the present disclosure;

[0018] Figure 2 A schematic diagram of a handheld gimbal provided in accordance with an embodiment of the present disclosure;

[0019] Figure 3 An application schematic diagram provided for an embodiment of the present disclosure;

[0020] Figure 4 This is a flow chart of an embodiment of a machine vision-based image processing method provided by the present disclosure;

[0021] Figure 5 A schematic diagram of a scenario provided in an embodiment of the present disclosure;

[0022] Figure 6 A schematic diagram of a scenario provided for another embodiment of the present disclosure;

[0023] Figure 7 A schematic diagram showing a comparison of network models according to an embodiment of the present disclosure;

[0024] Figure 8 is a flowchart of another embodiment of the image processing method provided by the present disclosure;

[0025] Figure 9 This is a flowchart of another embodiment of the image processing method disclosed herein;

[0026] Figure 10 A schematic structural diagram of a vehicle provided in one embodiment of the present disclosure;

[0027] Figure 11 A schematic diagram of the structure of a drone provided by an embodiment of the present disclosure;

[0028] Figure 12 A schematic structural diagram of an electronic device provided in one embodiment of the present disclosure;

[0029] Figure 13 A schematic diagram of the structure of a handheld gimbal provided in one embodiment of the present disclosure;

[0030] Figure 14 A schematic structural diagram of a mobile terminal provided in one embodiment of the present disclosure;

[0031] Figure 15 This is a schematic diagram of the memory loading status disclosed in the embodiments of this specification. DETAILED DESCRIPTION

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0033] First, the application scenarios involved in this disclosure are introduced:

[0034] The machine vision-based image processing method provided by the embodiments of the present disclosure is applied to scenarios such as autonomous driving and intelligent drones. It can detect and identify the position, category, and confidence level of objects of interest in the image, providing the necessary observation information for subsequent other functions.

[0035] In an optional embodiment, the method may be performed by the drone 10, such as Figure 1As shown, the drone 10 can be equipped with a camera device 1, and for example, the drone's processor can execute the corresponding software code to implement it, or the drone can execute the corresponding software code while interacting with the server for data, such as the server executing some operations to control the drone to execute the image processing method.

[0036] In an optional embodiment, the method can be performed by a handheld gimbal, such as Figure 2 As shown, the handheld gimbal 20 may include a camera device 2, which may be implemented, for example, by the processor of the handheld gimbal executing the corresponding software code, or by the drone executing the corresponding software code while interacting with the server for data, such as the server executing some operations to control the drone to execute the image processing method.

[0037] The camera device is used to obtain environmental images, such as environmental images around the drone or handheld gimbal.

[0038] In an optional embodiment, the method may be executed by an electronic device such as a mobile terminal, such as Figure 3 As shown, the electronic device can be installed on a vehicle or drone; or it can be executed by an onboard control device that communicates with the electronic device. The vehicle can be an autonomous vehicle or a conventional vehicle. For example, the electronic device, such as a processor of the electronic device, can execute corresponding software code to implement the image processing method. Alternatively, the electronic device can execute the corresponding software code while interacting with a server, such as the server executing some operations to control the electronic device to execute the image processing method.

[0039] In the consumer electronics market, electronic devices face computing power and bandwidth bottlenecks due to the different processor models they are equipped with.

[0040] The following specific embodiments are used to describe the technical solution of the present disclosure in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0041] Figure 4 This is a flow chart of an embodiment of the image processing method based on machine vision provided by the present disclosure. Figure 4 As shown, the method provided in this embodiment is applied to a movable platform equipped with an image acquisition device, and the method includes:

[0042] Step 101: Acquire an environment image.

[0043] In an optional embodiment, the environmental image may be image information captured by an image acquisition device. The image acquisition device is typically mounted on a movable object, which may be a vehicle, an unmanned aerial vehicle (UAV), a ground mobile robot, or the like. The image acquisition device may be a monocular camera, a binocular camera, a multi-camera camera, a fisheye lens, a compound eye lens, or the like. The camera device captures environmental image information surrounding the movable object, such as image information from the front, rear, or side of the movable object. In an optional embodiment, the camera device may also capture wide-band or panoramic information surrounding the movable object; it may capture multiple images, portions of an image, portions of an image, or a combination of images. The captured environmental image may be the original image output by the image sensor, or an image that has been processed but retains the original image's brightness information, such as an image in RGB or HSV format. The aforementioned environmental image may be environmental image information captured by the image acquisition device while a vehicle is driving or a UAV is flying.

[0044] Movable platforms include drones, vehicles, electronic equipment and other platforms.

[0045] Step 102: Use the preloaded environment detection model to determine the current scene based on the environment image.

[0046] In an optional embodiment, determining the current scene information includes extracting possible scenes in which the movable object is located based on the environment image acquired in the aforementioned step 101 .

[0047] This step can be implemented according to a judgment function, for example, reading the RGB or HSV distribution information of the environment image obtained in step 101 and judging the current scene according to the distribution.

[0048] This step may also be a statistical comparison process, for example, reading histogram information in HSV and then judging the scene based on the histogram information.

[0049] This step can also be performed through an environment detection model, which can be implemented based on a neural network. A neural network is constructed to output the current scene based on the input environment image.

[0050] In an optional embodiment, the scenes may include different time scenes, such as day and night; different weather scenes, such as sunny, rainy, foggy, snowy, etc.; different road conditions scenes, such as highways, urban roads, rural roads, etc.

[0051] In an optional embodiment, the current scene may include at least two scenes divided according to image brightness.

[0052] In an optional embodiment, the current scene divided according to image brightness may include a high-brightness scene and a low-brightness scene.

[0053] In an optional embodiment, the current scene divided according to image brightness may include a high-brightness scene, a medium-brightness scene, and a low-brightness scene.

[0054] In an optional embodiment, the current scene may include at least two scenes divided according to image visibility.

[0055] In an optional embodiment, the current scene divided according to image visibility may include a high-visibility scene and a low-visibility scene.

[0056] In an optional embodiment, the current scene divided according to image visibility may include a high-visibility scene, a medium-visibility scene, and a low-visibility scene.

[0057] In an optional embodiment, the at least two scenes divided according to image visibility may include a haze scene, a dusty scene, a snowy scene, a rainy scene, and the like.

[0058] In an optional embodiment, the current scene may include at least two scenes divided according to image texture information.

[0059] In an optional embodiment, the scenes divided according to the image texture information include weather information. In an optional embodiment, the weather information includes rain, snow, fog, blowing sand and other weather information.

[0060] Taking neural networks as an example, a network used for scene recognition only needs to output a small number of classification results. To achieve accurate output, the network layer does not require too many parameters. In other words, the neural network used for this judgment step only consumes a small amount of system computing power, and model loading also consumes only a small amount of system bandwidth.

[0061] In an optional embodiment, the environment detection model may be preloaded before determining the current scene, so that no loading operation is required when the model is used, thereby improving processing efficiency.

[0062] In an optional embodiment, the preloaded environment detection model is always in a loading state during the process of acquiring the environment image.

[0063] In order to ensure processing efficiency, the preloaded environment detection model is always in a loaded state during the environment image acquisition process, and can be used to determine the current scene at any time.

[0064] Step 103: Load a scene detection model that matches the current scene.

[0065] In an optional embodiment, this step loads a scene detection model that matches the current scene based on the current scene determined in step 102 .

[0066] The scene detection model can be established based on neural network models such as CNN, VGG, GoogleNet, etc., and trained based on training data of different scenes to obtain scene detection models matching different scenes.

[0067] The scenarios may include different time scenarios, such as day and night; different weather scenarios, such as sunny, rainy, foggy, snowy, etc.; different road conditions, such as highways, urban roads, rural roads, etc.

[0068] For example Figure 5 、 Figure 6 The scenes in which the vehicle is located are sunny scene and cloudy scene, or high brightness scene and low brightness scene.

[0069] The scene detection model corresponding to each scene may not require too many parameters and only consumes a small amount of system computing power. Small scene detection models corresponding to multiple scenes replace a large general detection model, enabling the device to work normally under limited computing power.

[0070] For example, if the computing power of the device is 500M, if the image processing function needs to load a 2.7G network model (for example Figure 7 Part a on the left), which is obviously impossible. In the solution of the embodiment of the present disclosure, the large network model is split into several small network models (i.e., scene detection models, such as Figure 7 Part b on the right) enables the device to work normally even when the computing power of the device is limited.

[0071] In an optional embodiment, the scene detection model may also be established based on other network models, which is not limited in this disclosure.

[0072] In an optional embodiment, the scene detection model matching the current scene is switched and loaded as the current scene changes.

[0073] In an optional embodiment, the scene detection model matching the current scene is not removed from the memory due to switching loading.

[0074] Specifically, a scene detection model matching the current scene is loaded based on the current scene. If the current scene changes, a scene detection model matching the changed scene is switched and loaded.

[0075] Furthermore, during the switching loading process, the scene detection model may not exit the memory, so as to improve the loading speed for the next use.

[0076] In an optional embodiment, the preloaded environment detection model and the scene detection model are in different threads.

[0077] Specifically, the preloaded environment detection model and scene detection model can be in different threads. For example, while processing the environment image using the scene detection model that matched the previously determined scene, the environment detection model can also be used to determine the current scene. The scene at this time may have changed and no longer match the scene detection model. After processing the environment image using the scene detection model, the scene detection model that matches the changed scene can be switched to process the environment image.

[0078] In an optional embodiment, the preloaded environment detection model performs inter-thread communication via a callback function.

[0079] For example, the information of the current scene determined by the environment detection model may be notified to the scene detection model through a callback function, or the environment image obtained by the image acquisition device may be obtained based on the callback function.

[0080] Step 104: Process the environment image based on the scene detection model.

[0081] In an optional embodiment, the environmental image is processed based on the scene detection model corresponding to the identified current scene, for example, the position of the target object in the environmental image, the category to which the target object belongs, and the confidence level in the category are identified.

[0082] In an optional embodiment, processing the environment image based on the scene detection model includes: acquiring object information in the environment image.

[0083] In an optional embodiment, the object information includes: position information of the target object in the environment image, category information of the target object, and confidence level of the target object in the corresponding category.

[0084] In an optional embodiment, a non-maximum suppression method is used to filter object information to obtain target detection results.

[0085] Specifically, the object information output by the scene detection model includes a large amount of target object information, including a lot of duplicate information. For example, there is a lot of location information, some of which overlaps. Methods such as non-maximum suppression can be used to filter the object information to obtain the final target detection results.

[0086] Ultimately, the location, category, and confidence level of the object of interest in the image can be obtained. This output can be provided as external observation information to downstream modules, such as state estimation and navigation control, to achieve more complex autonomous driving functions.

[0087] In an optional embodiment, the environmental image information is input into a loaded scene detection model corresponding to the current scene. The scene detection model then passes through several network layers to output target detection results, including, for example, the target object's location, category, and confidence level within that category. Target objects can be, for example, dynamic and / or static targets. Dynamic targets can include, for example, moving vehicles and drones, while static targets can include the number of surrounding objects, road signs, utility poles, and the like.

[0088] For example, Figure 5 As shown, the image acquisition device loaded on the vehicle acquires the environmental image around the vehicle. The vehicle uses a preloaded environmental detection model to determine the current scene based on the environmental image. For example, it determines that the current scene is a high-brightness scene, loads the scene detection model corresponding to the high-brightness scene, and processes the environmental image acquired by the image acquisition device based on the scene detection model.

[0089] For example, Figure 6 As shown, the image acquisition device loaded on the vehicle acquires the environmental image around the vehicle. The vehicle uses a preloaded environmental detection model to determine the current scene based on the environmental image. For example, if the current scene is determined to be a low-brightness scene, the scene detection model corresponding to the low-brightness scene is loaded, and the environmental image acquired by the image acquisition device is processed based on the scene detection model.

[0090] The method of this embodiment obtains an environmental image; uses a preloaded environmental detection model to confirm the current scene based on the environmental image; loads a scene detection model that matches the current scene; processes the environmental image based on the scene detection model, and when computing power is constrained, selects a lightweight scene detection model corresponding to the current scene, thereby improving the efficiency of image processing and the respective performance in different scenarios.

[0091] On the basis of the above embodiment, further, before processing the environment image or determining the scene based on the environment image, the environment image may be compressed.

[0092] Specifically, the acquired environmental image is generally color RGB image information, and the image resolution is generally large, for example, 1280×720. When processing the environmental image, the environmental image can be compressed, for example, the resolution can be compressed to 640×360, which can improve processing efficiency when computing power is constrained.

[0093] In an optional embodiment, a preloaded environment detection model is used to extract brightness information from an environment image to determine the current scene.

[0094] For example, the RGB or HSV information of the environment image can be obtained to extract the brightness information in the environment image and then determine the current scene, such as high brightness scene, medium brightness scene, and low brightness scene divided by image brightness. For example, high visibility scene, medium visibility scene, and low visibility scene divided by image visibility.

[0095] In an optional embodiment, a preloaded environment detection model is used to extract brightness information and images from an environment image to determine the current scene.

[0096] Furthermore, the pre-adapted environment detection model can not only extract the brightness information of the environment image, but also extract the image and combine the image and brightness information to determine the current scene.

[0097] Furthermore, a possible implementation of step 102 is as follows:

[0098] Obtain distribution information in the environment image, and use the distribution information to determine a current scene.

[0099] In an optional embodiment, the RGB or HSV distribution information of the environment image obtained in step 101 is read, and the current scene is determined based on the distribution information.

[0100] For RGB distribution information, in an optional embodiment, after obtaining the RGB distribution information in the environmental image, the information of the three R, G, and B channels of the pixel points in the environmental image can be averaged to obtain the average pixel value corresponding to each channel, or the proportion of pixels with brightness values greater than the preset brightness value can be obtained, so as to determine the current scene. For example, if the proportion of pixels with brightness values greater than the preset brightness value is greater than a certain value, it can be determined as a high-brightness scene, such as a daytime scene.

[0101] HSV distribution information is a method of representing points in the RGB color space within an inverted cone. HSV stands for Hue, Saturation, and Value. Hue is the basic attribute of color, commonly referred to as color names such as red and yellow. Saturation refers to the purity of the color; higher values indicate purer colors, while lower values indicate graying, with a value ranging from 0 to 100%. Brightness refers to the brightness of the color, with a value ranging from 0 to 100%.

[0102] In an optional embodiment, after obtaining the HSV distribution information in the environmental image information, the information of the three channels H, S, and V of the pixel points in the environmental image can be averaged to obtain the average pixel value corresponding to each channel, or the proportion of pixels with brightness values greater than the preset brightness value can be obtained, or the proportion of red and yellow light can be obtained to determine the current scene.

[0103] Furthermore, another possible implementation of step 102 is as follows:

[0104] Histogram information in the environment image is counted, and the current scene is determined using the histogram information.

[0105] In an optional embodiment, the RGB or HSV histogram information of the environment image obtained in step 101 is read, and the current scene is determined according to the RGB or HSV histogram.

[0106] In an optional embodiment, for RGB histogram information, in an optional embodiment, after obtaining the environmental image, the R, G, and B channels of the pixels in the environmental image are statistically analyzed to obtain histogram information, thereby determining the current scene based on the histogram information of the R, G, and B channels.

[0107] In an optional embodiment, for HSV histogram information, in an optional embodiment, after obtaining the environmental image, the H, S, and V channels of the pixel points in the environmental image are statistically analyzed to obtain histogram information, thereby determining the current scene based on the histogram information of the H, S, and V channels.

[0108] Furthermore, the current scene may be determined based on the distribution information or histogram information obtained in the aforementioned steps using a pre-trained environment detection model.

[0109] In an optional embodiment, the distribution information or histogram information obtained above may be input into a pre-trained environment detection model to output information of the current scene, thereby determining the current scene.

[0110] Furthermore, another possible implementation of step 102 is as follows:

[0111] According to the environment image, a pre-trained environment detection model is used to determine the current scene.

[0112] In an optional embodiment, the environment image may be directly input into the environment detection model, and the corresponding current scene information may be output.

[0113] Among them, the environment detection model can be established based on a neural network model such as CNN, and trained based on training data to obtain better parameters of the environment detection model.

[0114] This environmental detection model can output only a small number of classification results. To achieve accurate output, the network layer does not require too many parameters. In other words, the neural network used for this judgment step only consumes a small amount of system computing power, and model loading also consumes only a small amount of system bandwidth.

[0115] In other embodiments of the present disclosure, the environment detection model may also be established based on other network models, which is not limited in the embodiments of the present disclosure.

[0116] Furthermore, another possible implementation of step 102 is as follows:

[0117] Acquiring road sign information in the environment image;

[0118] The current scene is determined according to the road sign information.

[0119] Specifically, the road sign information in the environment image is obtained, and the current scene is determined based on the road sign information, such as a city road scene, a highway scene, etc. For example, the road sign information in the environment image information can be obtained through a recognition algorithm.

[0120] Based on the above embodiment, step 104 can be implemented in the following manner:

[0121] If the determined current scene includes multiple scenes, such as a daytime scene, a snowy scene, and a highway scene (for example, multiple scenes can be determined simultaneously based on one environmental image, such as a daytime scene, a snowy scene, and a highway scene), then the scene detection models corresponding to the above multiple scenes can be loaded in sequence, and the environmental image can be processed based on the scene detection models corresponding to the multiple scenes.

[0122] In an optional embodiment, it is assumed that, first, a scene detection model for daytime scene matching is loaded, and the environmental image is processed based on the scene detection model for daytime scene matching to obtain a first detection result; further, a scene detection model for snowy scene matching is loaded, and the information of the first detection result and the environmental image is input into the scene detection model for snowy scene matching, and the first detection result and the environmental image are processed based on the scene detection model for snowy scene matching. The first detection result can be used as prior information to make the obtained second detection result more accurate; further, a scene detection model for highway scene matching is loaded, and the information of the first detection result, the second detection result and the environmental image is input into the scene detection model for highway scene matching, and the first detection result, the second detection result and the environmental image are processed based on the scene detection model for highway scene matching. The first detection result and the second detection result can be used as prior information to make the obtained third detection result more accurate, and finally the target detection result is obtained according to the third detection result, or the target detection result is obtained according to the first detection result, the second detection result and the third detection result.

[0123] In an optional embodiment, obtaining the target detection result can be specifically achieved by:

[0124] The third detection result (or at least one of the first detection result, the second detection result and the third detection result) is filtered using a non-maximum suppression method to obtain the target detection result; the target detection result includes at least one of the following: position information of the target object in the environmental image information, category information of the target object and confidence of the target object in the corresponding category.

[0125] Specifically, the detection results output by the scene detection model include a large amount of target object information, including a lot of duplicate information. For example, there is a lot of location information, some of which overlaps. Methods such as non-maximum suppression can be used to filter the detection results to obtain the final target detection results.

[0126] Ultimately, the location, category, and confidence level of the object of interest in the image can be obtained. This output can be provided as external observation information to downstream modules, such as state estimation and navigation control, to achieve more complex autonomous driving functions.

[0127] On the basis of the above embodiment, further, the following operations may be performed before step 103:

[0128] Acquire training data corresponding to a scene detection model that matches the current scene; the training data includes environmental image data including location information and category information of target objects in different scenes;

[0129] The scene detection model is trained using the training data.

[0130] Specifically, the scene detection models corresponding to different scenes need to be pre-trained to obtain the optimal parameters of the scene detection model.

[0131] To obtain a scene detection model with better performance for different scenarios, such as daytime and nighttime environments, it is necessary to train the model separately using training data corresponding to different scenarios, such as daytime data and nighttime data. Specifically, a batch of training data is collected in advance for different scenarios, such as daytime and nighttime. Each training data set contains an image of the environment and the location and category annotations of the objects of interest in the image. Then, based on the training data corresponding to different scenarios, models are designed and trained separately to obtain the optimal scene detection model for different scenarios.

[0132] In the above specific implementation, during the model training process, the scene detection model is trained using the corresponding training set for each scene. In actual use, the current scene corresponding to the environment is first determined based on the environment image, and then the scene detection model corresponding to the current scene is loaded to perform target detection, thereby improving detection performance and increasing detection efficiency when computing power is constrained.

[0133] Figure 8 FIG. 1 is a flow chart of another embodiment of the target detection method provided by the present disclosure. Figure 8 As shown, the method provided in this embodiment includes:

[0134] Step 201: Acquire an environment image.

[0135] The environmental image may be image information collected by an image acquisition device, such as an environmental image around a vehicle. The environmental image may include multiple images, such as an image that triggers loading of a corresponding scene detection model, or an image used to determine the current scene.

[0136] Step 202: Extract feature information from the environment image.

[0137] Furthermore, before step 202 , the environment image may be compressed.

[0138] Step 203: Determine the current scene based on the feature information in the environment image.

[0139] Specifically, the current scene may be determined based on the environmental image information, for example, scenes at different times, such as a daytime scene or a nighttime scene.

[0140] The acquired environmental image is generally color RGB image information, and the image resolution is generally large, for example, 1280×720. When processing the environmental image information, the environmental image information can be compressed, for example, the resolution can be compressed to 640×360, which can improve processing efficiency when computing power is constrained.

[0141] In an optional embodiment, the current scene, such as a daytime scene or a nighttime scene, can be determined by using feature information extracted from the environment image and an environment detection model.

[0142] The feature information includes at least one of the following: average pixel value, proportion of high brightness value, proportion of red and yellow light, and HSV three-channel statistical histogram.

[0143] The following describes the process of extracting feature information:

[0144] A color image can be composed of three channels (R, G, and B), and a histogram can be extracted for each channel. The average pixel value can be calculated by averaging the three channels. The high brightness value percentage refers to the percentage of pixels with brightness values greater than a preset high brightness value.

[0145] HSV is a method of representing points in the RGB color space within an inverted cone. HSV stands for Hue, Saturation, and Value. Hue is the basic attribute of color, commonly referred to as color names such as red and yellow. Saturation refers to the purity of the color; higher values indicate a purer color, while lower values indicate a grayer color, ranging from 0 to 100%. Brightness refers to the brightness of the color, ranging from 0 to 100%.

[0146] The method of extracting HSV color space features is similar to RGB. The key point is to convert the original image into an image in the HSV color space, and then perform histogram drawing operations on the three channels respectively.

[0147] After converting the image information into the HSV color space, the proportion of red and yellow light can also be obtained.

[0148] The number of feature information of the HSV three-channel statistical histogram may be 3×20=60. In one embodiment, the above four features may be concatenated together to form feature information with a length of 63.

[0149] Furthermore, a pre-trained environment detection model can be used to input the extracted feature information into the environment detection model, and output the corresponding current scene information;

[0150] In other embodiments of the present disclosure, the environment image may be directly input into the environment detection model to output the corresponding current scene information.

[0151] Furthermore, for different time scenes such as daytime and nighttime, or weather scenes such as snowy days, foggy days, rainy days, and sunny days, step 203 can be implemented in the following manner:

[0152] The ambient light intensity of the current scene is determined based on the feature information in the environmental image.

[0153] The current scene is determined based on the ambient light intensity of the current scene.

[0154] In an optional embodiment, a pre-trained environment detection model can be used to input the extracted feature information into the environment detection model, output the ambient light intensity of the current scene, and determine the current scene based on the ambient light intensity. Since the ambient light intensity of different time scenes, such as daytime scenes and nighttime scenes, is different, the current scene can be determined based on the ambient light intensity.

[0155] In one embodiment of the present disclosure, the environment detection model may be trained in advance, which may be achieved in the following manner:

[0156] Acquire training data; the training data includes feature information of multiple environmental images and scene information corresponding to each environmental image, or multiple environmental images and scene information corresponding to each environmental image;

[0157] The pre-established environment detection model is trained using the training data to obtain a trained environment detection model.

[0158] Specifically, the environment detection model can be established through deep learning algorithms, such as convolutional neural network CNN model, VGG model, GoogleNet model, etc. In order to obtain an environment detection model with better recognition performance for different scenes such as daytime scenes and nighttime scenes, it is necessary to train the environment detection model with training data corresponding to different scenes such as daytime scenes and nighttime scenes to obtain better parameters of the environment detection model.

[0159] Step 204: Load a scene detection model that matches the current scene.

[0160] Specifically, this step loads the corresponding scene detection model into the memory of the device based on the current scene determined in step 203 .

[0161] Step 205: Process the environment image based on the scene detection model to obtain a first detection result.

[0162] Specifically, the environment image is processed based on the scene detection model corresponding to the current scene, for example, the position of the target object in the environment image, the category to which the target object belongs, and the confidence level in the category are identified.

[0163] The scene detection model can be a pre-trained machine learning model, such as a convolutional neural network model. During model training, the scene detection model is trained for each scene using a corresponding training dataset. During detection, the environment image information is input into the scene detection model corresponding to the current scene. After being processed through several convolutional layers and pooling layers, a first detection result is output.

[0164] Step 206: Filter the first detection result using a non-maximum suppression method to obtain a target detection result; the target detection result includes at least one of the following: position information of the target object in the environment image, category information of the target object, and confidence of the target object in the corresponding category.

[0165] Specifically, the detection results output by the scene detection model include a large amount of target object information, including a lot of duplicate information. For example, there is a lot of location information, some of which overlaps. Methods such as non-maximum suppression can be used to filter the detection results to obtain the final target detection results.

[0166] Ultimately, the location, category, and confidence level of the object of interest in the image can be obtained. This output can be provided as external observation information to downstream modules, such as state estimation and navigation control, to achieve more complex autonomous driving functions.

[0167] Furthermore, in one embodiment of the present disclosure, Figure 5 As shown, if the current scene includes the first scene and the second scene, step 205 can be implemented as follows:

[0168] Step 2051: Process the environment image based on the scene detection model matched with the first scene to obtain a first detection result;

[0169] Step 2052: Process the first detection result based on the scene detection model matched with the second scene to obtain a second detection result;

[0170] Step 2053: Obtain a target detection result based on the second detection result.

[0171] Specifically, the scene can be determined based on the environmental image. For example, the current scene includes different time scenes such as day and night, or weather scenes such as snowy, foggy, rainy, and sunny days, or road conditions such as highways, rural roads, and urban roads.

[0172] It is assumed that it is determined based on the environment image that the current scene includes at least two scenes, for example, a first scene and a second scene.

[0173] Assuming the first scene is a daytime scene within the time scenario, the environment image is processed based on the scene detection model matched to the first scene to obtain a first detection result. Furthermore, the first detection result is input into a second scene, for example, a snowy scene within the weather scenario. The first detection result is processed based on the scene detection model matched to the second scene to obtain a second detection result. Finally, a target detection result is obtained based on the second detection result. Because the environment image has already been processed using the scene detection model matched to the first scene when the detection model matched to the second scene is used for target detection, prior information is obtained, making the final target detection result more accurate.

[0174] In an optional embodiment, the first scene and the second scene may be a high-brightness scene and a low-brightness scene, respectively.

[0175] In other embodiments of the present disclosure, the scene detection model based on the second scene matching may be used for processing first, and then the scene detection model based on the first scene matching may be used for processing. The embodiments of the present disclosure are not limited to this.

[0176] Figure 9 For the rest of the steps, see Figure 8 Explanation, no further elaboration is given here.

[0177] The method of this embodiment obtains an environmental image; confirms the current scene based on the environmental image; loads a scene detection model that matches the current scene; processes the environmental image based on the scene detection model, and when computing power is constrained, selects a lightweight scene detection model corresponding to the current scene, thereby improving the efficiency of image processing and the respective detection performance in different scenarios.

[0178] like Figure 10 As shown, a vehicle is also provided in an embodiment of the present disclosure, which is equipped with a camera 11, a memory 12, and a processor 13. The memory 12 is used to store instructions, and the instructions are executed by the processor 13 to implement any one of the methods in the aforementioned method embodiments.

[0179] The vehicle provided in this embodiment is used to execute the image processing method provided in any of the aforementioned embodiments. The technical principles and technical effects are similar and will not be described in detail here.

[0180] like Figure 11 As shown, an embodiment of the present disclosure further provides a drone, which is equipped with a camera 21, a memory 22, and a processor 23. The memory 22 is used to store instructions, and the instructions are executed by the processor 23 to implement any one of the methods in the aforementioned method embodiments.

[0181] The drone provided in this embodiment is used to execute the image processing method provided in any of the aforementioned embodiments. The technical principles and technical effects are similar and will not be described in detail here.

[0182] like Figure 12 As shown, an embodiment of the present disclosure also provides an electronic device that can be communicatively connected to a camera device. The electronic device includes a memory 32 and a processor 31. The memory 32 is used to store instructions, and the instructions are executed by the processor 31 to implement any one of the methods in the aforementioned method embodiments.

[0183] The electronic device provided in this embodiment is used to execute the image processing method provided in any of the aforementioned embodiments. The technical principles and technical effects are similar and will not be described in detail here.

[0184] like Figure 13 As shown, a handheld gimbal is also provided in an embodiment of the present disclosure, and the handheld gimbal includes: a camera device 41, a memory 42, and a processor 43, the memory 42 is used to store instructions, and the instructions are executed by the processor 43 to implement any one of the methods in the aforementioned method embodiments.

[0185] The handheld gimbal provided in this embodiment is used to execute the image processing method provided in any of the aforementioned embodiments. The technical principles and technical effects are similar and will not be described in detail here.

[0186] like Figure 14 As shown, a mobile terminal is also provided in an embodiment of the present disclosure, which includes: a camera 51, a memory 52, and a processor 53, wherein the memory 52 is used to store instructions, and the instructions are executed by the processor 53 to implement any one of the methods in the aforementioned method embodiments.

[0187] The mobile terminal provided in this embodiment is used to execute the image processing method provided in any of the aforementioned embodiments. The technical principles and technical effects are similar and will not be described in detail here.

[0188] A computer-readable storage medium is also provided in an embodiment of the present disclosure, on which a computer program is stored. When the computer program is executed by a processor, the corresponding method in the aforementioned method embodiment is implemented. The specific implementation process can be found in the aforementioned method embodiment. The implementation principle and technical effects are similar and will not be repeated here.

[0189] The present disclosure also provides a program product comprising a computer program (i.e., execution instructions) stored in a readable storage medium. A processor can read the computer program from the readable storage medium and execute the computer program to perform the target detection method provided in any of the aforementioned method embodiments.

[0190] The present disclosure also provides a vehicle, including:

[0191] vehicle body; and

[0192] The electronic device described in any of the above embodiments is mounted on the vehicle body. Its implementation principle and technical effects are similar to those of the method embodiment and will not be repeated here.

[0193] The present disclosure also provides a drone, including:

[0194] fuselage; and

[0195] The electronic device described in any of the above embodiments is mounted on the vehicle body. Its implementation principle and technical effects are similar to those of the method embodiment and will not be repeated here.

[0196] Figure 15: This is a schematic diagram of the memory usage ratio during the model loading process provided in the embodiment of this specification. The environment detection model is always loaded. For example, it can be always loaded in the processor memory during the operation of the mobile platform. It only needs to judge the current environment, and the system resources required are relatively small. The environment detection model only needs to identify and output the category information of the current environment, and the category information is used to load the scene detection model. The scene detection model is used to detect objects around the mobile platform. On the one hand, the environment detection model and the scene model can greatly reduce the resources occupied by the loaded model; on the other hand, the resources occupied by the scene model will be greater than the environment detection model. As an optional embodiment, the environment detection model can be a trained neural network model that can output the recognition classification results based on the input image information, such as daytime, nighttime, rain, snow, and fog. As an optional embodiment, the environment detection model can be a trained neural network model that can output the recognition two-dimensional classification results based on the input image information, such as daytime-rain, nighttime-rain, and daytime-fog. As an optional embodiment, the environment detection model can be a trained neural network model that can output a three-dimensional classification result based on the input image information, and the dimensions include but are not limited to weather-climate brightness, such as daytime-rain-darkness, nighttime-rain-darkness, and daytime-sunny-brightness. As an optional embodiment, the environment detection model can be a trained neural network model that can output a four-dimensional or even high-dimensional classification result based on the input image information, and the dimensions include but are not limited to weather-climate brightness, such as daytime-rain-darkness-road, nighttime-rain-darkness-road, and daytime-sunny-brightness-tunnel. As an optional embodiment, the environment detection model can be a judgment function based on the output parameters of the image sensor, for example, judging whether it is daytime or night based on the brightness information of the image.

[0197] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0198] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present disclosure, rather than to limit them. Although the embodiments of the present disclosure have been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present disclosure.

Claims

1. A machine vision-based image processing method, applied to a movable platform equipped with an image acquisition device, characterized in that: The method comprises: Get the environment image; Using a preloaded environment detection model, determining a current scene based on the environment image; Loading a scene detection model that matches the current scene; processing an environment image based on the scene detection model; If the current scene includes a first scene and a second scene, processing the environment image based on the scene detection model includes: Processing the environment image based on the scene detection model of the first scene matching to obtain a first detection result; the first detection result includes a position of a target object in the environment image, a category to which the target object belongs, and a confidence level in the category; Processing the first detection result based on the scene detection model matched with the second scene to obtain a second detection result; According to the second detection result, a target detection result is obtained; the target detection result includes at least one of the following: position information of the target object in the environmental image information, category information of the target object, and confidence of the target object in the corresponding category.

2. The method according to claim 1, characterized in that The current scene includes at least two scenes divided according to image brightness.

3. The method according to claim 2, characterized in that The current scene includes a high-brightness scene and a low-brightness scene.

4. The method according to claim 2, characterized in that The current scene includes a high-brightness scene, a medium-brightness scene, and a low-brightness scene.

5. The method according to claim 1, wherein The current scene includes at least two scenes divided according to image visibility.

6. The method according to claim 5, characterized in that The current scene includes a high-visibility scene and a low-visibility scene.

7. The method according to claim 5, characterized in that The current scene includes a high-visibility scene, a medium-visibility scene, and a low-visibility scene.

8. The method according to claim 5, characterized in that The at least two scenes divided according to image visibility include a haze scene and a dust scene.

9. The method according to claim 1, characterized in that The current scene includes at least two scenes divided according to image texture information.

10. The method according to claim 9, characterized in that The scenes divided according to image texture information include weather information.

11. The method according to claim 10, characterized in that The weather information includes rain, snow, fog and sandstorm weather information.

12. The method according to claim 1, characterized in that The preloaded environment detection model is used to extract brightness information from the environment image and determine the current scene.

13. The method according to claim 1, wherein The preloaded environment detection model is used to extract brightness information and images from the environment image to determine the current scene.

14. The method according to claim 1, wherein The preloaded environment detection model is always in a loading state during the image acquisition process.

15. The method according to claim 14, characterized in that The scene detection model matching the current scene is switched and loaded as the current scene changes.

16. The method according to claim 15, characterized in that The scene detection model matching the current scene is not removed from the memory due to switching loading.

17. The method according to claim 1, wherein The preloaded environment detection model and the scene detection model are in different threads.

18. The method according to claim 17, characterized in that The preloaded environment detection model performs inter-thread communication via a callback function.

19. The method according to claim 1, wherein Processing the environment image based on the scene detection model includes: acquiring object information in the environment image.

20. The method according to claim 19, wherein The obtained object information is filtered using a non-maximum suppression method to obtain a target detection result.

21. The method according to claim 19, wherein The object information includes: position information of the target object in the environment image, category information of the target object, and confidence of the target object in the corresponding category.

22. The method according to claim 21, characterized in that The determining the current scene according to the environment image includes: extracting feature information from the environment image; The current scene is determined according to feature information in the environment image.

23. The method according to claim 22, characterized in that The determining the current scene according to the feature information in the environment image includes: Determining the ambient light intensity of the current scene based on feature information in the environmental image; The current scene is determined according to the ambient light intensity of the current scene.

24. The method according to claim 22, characterized in that Before extracting the feature information from the environment image, the method further includes: The environment image is compressed.

25. The method according to claim 22, wherein The feature information includes at least one of the following: average pixel value, proportion of high brightness value, proportion of red and yellow light, and HSV three-channel statistical histogram.

26. The method according to claim 22, characterized in that Determining the current scene according to the environment image includes: Acquiring road sign information in the environment image; The current scene is determined according to the road sign information.

27. The method according to claim 1, wherein Before processing the environment image based on the scene detection model, the method further includes: Acquire training data corresponding to a scene detection model that matches the current scene; the training data includes environmental image data including location information and category information of target objects in different scenes; The scene detection model is trained using the training data.

28. A vehicle, characterized in that: The vehicle is equipped with a camera device, a memory, and a processor, wherein the memory is used to store instructions, and the instructions are executed by the processor to implement the method according to any one of claims 1 to 27.

29. A drone, characterized in that: The drone is equipped with a camera, a memory, and a processor, wherein the memory is used to store instructions, and the instructions are executed by the processor to implement the method according to any one of claims 1 to 27.

30. An electronic device, characterized in that: The electronic device is communicatively connected to the camera device, and includes a memory and a processor, wherein the memory is used to store instructions, and the instructions are executed by the processor to implement the method according to any one of claims 1 to 27.

31. A handheld gimbal, characterized in that: The handheld gimbal includes: a camera device, a memory, and a processor, wherein the memory is used to store instructions, and the instructions are executed by the processor to implement the method according to any one of claims 1-27.

32. A mobile terminal, characterized in that: The mobile terminal includes: a camera device, a memory, and a processor, wherein the memory is used to store instructions, and the instructions are executed by the processor to implement the method according to any one of claims 1 to 27.

Citation Information

Patent Citations

  • Method and device for scene recognition

    CN103945088A

  • Object identification method and device based on deep learning neural network

    CN107316035A

  • Environmental perception adaptive image recognition method and device

    CN110059594A