Unlearned region estimation device and integrated program

Through the unlearned area estimation device, the first area estimation unit and the second area estimation unit combine the photography scene to integrate the area, the accuracy problem of detecting and learning objects in the prior art is solved, and the accuracy and safety of detection are improved.

CN120339993APending Publication Date: 2025-07-18HONDA MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510396641.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect drops or obstacles in learned object types, resulting in reduced detection accuracy, especially in advanced driving assistance systems and autonomous driving, which are prone to missed inspections.

Method used

The unlearned area estimation device is used to estimate the object area by the first area estimation unit and the second area estimation unit, and detect it using different features. The estimation area synthesis unit combines the area estimation of different features according to the photography scene, so as to improve detection accuracy.

Benefits of technology

It improves the accuracy of detection of obstacles or drops on the road, reduces the missed detection rate, and enhances the safety of autonomous driving or assisted driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339993A_ABST
    Figure CN120339993A_ABST
Patent Text Reader

Abstract

The invention discloses an unlearned area estimation device and an integrated program, which can improve the accuracy of detecting obstacles or falling objects on a road. The method includes: estimating, by a first region estimation unit, an object region of a dropped object from an input signal; estimating, by a second region estimation unit, a region having a characteristic different from that of the first region estimation unit on the basis of a signal identical to or different from the input signal; an estimated region integration unit integrates the regions estimated by the first region estimation unit and the second region estimation unit on the basis of the photographic scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to an unlearned area estimation device and an integrated program. Background Art

[0002] In advanced driving assistance systems (ADAS) and autonomous driving, when there are falling objects or obstacles on the road, it is necessary to identify the falling objects or obstacles to avoid driving. However, since the number of samples of data on falling objects that can be collected is small and diverse, it is difficult to learn about falling objects in advance.

[0003] Currently, the detection method for falling objects or obstacles on the road is to use a neural network to perform region division on the collected images for detection. However, for learned object categories, the neural network will not identify them as falling objects or obstacles. Therefore, the current method is difficult to detect falling objects or dropped objects of learned object categories, resulting in missed detections and reducing the detection accuracy. Summary of the Invention

[0004] This application expects to provide an unlearned area estimation device and an integrated program that can improve the accuracy of detecting obstacles or falling objects on the road.

[0005] The technical solution of this application is implemented as follows:

[0006] In a first aspect, this application provides an unlearned area estimation device, and the device includes:

[0007] A first area estimation unit, configured to estimate the object area of a falling object according to an input signal;

[0008] A second area estimation unit, configured to estimate an area different from the characteristics of the first area estimation unit according to a signal that is the same as or different from the input signal;

[0009] An estimated area integration unit, configured to integrate the areas estimated by the first area estimation unit and the second area estimation unit according to the shooting scene.

[0010] Optionally, the unlearned area estimation device includes an area estimation unit, and the estimated area integration unit of the estimated area integration unit includes:

[0011] The area estimation unit is configured to perform area estimation with two or more different characteristics according to claim 1;

[0012] The estimated area integration unit is configured to integrate multiple estimated areas.

[0013] Optionally, the estimated area integration unit further includes: a photography scene determination unit,

[0014] The photography scene determination unit is configured to determine the scene for obtaining the input signal according to claim 1;

[0015] The estimated area integration unit is configured to integrate the estimated area according to the output from the photography scene determination unit.

[0016] Optionally, the first area estimation unit is configured to perform image segmentation according to the input signal and estimate the object area of the dropped object;

[0017] The second area estimation unit is configured to perform depth estimation according to a signal that is the same as or different from the input signal and estimate an area with characteristics different from those of the first area estimation unit.

[0018] Optionally, the second area estimation unit is configured to perform depth estimation according to a signal that is the same as or different from the input signal, determine the depth value corresponding to the pixel;

[0019] Determine the target object according to the depth value corresponding to the pixel, and predict the normal vector information corresponding to the pixel;

[0020] Use the normal vector information to determine the area corresponding to the target object.

[0021] Optionally, the device includes: a mixture of experts model;

[0022] The mixture of experts model is configured to integrate the estimated areas output by multiple area estimation units according to the photography scene through the estimated area integration unit; the multiple area estimation units include the first area estimation unit and the second area estimation unit.

[0023] Optionally, the photography scene determination unit is configured to determine the scene for obtaining the input signal according to claim 1 according to a scene understanding algorithm.

[0024] Optionally, the unlearned area estimation device includes multiple area estimation units;

[0025] Each area estimation unit among the multiple area estimation units is configured to perform area estimation with two or more different characteristics according to claim 1 based on its own input signal.

[0026] Optionally, the input signal includes: at least one signal data collected by at least one sensor; the at least one sensor includes at least one of an image sensor and a lidar sensor.

[0027] Second aspect, the present application provides an unlearned area estimation method, the method comprising:

[0028] Through a first area estimation unit, estimating an object area of a dropped object based on an input signal;

[0029] Through a second area estimation unit, estimating an area different from the characteristics of the first area estimation unit based on a signal that is the same as or different from the input signal;

[0030] Through an estimated area integration unit, integrating the areas estimated by the first area estimation unit and the second area estimation unit according to a photographing scene.

[0031] Optionally, the method further comprises:

[0032] Through an area estimation unit, performing area estimation with two or more different characteristics according to claim 1;

[0033] Through an estimated area integration unit in the estimated area integration unit, by integrating a plurality of estimated areas.

[0034] Optionally, the estimated area integration unit further comprises: a photographing scene determination unit; the integrating a plurality of estimated areas includes:

[0035] Determining a scene for obtaining the input signal according to claim 1;

[0036] Through the estimated area integration unit, integrating the estimated areas according to the output from the photographing scene determination unit.

[0037] Optionally, the estimating an object area of a dropped object based on the input signal includes:

[0038] Performing image segmentation according to the input signal to estimate an object area of a dropped object;

[0039] The estimating an area different from the characteristics of the first area estimation unit based on a signal that is the same as or different from the input signal includes:

[0040] Performing depth estimation according to a signal that is the same as or different from the input signal to estimate an area different from the characteristics of the first area estimation unit.

[0041] Optionally, performing depth estimation according to a signal that is the same as or different from the input signal to estimate an area different from the characteristics of the first area estimation unit includes:

[0042] Performing depth estimation according to a signal that is the same as or different from the input signal to determine a depth value corresponding to a pixel;

[0043] Determine a target object based on the depth value corresponding to a pixel, and predict the normal vector information corresponding to the pixel;

[0044] Use the normal vector information to determine the region corresponding to the target object.

[0045] Optionally, the determining the scene for obtaining the input signal according to claim 1 includes:

[0046] Determine the scene of the input signal according to claim 1 according to a scene understanding algorithm.

[0047] In a third aspect, the present application provides an electronic device, including a memory and a processor; wherein,

[0048] The memory is used to store executable instructions;

[0049] When the processor is used to execute the executable instructions stored in the memory, the method for estimating an unlearned area provided in the embodiments of the present application is implemented.

[0050] In a fourth aspect, the present application provides a computer-readable storage medium, storing executable instructions, which are used to cause a processor to implement the method for estimating an unlearned area provided in the embodiments of the present application when executed.

[0051] In a fifth aspect, the embodiments of the present application provide an integrated program, including a computer program or instructions, and when the computer program or instructions are executed by a processor, the method for estimating an unlearned area provided in the embodiments of the present application is implemented.

[0052] The present application provides an unlearned area estimation device and an integrated program, which can estimate the object area of a dropped object according to an input signal through a first area estimation unit; estimate an area different from the characteristics of the first area estimation unit according to a signal same as or different from the input signal through a second area estimation unit; and integrate the areas estimated by the first area estimation unit and the second area estimation unit according to a shooting scene through an estimated area integration unit. In the embodiments of the present application, through the first area estimation unit and the second area estimation unit, the object areas of the dropped objects are estimated respectively according to the input signal and a signal same as or different from the input signal to obtain areas with different characteristics, and through the estimated area integration unit, according to the shooting scene, the area estimated by the first area estimation unit and the area estimated by the second area estimation unit are integrated, so as to realize the estimation of the object area by adaptively combining areas with different characteristics estimated by different estimation parts according to the shooting scene, thereby enabling complementary use of different feature estimation methods to detect obstacles or dropped objects on the road, reducing missed detections, and improving the accuracy of obstacle or dropped object detection. Description of the Drawings

[0053] Figure 1 It is an example of a front road surface image collected by a vehicle camera;

[0054] Figure 2 It is the classification prediction diagram of the current neural network for the front road surface image;

[0055] Figure 3 It is an optional process schematic diagram of the unlearned area estimation method provided by the embodiment of the present application;

[0056] Figure 4 It is an optional schematic diagram of the depth map output by the second area estimation unit provided by the embodiment of the present application;

[0057] Figure 5 It is an optional network processing flow schematic diagram of the unlearned area estimation method provided by the embodiment of the present application;

[0058] Figure 6 It is an optional network processing flow schematic diagram of the unlearned area estimation method provided by the embodiment of the present application;

[0059] Figure 7 It is an optional network processing flow schematic diagram of the unlearned area estimation method provided by the embodiment of the present application;

[0060] Figure 8 It is an optional network processing flow schematic diagram of the unlearned area estimation method provided by the embodiment of the present application;

[0061] Figure 9 It is an optional network processing flow schematic diagram of the unlearned area estimation method provided by the embodiment of the present application;

[0062] Figure 10 It is an optional structural schematic diagram of the unlearned area estimation device provided by the embodiment of the present application;

[0063] Figure 11 It is an optional structural schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0064] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0065] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0066] In the following description, the terms "first", "second", and "third" only distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first", "second", and "third" can be interchanged in a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0068] Currently, the detection method for objects or obstacles dropped on the road is to use a neural network to perform region division on the collected images for detection. However, for the object categories that have been learned, the neural network will not recognize them as dropped objects or obstacles. Exemplarily, for the Figure 1 front road surface image collected by the vehicle's camera as shown, the obstacles on the road are bicycle 11 and bicycle 12. Since the neural network has learned the features of bicycles, it will recognize bicycle 11 with obvious appearance features and classify the pixels of bicycle 11 as bicycles in the classification prediction map output as Figure 2 shown, forming the image region 21 classified as bicycles in Figure 2 . And because the appearance of bicycle 12 is very different from the appearance of bicycles learned by the neural network, it will be regarded as a dropped object or obstacle, forming the image region 22 classified as an obstacle in Figure 2 . It can be seen that the current method is difficult to detect dropped objects or fallen objects of the learned object categories, which may cause missed detections, reduce the accuracy of detection, and further reduce the safety of autonomous driving or assisted driving.

[0069] The embodiments of the present application provide an unlearned region estimation device and an integrated program, which can improve the accuracy of detecting obstacles or dropped objects on the road, such as dropped objects or obstacles on the road. The unlearned region estimation method of the embodiments of the present application can be as Figure 3 shown, including S101 - S103, as follows:

[0070] S101. Through the first region estimation unit, estimate the object region of the dropped object according to the input signal.

[0071] S102. Through the second region estimation unit, estimate the region different from the feature of the first region estimation unit according to a signal that is the same as or different from the input signal.

[0072] In the embodiments of the present application, the input signal may include an image collected by an image sensor or point cloud data scanned by a lidar. The first region estimation unit or the second estimation unit may include: a neural network for image segmentation tasks, a neural network for depth estimation, or any one of at least one neural network for object detection based on the signals of at least one sensor. It should be noted that the first region estimation unit and the second estimation unit are neural networks for estimating the object region of the road debris through different feature estimation methods. That is to say, the first region estimation unit and the second region estimation unit are neural networks for different feature processing processes respectively.

[0073] In some embodiments, taking the first region estimation unit as a neural network for image segmentation tasks as an example, the first region estimation unit may extract features from the image, distinguish at least one image region corresponding to at least one type of object in the image according to the extracted features, and output the category corresponding to each pixel point in the image, so as to implement the image segmentation task. Exemplarily, the first region estimation unit may perform pixel-level semantic segmentation on the image and output a semantic segmentation map, and the semantic segmentation map includes the semantic category labels corresponding to each pixel point in the input image. Exemplarily, the first region estimation unit may include: a neural network with a Mask2Former structure.

[0074] Through the first region estimation unit of the image segmentation task, the image region corresponding to the learned object can be recognized and divided from the input image. If the image contains an object that the first region estimation unit has not learned, the first region estimation unit can segment the image region corresponding to the unlearned object and mark the category of the pixels included in the image region as a category representing unlearning (or a category representing abnormality). That is to say, through the first region estimation unit, the object region corresponding to the unlearned object in the image can be detected, so as to realize the detection of unlearned obstacles or debris on the road.

[0075] In the embodiments of the present application, the signals processed by the second region estimation unit may include signals that are the same as or different from the input signals of the first region estimation unit. Exemplarily, the same image may be processed by the first region estimation unit and the second region estimation unit, or images collected from different perspectives of the road may be processed by the first region estimation unit and the second region estimation unit respectively.

[0076] In some embodiments, taking the second region estimation unit as an example of a neural network that performs depth estimation tasks, the second region estimation unit performs depth estimation on the input image and outputs depth information corresponding to each pixel in the image. Thus, three-dimensional objects in the image can be recognized based on the depth information corresponding to each pixel, and then the three-dimensional objects existing on the road can be determined, thereby realizing the detection of obstacles or dropped objects on the road. Exemplarily, the depth information corresponding to a pixel may include the relative distance between the pixel and the shooting source.

[0077] Exemplarily, the second region estimation unit may include a depth estimation network based on monocular vision or binocular vision. The second region estimation unit can output a depth map according to the depth information of each pixel. Taking the second region estimation unit's Figure 1 depth estimation as an example, the output depth map can be as Figure 4 shown. According to Figure 4 the depth information corresponding to each pixel in, three-dimensional objects 41 and 42 on the vehicle driving route can be detected or recognized, thereby realizing the detection of obstacles or dropped objects on the driving route.

[0078] In the embodiments of the present application, the region output by the second region estimation unit is a region with different characteristics from the region output by the first region estimation unit. Exemplarily, the object region output by the first region estimation unit is a region representing the semantic characteristics of the dropped object. The region output by the second region estimation unit is a region representing the depth characteristics of the dropped object. Or, the first region estimation unit and the second region estimation unit can respectively perform object detection on the two-dimensional image and three-dimensional point cloud data collected by different sensors, and output different types of object detection regions (bounding box) to represent the object regions of dropped objects with different characteristics.

[0079] S103. Through the estimated region integration unit, integrate the regions estimated by the first region estimation unit and the second region estimation unit according to the shooting scene.

[0080] In the embodiments of the present application, the estimated region integration unit is used to integrate the regions estimated by the first region estimation unit and the second region estimation unit according to the shooting scene. Among them, the shooting scene includes scene factors for signal collection of the dropped object. Here, the scene factors may include object feature factors of the collected object, that is, the dropped object, or the collection environment factors where the signal collection is located. In some embodiments, the shooting scene may include: the appearance of the dropped object (such as the shape and / or size of the dropped object), or the distance between the dropped object and the electronic device, or the current shooting environment where the dropped object is located (such as weather, light brightness, etc.). Specific selection is made according to the actual situation, and the embodiments of the present application do not make limitations.

[0081] In some embodiments, the presumption area integration unit may include a photography scene determination unit. Exemplarily, the photography scene determination unit may include a neural network for determining the photography scene. The input signal of the presumption area integration unit may be the same signal as that of the first area presumption unit and / or the second area presumption unit. The photography scene determination unit in the presumption area integration unit is used to determine the photography scene by using a scene understanding algorithm according to the signal input to the first area presumption unit and / or the signal input to the second area presumption unit. Alternatively, the presumption area integration unit may also be connected to an independent photography scene determination unit. Exemplarily, the input signal of the presumption area integration unit may be the output signal of the photography scene determination unit, and the input signal of the photography scene determination unit may be the same signal as that of the first area presumption unit and / or the second area presumption unit. Thus, the photography scene can be determined by using a scene understanding algorithm according to the signal input to the first area presumption unit and / or the signal input to the second area presumption unit, and then output to the presumption area integration unit.

[0082] In the current related technologies, for the problem of merging the inference results of two different networks into one, generally a simple threshold is used for merging. However, the threshold belongs to the hyperparameters of the neural network, so it is very difficult to determine a suitable threshold for merging. Moreover, different input environments require different applicable thresholds. Therefore, the simple threshold merging method in the current related technologies is usually not applicable to actual applications.

[0083] In the embodiments of the present application, the presumption area integration unit may be implemented in the form of a neural network. The presumption area integration unit may be trained together with the untrained first area presumption unit and the untrained second area presumption unit, or the presumption area integration unit may be trained based on the pre-trained first area presumption unit and the second area presumption unit. The specific selection is made according to the actual situation, and the embodiments of the present application do not make any limitations. During the network training process, the presumption area integration unit learns the optimal parameters for integrating the areas output by the first area presumption unit and the areas output by the second area presumption unit under various photography scenes. In this way, after the training is completed, the presumption area integration unit can be used to automatically perform the optimal integration of the areas presumed by the first area presumption unit and the areas presumed by the second area presumption unit according to the photography scene, so as to accurately detect the area where the dropped object or obstacle is located. Moreover, under different photography scenes, the optimal parameters for area integration can be adaptively adjusted to improve the system robustness.

[0084] It can be understood that in the embodiments of the present application, the first region estimation unit and the second region estimation unit respectively perform object region estimation of the dropped object based on the input signal and a signal that is the same as or different from the input signal, obtain regions with different features, and the estimated region integration unit integrates the regions estimated by the first region estimation unit and the regions estimated by the second region estimation unit according to the photography scene, realizing the estimation of the object region by adaptively combining regions with different features estimated by different estimation parts according to the photography scene. Thus, different feature estimation methods can be complementarily used to detect obstacles or dropped objects on the road, reducing missed detections and improving the accuracy of obstacle or dropped object detection.

[0085] In some embodiments, the neural network model including the first region estimation unit, the second region estimation unit, and the estimated region integration unit can be a Mixture of Experts (MoE). The first region estimation unit and the second region estimation unit are equivalent to the expert models in the MoE, and the estimated region integration unit is equivalent to the gating network or gating model in the MoE. The gating network or gating model is used to balance and combine the outputs of the expert models, determine the weights of each expert model for the final prediction, and perform result combination. Exemplarily, the first region estimation unit may include: a semantic segmentation neural network, and the second region estimation unit may include: a depth estimation neural network. For the case where the dropped object on the road is as shown by the object 50 in Figure 5 the object 50 in the figure, the appearance of the object 50 is quite different from the objects learned by the semantic segmentation neural network 52. The estimated region integration unit 51 can assign a higher weight to the region output by the semantic segmentation neural network 52 between the semantic segmentation neural network 52 and the depth estimation neural network 53, and integrate the regions output by the semantic segmentation neural network 52 and the depth estimation neural network 53 according to the assigned weights, and output the finally determined object region 54 of the dropped object. For the case where the dropped object on the road is as shown by the objects 60-1 and 60-2 in Figure 6 the figure, the appearances of the objects 60-1 and 60-2 are quite similar to the objects learned by the semantic segmentation neural network 62. The estimated region integration unit 61 can assign a higher weight to the region output by the depth estimation neural network 63 between the semantic segmentation neural network 62 and the depth estimation neural network 63, and integrate the regions output by the semantic segmentation neural network 62 and the depth estimation neural network 63 according to the assigned weights, and output the finally determined object regions 64-1 and 64-2 of the dropped object.

[0086] In some embodiments, the photography scene may include: the shape and / or size of the dropped object.

[0087] In some embodiments, the presumed area integration unit may adaptively select, in terms of pixels, the feature information output by the area presumption unit with a higher weight from the areas output by the first area presumption unit and the second presumption unit according to the shape and / or size of the dropped object, and perform integration using a weighted sum to output the object area of the dropped object.

[0088] Exemplarily, as Figure 7 shown, the input signal of the first area presumption unit and the signal of the second area presumption unit may be the same color (RGB) image. The first area presumption unit may include a semantic segmentation network 71. The semantic segmentation network 71 determines, through semantic segmentation, the pixel label corresponding to each pixel in the RGB image. The pixel label indicates whether the pixel belongs to the pixel of the dropped object (an object not learned). In some embodiments, the pixel label may include the anomaly score of the pixel. Exemplarily, the anomaly score corresponding to the pixel represents the probability that the pixel is a dropped object. The greater the difference between the appearance of the pixel and the object categories learned by the semantic segmentation network, the higher the probability that the pixel is a dropped object. Here, the pixel label corresponding to each pixel in the RGB image may be regarded as the pixel label image output by the semantic segmentation network.

[0089] As Figure 7 shown, the second area presumption unit may include an object detection network 72 based on monocular depth estimation and normal vector. The object detection network 72 obtains, through a monocular depth estimation (such as MonoDepth depth estimation) algorithm, the depth value corresponding to each pixel in the RGB image, and determines the depth information corresponding to each pixel based on the depth value. In some embodiments, the object detection network 72 may predict the normal vector information corresponding to each pixel in the image. Exemplarily, the normal vector information may include the surface normal vector of the object where the pixel is located. Using the depth value corresponding to each pixel, the target object can be preliminarily determined. Using the normal vector information, the area corresponding to the target object can be further extracted. That is, the depth information corresponding to each pixel can be determined by combining the normal vector information and the depth value, so as to determine the depth information image output by the object detection network 72. The depth information image includes the estimated area obtained by the depth estimation of the second area unit. In this way, the depth value is post-processed or optimized using the surface normal vector. For example, the depth value can be optimized using the normal vector information through smoothness constraints and surface continuity constraints to ensure the consistency between the surface normal vector and the depth value.

[0090] As Figure 7As shown, the estimated area integration unit may include a neural network (such as UNet) with an Encoder-Decoder structure, like the gate model 73. Exemplarily, during the training process, the gate model 73 can obtain the weight w1 corresponding to the semantic segmentation network 71 and the weight w2 corresponding to the object detection network 72 through supervised learning. In this way, through the gate model 73, based on the pixel label image output by the semantic segmentation network 71 and the depth information image output by the object detection network 72, the pixel labels and depth information can be weighted and merged pixel by pixel with w1 and w2 respectively, and the merged result is output. The merged result includes the object area of the finally determined dropped object.

[0091] It can be understood that the embodiment of the present application realizes regarding objects not included in the learning as dropped objects and identifying them. Therefore, it is completely unnecessary to have a large amount of learning data for a large number of different types of dropped objects. At the same time, by integrating object detection information based on depth estimation, the problem of misdetecting objects similar to those learned in the unlearned object detection model is avoided, and the accuracy of obstacle or dropped object detection is improved.

[0092] In some embodiments, the first area estimation unit and the second area estimation unit may be area estimation units among multiple area estimation units, and the multiple area estimation units may further include other area estimation units. The multiple area estimation units can output estimated areas with multiple different features. The estimated area integration unit can also integrate the areas with multiple different features output by the multiple area estimation units. Exemplarily, the estimated area integration unit performs integration through multiple estimated areas, executes area estimation with two or more different features, and determines the final object area of the dropped object.

[0093] In some embodiments, the photography scene may include: the photography environment. Exemplarily, the photography environment may include the weather environment and / or the light environment during photography, such as sunny, night, day, rainy, etc. Specifically, it is selected according to the actual situation, and the embodiments of the present application do not make limitations.

[0094] In some embodiments, the estimated region integration unit may adaptively select the output of a computationally efficient region estimation unit from multiple region estimation units according to the photography scene to achieve high-performance prediction. Each region estimation unit among the multiple region estimation units may be a computationally efficient lightweight model (such as MobileNet), and each region estimation unit is good at detecting falling objects in at least one photography environment. Among them, the estimated region integration unit may be a lightweight model with the same structure (such as a gate model). The estimated region integration unit can automatically learn the photography environments that each region estimation unit among the multiple region estimation units is good at processing during the training process. Thus, in different photography environments, the estimated region integration unit can output the weights corresponding to each region estimation unit in that photography environment. Exemplarily, the weight corresponding to this photography environment may represent the importance level (which can be a probability value) of the output result of the corresponding estimated region integration unit in this photography environment. Exemplarily, as Figure 8 shown, the gate model can output the weights corresponding to each region estimation unit for the image in each photography environment according to various different photography environments (such as daytime on a sunny day, dusk on a rainy day, night on a sunny day, night on a rainy day, daytime on a foggy day, etc.) and the photography environments that each learned region estimation unit (such as region estimation unit 81, region estimation unit 82, region estimation unit 83) is good at processing. For example, for the image in the daytime on a sunny day environment, it outputs the weight w1 corresponding to region estimation unit 81, the weight w2 corresponding to region estimation unit 82, and w3 corresponding to region estimation unit 83. In some embodiments, the estimated region not used corresponding to the region estimation unit can be represented by setting the weight to 0. The estimated regions output by each region estimation unit are weighted according to the weights corresponding to each region estimation unit, and the multiple weighted estimated regions corresponding to the multiple region estimation units are merged. Exemplarily, the estimated region output by region estimation unit 81 is weighted by w1 to obtain the weighted estimated region 1 corresponding to region estimation unit 81; the estimated region output by region estimation unit 82 is weighted by w2 to obtain the weighted estimated region 2 corresponding to region estimation unit 82; the estimated region output by region estimation unit 83 is weighted by w3 to obtain the weighted estimated region 3 corresponding to region estimation unit 83; the weighted estimated region 1, the weighted estimated region 2, and the weighted estimated region 3 are integrally merged, and the object region of the falling object in this photography scene is determined according to the merging result, and then the detection of falling objects in various different photography environments can be processed.

[0095] It should be noted that for the case where the photography scene includes the photography environment, the MoE model can also be used to implement the multiple region estimation units and the estimated region integration unit; the multiple region estimation units are equivalent to the expert models in the MoE model, and the estimated region integration unit is equivalent to the gate model or the gating network in the MoE model.

[0096] It can be understood that, compared with the related technologies currently in use which usually use a large model to implement object detection in a multi-photography environment, the embodiments of the present application use multiple small models, which can improve the overall computing efficiency and thus improve the efficiency of detecting dropped objects. Moreover, since the best model output can be selected according to the input photography environment, there is no need to reset the threshold. In addition, it can also prevent the accuracy degradation caused by the change of the photography environment and achieve more robust detection.

[0097] In some embodiments, the photography scene may include: the distance between the dropped object and the electronic device. The multiple region estimation units may include: multiple region estimation units corresponding to multiple sensor types. Exemplarily, the sensor types may include image sensors such as cameras, fish-eye cameras, lidars, and so on. Correspondingly, the input signals may include: RGB images from cameras, point clouds from LiDARs, images from fish-eye cameras, sensor signals corresponding to multiple sensor types such as 4D ladar. For the sensor signals of each sensor type, the target detection of the dropped object is performed by the corresponding region estimation unit in the multiple region estimation units, and the target detection region corresponding to each region estimation unit is output as the estimated region corresponding to each region estimation unit. It can be understood that when performing target detection on different types of input signals, the formats of the target detection regions output by different region estimation units may be different. For example, the target detection region corresponding to the RGB image of the camera may be a two-dimensional (2D) target detection box and a three-dimensional (3D) target detection box, and the target detection region corresponding to the radar point cloud data may be a 3D target detection box. In the embodiments of the present application, the target detection regions output by different region estimation units can be integrated by the estimated region integration unit to determine the final object region of the dropped object.

[0098] In some embodiments, the estimated region integration unit can automatically determine the importance degree of different types of sensor signals according to the distance between the dropped object and the electronic device, thereby determining the weight corresponding to each region estimation unit, and weighted integrating the multiple target detection regions output by the multiple region estimation units according to the weight corresponding to each region estimation unit to determine the final object region of the dropped object.

[0099] Exemplarily, such as Figure 9As shown in the figure, the multiple region estimation units may include a first network model 91, a second network model 92, and a third network model 93. The estimated region integration unit may include a gate model 94. Among them, the first network model 91 is used to process the RGB image collected by the monocular camera; the second network model 92 is used to process the point cloud data collected by the lidar; the third network model 93 is used to process the fisheye image collected by the fisheye camera. That is to say, each region estimation unit is a neural network model corresponding to the input signal, and the estimated region integration unit is a model capable of processing multi-modal inputs. Through the first network model 91, object detection of falling objects is performed on the RGB image collected by the monocular camera, and a 2D object detection box corresponding to the RGB image is output. Through the second network model 92, object detection of falling objects is performed on the point cloud data collected by the lidar, and a 3D object detection box corresponding to the point cloud data is output. Through the third network model 93, object detection of falling objects is performed on the fisheye image collected by the fisheye camera, and a 2D object detection box corresponding to the fisheye image is output. The multi-modal input of the gate model 94 is the RGB image collected by the monocular camera, the point cloud data collected by the lidar, and the fisheye image collected by the fisheye camera. Through the gate model 94, the distance between the electronic device and the falling object is determined according to the multi-modal input, and one or more object detection boxes are selected from the object detection boxes output by the multiple models for merging according to the distance, and the object region of the final falling object is output.

[0100] It should be noted that, in some embodiments, the first network model 91 may also output both the 2D object detection box and the 3D object detection box corresponding to the RGB image to participate in the decision of the object detection box, and the specific selection is made according to the actual situation, which is not limited in the embodiments of the present application.

[0101] It can be understood that the output of each sensor model has its advantages and disadvantages. The estimated region integration unit can automatically select the appropriate model output according to the distance between the electronic device and the falling object to determine the object region of the final falling object, thereby improving the accuracy of obstacle or falling object detection. Moreover, since each region estimation unit in the multiple region estimation units, such as each sensor model, can be implemented as a lightweight model with high computational efficiency, and the estimated region integration unit, such as the gate model, can also be implemented as a lightweight model. In this way, the multiple region estimation units and the estimated region integration unit can perform synchronous learning, for example, all models are trained at one time by means of supervised training, so as to reduce the consumption of computing resources and improve the network training efficiency.

[0102] The embodiments of the present application also provide an unlearned region estimation device. Figure 10 It is a schematic structural diagram of the unlearned region estimation device provided by the embodiments of the present application. As Figure 10As shown, the unlearned area estimation device 100 includes: a first area estimation unit 110, a second area estimation unit 120, and an estimated area integration unit 130.

[0103] Among them:

[0104] The first area estimation unit 110 is configured to estimate the object area of the dropped object according to the input signal;

[0105] The second area estimation unit 120 is configured to estimate an area different from the characteristics of the first area estimation unit 110 according to a signal that is the same as or different from the input signal;

[0106] The estimated area integration unit 130 is configured to integrate the areas estimated by the first area estimation unit 110 and the second area estimation unit 120 according to the shooting scene.

[0107] In some embodiments, the unlearned area estimation device 100 includes an area estimation unit, and the estimated area integration unit 130 includes an estimated area integration unit.

[0108] The area estimation unit is configured to perform the above-mentioned area estimation with two or more different characteristics;

[0109] The estimated area integration unit is configured to integrate multiple estimated areas.

[0110] In some embodiments, the estimated area integration unit 130 further includes: a shooting scene determination unit.

[0111] The shooting scene determination unit is configured to determine the scene for obtaining the above input signal;

[0112] The estimated area integration unit is configured to integrate the estimated areas according to the output from the shooting scene determination unit.

[0113] In some embodiments, the first area estimation unit 110 is configured to perform image segmentation according to the input signal and estimate the object area of the dropped object;

[0114] The second area estimation unit 120 is configured to perform depth estimation according to a signal that is the same as or different from the input signal and estimate an area different from the characteristics of the first area estimation unit.

[0115] In some embodiments, the second area estimation unit 120 is configured to perform depth estimation according to a signal that is the same as or different from the input signal, determine the depth value corresponding to the pixel; determine the target object according to the depth value corresponding to the pixel, and predict the normal vector information corresponding to the pixel; use the normal vector information to determine the area corresponding to the target object.

[0116] In some embodiments, the unlearned region estimation device 100 includes: a mixture of experts model;

[0117] The mixture of experts model is configured to integrate, by the estimation region integration unit 130, the estimated regions output by a plurality of region estimation units according to the photographing scene; the plurality of region estimation units include the first region estimation unit 110 and the second region estimation unit 120.

[0118] In some embodiments, the photographing scene determination unit is configured to determine the scene of the above input signal according to a scene understanding algorithm.

[0119] In some embodiments, the unlearned region estimation device 100 includes a plurality of region estimation units;

[0120] Each region estimation unit in the plurality of region estimation units is configured to perform the above-mentioned region estimation having two or more different features according to the signal input to itself.

[0121] In some embodiments, the input signal includes: at least one signal data collected by at least one sensor; the at least one sensor includes at least one of an image sensor and a lidar sensor.

[0122] It should be noted that the description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0123] An embodiment of the present application further provides an electronic device, Figure 11 which is an optional structural schematic diagram of the electronic device provided by the embodiment of the present application. As Figure 11 shown, the electronic device 3 includes: a memory 32 and a processor 33. Among them, the memory 32 and the processor 33 are connected through a communication bus 34; the memory 32 is used to store executable instructions; the processor 33 is used to implement the unlearned region estimation method provided by the embodiment of the present application when executing the executable instructions stored in the memory 32.

[0124] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, where the executable instructions, when executed by the above-mentioned processor, will cause the above-mentioned processor to execute the unlearned region estimation method provided by the embodiment of the present application.

[0125] In some embodiments, the computer-readable storage medium may be a ferroelectric random access memory (FRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disc, or a memory such as a CD-ROM; it may also be various devices including one or any combination of the above memories.

[0126] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0127] As an example, the executable instructions may or may not correspond to a file in the file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a hypertext markup language (HTML) document, stored in a single file dedicated to the program in question, or, stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or portions of code). As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one site, or, on multiple computing devices distributed at multiple sites and interconnected by a communication network.

[0128] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, system, or integrated program (computer program product). Therefore, the present application may take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) that contain computer-usable program code.

[0129] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0130] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0131] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0132] As mentioned above, it is only a preferred embodiment of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. An unlearned area estimation device, comprising: A first area estimation unit configured to estimate an object area of a dropped object based on an input signal; A second area estimation unit configured to estimate an area different from the characteristics of the first area estimation unit based on a signal that is the same as or different from the input signal; An estimated area integration unit configured to integrate the areas estimated by the first area estimation unit and the second area estimation unit according to a photographing scene.

2. The apparatus according to claim 1, wherein, The unlearned area estimation device includes an area estimation unit, and the estimated area integration unit includes an estimated area integration unit, The area estimation unit is configured to perform area estimation with two or more different characteristics according to claim 1; The estimated area integration unit is configured to integrate a plurality of estimated areas.

3. The device according to claim 2, wherein, The estimated area integration unit further includes: a photographing scene determination unit, The photographing scene determination unit is configured to determine a scene for obtaining the input signal according to claim 1; The estimated area integration unit is configured to integrate the estimated areas according to the output from the photographing scene determination unit.

4. The device according to any one of claims 1-3, wherein, The first area estimation unit is configured to perform image segmentation based on the input signal and estimate an object area of a dropped object; The second area estimation unit is configured to perform depth estimation based on a signal that is the same as or different from the input signal and estimate an area different from the characteristics of the first area estimation unit.

5. The device according to claim 4, wherein, The second area estimation unit is configured to perform depth estimation based on a signal that is the same as or different from the input signal, determine a depth value corresponding to a pixel; Determine a target object according to the depth value corresponding to the pixel, and predict normal vector information corresponding to the pixel; Use the normal vector information to determine an area corresponding to the target object.

6. The device according to any one of claims 1-5, wherein, The device includes: a mixture of experts model; The mixture of experts model is configured to, through the estimated area integration unit, integrate the estimated areas output by a plurality of area estimation units according to a photographing scene; the plurality of area estimation units include the first area estimation unit and the second area estimation unit.

7. The device according to any one of claims 3-6, wherein, The photographing scene determination unit is configured to determine a scene for obtaining the input signal according to claim 1 according to a scene understanding algorithm.

8. The device according to any one of claims 2-7, wherein, The unlearned area estimation device includes a plurality of area estimation units; Each area estimation unit of the plurality of area estimation units is configured to perform area estimation with two or more different characteristics according to claim 1 based on its own input signal.

9. The device according to any one of claims 1-8, wherein, The input signal includes: at least one signal data collected by at least one sensor; the at least one sensor includes at least one of an image sensor and a lidar sensor.

10. An unlearned area estimation method, comprising: Estimating an object area of a dropped object through a first area estimation unit based on an input signal; Through the second region estimation unit, a region different from the feature of the first region estimation unit is estimated based on a signal that is the same as or different from the input signal; Through the estimated region integration unit, the regions estimated by the first region estimation unit and the second region estimation unit are integrated according to the shooting scene.

11. The method according to claim 10, wherein, The method further includes: Through the region estimation unit, perform region estimation having two or more different features according to claim 1; Through the estimated region integration unit in the estimated region integration unit, by integrating a plurality of estimated regions.

12. The method according to claim 11, wherein, The estimated region integration unit further includes: a shooting scene determination unit; the integrating of the plurality of estimated regions includes: Determine the scene for obtaining the input signal according to claim 1; Through the estimated region integration unit, the estimated regions are integrated according to the output from the shooting scene determination unit.

13. The method according to any one of claims 10-12, wherein The estimating the object region of the dropped object according to the input signal includes: Performing image segmentation according to the input signal to estimate the object region of the dropped object; The estimating a region different from the feature of the first region estimation unit according to a signal that is the same as or different from the input signal includes: Performing depth estimation according to a signal that is the same as or different from the input signal to estimate a region different from the feature of the first region estimation unit.

14. The method according to claim 13, wherein, Performing depth estimation according to a signal that is the same as or different from the input signal to estimate a region different from the feature of the first region estimation unit includes: Performing depth estimation according to a signal that is the same as or different from the input signal to determine the depth value corresponding to the pixel; Determining the target object according to the depth value corresponding to the pixel and predicting the normal vector information corresponding to the pixel; Using the normal vector information to determine the region corresponding to the target object.

15. The method according to any one of claims 12 - 14, wherein, The determining the scene for obtaining the input signal according to claim 1 includes: Determining the scene for obtaining the input signal according to claim 1 according to the scene understanding algorithm.

16. An electronic device, comprising: A memory and a processor; wherein The memory is used to store executable instructions; The processor is used to implement the method according to any one of claims 10 to 15 when executing the executable instructions stored in the memory.

17. A computer-readable storage medium storing executable instructions, which when executed by a processor, implement the method according to any one of claims 10 to 15.

18. An integrated program including computer-executable instructions, which when executed by a processor, implement the method according to any one of claims 10 to 15.