Method and apparatus for determining sensing result, medium, and device
By combining wide-angle and narrow-angle cameras with distance-specific sensing task models, the method addresses high computational complexity and low recall rates in ADAS systems, achieving efficient and accurate object detection from close to long distances.
Patent Information
- Application Number
- JP2025106934
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-26
- Filing Date
- 2025-06-25
- Publication Date
- 2026-01-15
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Advanced driver assistance systems face high computational complexity and low recall rates for long-distance object recognition due to multi-scale feature extraction based on high-resolution images, making it difficult to deploy models on in-vehicle terminals.
Utilize a combination of wide-angle and narrow-angle cameras to cover different distance ranges, with wide-angle cameras for close-range accuracy and narrow-angle cameras for long-range accuracy, employing sensing task models specific to each range to reduce computational complexity and improve recall rates.
This approach enables accurate object detection across all distance ranges with reduced computational demands, facilitating deployment on in-vehicle terminals and enhancing long-distance object recognition rates.
Smart Images

Figure 2026005225000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to computer vision technology, and more particularly to a method, apparatus, medium and device for determining sensing results. [Background technology]
[0002] In advanced driver assistance systems, it is generally necessary to recognize objects at various distances in front of the vehicle. Related technologies generally perform multi-scale feature extraction based on high-resolution images and recognize objects in various distance ranges based on the multi-scale features. However, multi-scale feature extraction and object recognition tend to require a relatively large amount of computation for the sensing task model, which is unfavorable for deploying the model on an in-vehicle terminal, and the recall rate for distant objects is relatively low. Summary of the Invention [Problem to be solved by the invention]
[0003] The embodiments of the present disclosure provide a method, apparatus, medium and device for determining sensing results, which can reduce the computational complexity of sensing task models and improve the recall rate of long-distance objects. [Means for solving the problem]
[0004] A method for determining a sensing result according to a first aspect of the present disclosure includes the steps of: determining a first image collected by a wide-angle camera and a second image collected by a narrow-angle camera, wherein the field of view of the narrow-angle camera is smaller than the field of view of the wide-angle camera; determining a first sensing result corresponding to each of the distance ranges based on the first image, the second image, and a sensing task model corresponding to at least one distance range; and determining a target sensing result based on the first sensing result corresponding to each of the distance ranges.
[0005] An apparatus for determining a sensing result according to a second aspect of the present disclosure includes a first processing module for determining a first image collected by a wide-angle camera and a second image collected by a narrow-angle camera, wherein the field of view of the narrow-angle camera is smaller than the field of view of the wide-angle camera; a second processing module for determining a first sensing result corresponding to each of the distance ranges based on the first image, the second image and a sensing task model corresponding to at least one distance range; and a third processing module for determining a target sensing result based on the first sensing result corresponding to each of the distance ranges.
[0006] A computer-readable storage medium according to a third aspect of the present disclosure stores a computer program for executing the method for determining a sensing result according to any of the above embodiments of the present disclosure.
[0007] An electronic device according to a fourth aspect of the present disclosure includes a processor and a memory for storing instructions executable by the processor, the processor being used to read and execute the executable instructions from the memory to realize a method for determining a sensing result described in any of the above embodiments of the present disclosure.
[0008] A fifth aspect of the present disclosure provides a computer program product, wherein instructions in the computer program product, when executed by a processor, perform a method for determining a sensing result according to any of the above embodiments of the present disclosure. [Effects of the Invention]
[0009] The method, device, medium, and apparatus for determining a sensing result according to the above embodiments of the present disclosure can realize object sensing in each distance range based on a wide-angle image (first image) collected by a wide-angle camera and a narrow-angle image (second image) collected by a narrow-angle camera, by combining them with sensing task models for different distance ranges to obtain first sensing results corresponding to each distance range, and then determine the target sensing result by combining the first sensing results for each distance range. Since the wide-angle camera has high sensing accuracy for close-distance objects, while the narrow-angle camera has high sensing accuracy for long-distance objects, the wide-angle image and the narrow-angle image can cover objects in the entire distance range, from close to long distances. Furthermore, by combining them with sensing task models for different distance ranges, accurate reproduction of objects in the entire distance range can be realized, thereby ensuring the reproduction rate of close-distance objects and significantly improving the reproduction rate of long-distance objects. In addition, by realizing object detection in different distance ranges using sensing task models in different distance ranges, there is no need to rely on multi-scale feature extraction of high-resolution images, which effectively reduces the computational complexity of the sensing task model and makes it easier to deploy the model on the terminal. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is an exemplary application scenario of the method for determining sensing results according to the present disclosure; [Figure 2] 1 is a flowchart of a method for determining a sensing result according to one exemplary embodiment of the present disclosure. [Figure 3] 10 is a flowchart of a method for determining a sensing result according to another exemplary embodiment of the present disclosure. [Figure 4] 10 is a flowchart of a method for determining a sensing result according to yet another exemplary embodiment of the present disclosure. [Figure 5] 10 is a flowchart of a method for determining a sensing result according to yet another example embodiment of the present disclosure. [Figure 6] 10 is a flowchart of a method for determining a sensing result according to yet another exemplary embodiment of the present disclosure. [Figure 7] 10 is a flowchart of determining a sensing result according to one exemplary embodiment of the present disclosure. [Figure 8] FIG. 2 is a schematic diagram of the cropping principle of an input image corresponding to a sensing task model D according to one exemplary embodiment of the present disclosure. [Figure 9] FIG. 2 is a schematic diagram of the cropping principle of an input image corresponding to a sensing task model E according to one exemplary embodiment of the present disclosure. [Figure 10] FIG. 2 is a schematic diagram of the cropping principle of an input image corresponding to a sensing task model F according to one exemplary embodiment of the present disclosure. [Figure 11] 1 is a structural schematic diagram of an apparatus for determining a sensing result according to one exemplary embodiment of the present disclosure; [Figure 12] FIG. 10 is a structural schematic diagram of an apparatus for determining sensing results according to another exemplary embodiment of the present disclosure. [Figure 13] FIG. 10 is a structural schematic diagram of an apparatus for determining sensing results according to yet another exemplary embodiment of the present disclosure. [Figure 14] 1 is a structural diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0011] In order to explain the present disclosure, exemplary embodiments of the present disclosure will be described in detail below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, and are not all of the embodiments, and the present disclosure is not limited to the exemplary embodiments.
[0012] Unless otherwise specifically stated, the relative arrangement of components and steps, formulas and numerical values described in these examples do not limit the scope of the present disclosure.
[0013] Summary of the Disclosure In the process of realizing the present disclosure, the inventors discovered the following: Advanced driver assistance systems generally need to recognize objects at various distances ahead of the vehicle. Related technologies typically perform multi-scale feature extraction based on high-resolution images and recognize objects in each distance range based on the multi-scale features. However, performing multi-scale feature extraction and object recognition based on high-resolution images tends to result in a complex network structure for the sensing task model, which increases the amount of computation required for the sensing task model and makes it difficult to deploy the model on an in-vehicle device. Furthermore, related technologies typically rely on wide-angle images collected by a camera with a relatively large field of view. Wide-angle images have a relatively high recall rate for close-range objects but a relatively low recall rate for long-range objects.
[0014] Illustrative Overview FIG. 1 illustrates an exemplary application scenario of the method for determining a sensing result according to the present disclosure. As shown in FIG. 1, while a vehicle 11 is traveling on a road, a wide-angle camera 12 of the vehicle 11 may collect a wide-angle image (referred to as a first image) of the area in front of the vehicle 11, and a narrow-angle camera 13 may collect a narrow-angle image (referred to as a second image) of the area in front of the vehicle 11. The field of view (FOV) of the wide-angle camera 12 is larger than that of the narrow-angle camera 13, i.e., FOV1 in the figure is larger than FOV2. For example, FOV1 may be 120 degrees, and FOV2 may be 30 degrees. Since the wide-angle camera 12 and the narrow-angle camera 13 have overlapping coverage areas, the narrow-angle camera 13 can be used as a supplement to the wide-angle camera 12 to improve the recall rate of distant objects. Sensing targets may include, but are not limited to, a curb 14, lane markings 15, other vehicles 16, other objects 17, etc., located in front of the vehicle 11. Other objects 17 may include, for example, pedestrians, bicyclists, traffic lights, signboards, traffic cones, road arrows, crosswalks, stop lines, etc. According to the method for determining a sensing result of the present disclosure, when a first image collected by wide-angle camera 12 and a second image collected by narrow-angle camera 13 are determined, a first sensing result corresponding to each distance range can be determined based on the first image, the second image, and a sensing task model corresponding to at least one distance range. Based on the first sensing result corresponding to each distance range, a target sensing result can be determined. Because wide-angle camera 12 has high sensing accuracy for close-distance objects, while narrow-angle camera 13 has high sensing accuracy for long-distance objects, the wide-angle image and narrow-angle image can cover all distance ranges of objects, from close to long. Furthermore, by combining sensing task models for different distance ranges, accurate reproduction of objects within all distance ranges can be achieved. While ensuring the reproduction rate of close-distance objects, the reproduction rate of long-distance objects can be significantly improved. In addition, by realizing object detection in different distance ranges using sensing task models in different distance ranges, there is no need to rely on multi-scale feature extraction of high-resolution images, which effectively reduces the computational complexity of the sensing task model and makes it easier to deploy the model on the terminal.The method for determining the sensing result of the present disclosure is not limited to being used in intelligent driving fields or scenes, but may also be applied to other fields or scenes, such as security monitoring fields.
[0015] Exemplary Methods 2 is a flowchart of a method for determining a sensing result according to an exemplary embodiment of the present disclosure. This embodiment is particularly applicable to electronic devices such as in-vehicle computing platforms. As shown in FIG. 2, the method includes steps 201 to 203.
[0016] In step 201, a first image collected with a wide-angle camera and a second image collected with a narrow-angle camera are determined.
[0017] Here, the field of view of the narrow-angle camera is smaller than that of the wide-angle camera, as shown in FIG. 1. The field of view of the wide-angle camera and the narrow-angle camera have an overlapping area. For example, both the wide-angle camera and the narrow-angle camera are forward-view cameras used to detect objects in front of the vehicle, which may include, for example, curbs, lane markings, other vehicles, pedestrians, cyclists, traffic lights, signs, road markings, etc. The first image collected by the wide-angle camera corresponds to the field of view of the wide-angle camera, i.e., a wide-angle image. The second image collected by the narrow-angle camera corresponds to the field of view of the narrow-angle camera, i.e., a narrow-angle image.
[0018] In step 202, first sensing results corresponding to each distance range are determined based on the first image, the second image, and a sensing task model corresponding to at least one distance range.
[0019] Here, the at least one distance range may include one or more distance ranges. Each distance range may correspond to one or more sensing task models. That is, the number of sensing task models corresponding to each distance range may be one or more. By performing object sensing processes using sensing task models corresponding to different distance ranges, first sensing results corresponding to each distance range may be obtained. The first sensing results may include sensing results corresponding to one or more types of objects. The number of objects of each type may be one or more. The sensing results corresponding to each type of object may include sensing results for each object of that type. The sensing results for each object may include target detection results, semantic segmentation results, etc. The task types included in specific sensing results may be set according to actual needs, and the embodiments of the present disclosure are not limited thereto.
[0020] In some alternative embodiments, the at least one distance range generally includes multiple distance ranges to cover the entire distance range ahead of the vehicle. For example, a sensing task model for at least one distance range (referred to as a first distance range) may be set for a first image with a wide angle. A sensing task model for at least one distance range (referred to as a second distance range) may be set for a second image with a narrow angle.
[0021] In some alternative embodiments, any one distance range may include distance ranges corresponding to one or more types of objects. For example, taking a wide-angle camera as an example, distance range a corresponds to sensing task model A. If sensing task model A is a multi-task sensing model capable of simultaneously sensing multiple types of objects, sensing task model A has different effective sensing distance ranges for different objects, and distance range a includes the distance ranges of the different objects. For example, for vehicles (other vehicles surrounding the host vehicle), sensing task model A can effectively detect other vehicles within a range of 0 to 55 meters. For pedestrians or cyclists, sensing task model A can effectively detect pedestrians or cyclists within a range of 0 to 21 meters. For traffic lights, sensing task model A can effectively detect traffic lights within a range of 0 to 28 meters. That is, distance range a includes the range of 0 to 55 meters for other vehicles, the range of 0 to 21 meters for pedestrians or cyclists, and the range of 0 to 28 meters for traffic lights. For example, sensing task model B can effectively detect other vehicles within a range of 55 meters to 110 meters, pedestrians or cyclists within a range of 21 meters to 43 meters, and traffic lights within a range of 28 meters to 57 meters. Sensing task model C can effectively detect other vehicles within a range of 110 meters to 220 meters, pedestrians or cyclists within a range of 43 meters to 87 meters, and traffic lights within a range of 57 meters to 114 meters. In short, a single sensing task model may have different effective sensing distance ranges for different types of objects, and for a single type of object, multiple sensing task models can be used to cover effective sensing across multiple distance ranges. This allows for the combination of sensing task models with different distance ranges corresponding to wide-angle images and narrow-angle images to achieve sensing reproduction of objects across the entire distance range, ensuring the sensing reproduction rate and accuracy of close-range objects while significantly improving the sensing reproduction rate and accuracy of long-range objects.
[0022] In step 203, a target sensing result is determined based on the first sensing results respectively corresponding to each distance range.
[0023] Here, after obtaining first sensing results corresponding to each distance range, the target detection result can be determined by combining the first sensing results corresponding to each distance range. For example, by combining the vehicle's front-view wide-angle sensing result and the front-view narrow-angle sensing result to obtain the vehicle's front-view sensing result, the target detection result can include the sensing results of all objects within the entire distance range. The entire distance range refers to the entire distance range covered by the wide-angle image and the narrow-angle image. The distance may refer to the vertical distance from the camera (wide-angle camera or narrow-angle camera). In a vehicle's front-view sensing scene, the distance may refer to the vertical distance from the host vehicle.
[0024] The method for determining a detection result according to this embodiment uses wide-angle images collected by a wide-angle camera and narrow-angle images collected by a narrow-angle camera, combined with detection task models for different distance ranges to achieve object detection in each distance range, obtain first detection results corresponding to each distance range, and then combine the first detection results for each distance range to determine the target detection result. Because wide-angle cameras have high detection accuracy for close-range objects, while narrow-angle cameras have high detection accuracy for long-range objects, the wide-angle images and narrow-angle images can cover the entire range of objects from close to long distances. Furthermore, by combining the detection task models for different distance ranges, accurate reproduction of objects in all distance ranges can be achieved, ensuring the reproduction rate of close-range objects while significantly improving the reproduction rate of long-range objects. Furthermore, using detection task models for different distance ranges to detect objects in different distance ranges eliminates the need to rely on multi-scale feature extraction from high-resolution images, thereby effectively reducing the computational complexity of the detection task model and facilitating deployment of the model on a terminal.
[0025] FIG. 3 is a flowchart of a method for determining a sensing result according to another exemplary embodiment of the present disclosure.
[0026] In some alternative embodiments, in addition to the embodiment shown in FIG. 2 above, as shown in FIG. 3, step 202 of determining first sensing results corresponding to each distance range based on a first image, a second image, and a sensing task model corresponding to at least one distance range may specifically include steps 2021 to 2023.
[0027] In step 2021, based on a first image and a sensing task model of at least one first distance range corresponding to the first image, a wide-angle sensing result corresponding to each first distance range is determined.
[0028] Here, each first distance range may include a distance range corresponding to one or more types of objects. Each first distance range and the corresponding sensing task model may be set based on the effective sensing distance range of the wide-angle camera. For the same type of object, multiple first distance ranges correspond to different longitudinal distance ranges of that type of object relative to the host vehicle. For example, the above-mentioned ranges of other vehicles may be 0 to 55 meters, 55 to 110 meters, and 110 to 220 meters. The sensing task models for the three first distance ranges can respectively realize effective sensing reproduction of other vehicles within these three distance ranges, thereby covering the effective reproduction of other vehicles within the range of 0 to 220 meters. The wide-angle effective sensing ranges for different types of objects may be different. For example, for pedestrians or cyclists, the ranges of 0 to 21 meters, 21 to 43 meters, and 43 to 87 meters can realize effective sensing reproduction of pedestrians or cyclists within the range of 0 to 87 meters. This is mainly related to the relationship between the size, distance, and minimum pixels required to detect different types of objects in an image. For example, for the same camera, when objects are at the same distance, the larger the size of the object, the easier it is to detect and reproduce it, and the smaller the size, the harder it is to detect and reproduce it. In practical application, multiple first distance ranges and sensing task models corresponding to each first distance range may be set according to the camera's performance parameters, distance, and the actual physical size of the object. The wide-angle sensing results corresponding to each first distance range include the sensing results of each object sensed from the input image by the sensing task model of that first distance range.
[0029] In some optional embodiments, for a sensing task model corresponding to any one first distance range, the first image may be used as an input image of the sensing task model, i.e., sensing processing is performed on the first image using the sensing task model to obtain a wide-angle sensing result corresponding to the first distance range.
[0030] In some alternative embodiments, the first image may be preprocessed for a sensing task model corresponding to any one of the first distance ranges, and the preprocessed image may be used as an input image for the sensing task model. The preprocessing may include at least one of a scale conversion process, a cropping process, etc. The scale conversion process may be achieved by downsampling (subsampling), i.e., downsampling the first image to obtain an image with a relatively low scale, such as 1 / 2 scale, 1 / 4 scale, or 1 / 8 scale, thereby reducing the size of the input image for the sensing task model and improving the efficiency of the sensing process. The cropping process may be achieved by setting cropping parameters based on the effective distance ranges covered by different sensing task models. For example, if the sensing task model needs to cover a distance range of 0 to 55, a corresponding area may be cropped from the first image based on the area occupied by the distance range of 0 to 55 in the first image, and used as the input image for the sensing task model. This improves the effectiveness of the input image and improves the accuracy and precision of the model sensing results, while further reducing the size of the input image, reducing the computational complexity of the sensing task model, and further improving the efficiency of the sensing process.
[0031] In step 2022, determine narrow-angle sensing results respectively corresponding to each second distance range based on the second image and the sensing task model of at least one second distance range corresponding to the second image.
[0032] Here, each second distance range may include distance ranges corresponding to one or more types of objects. For the same type of object, the distance covered by at least some sub-ranges of the second distance ranges is greater than the distance covered by the first distance range. That is, the entire distance of the second distance range may be greater than the distance within the first distance range, or may have a range that partially overlaps with the first distance range. For example, if the object is another vehicle and the widest first distance range is between 110 meters and 220 meters, at least one second distance range may include a range between 134 meters and 551 meters or a range between 220 meters and 551 meters.
[0033] In some alternative embodiments, the at least one second distance range may include one or more second distance ranges. For example, taking other vehicles as an example, one second distance range may include a range of 134 meters to 300 meters from the other vehicle, and another second distance range may include a range of 300 meters to 600 meters from the other vehicle. The number of second distance ranges may be set according to actual detection requirements, and the embodiments of the present disclosure are not limited thereto.
[0034] In step 2023, a first sensing result corresponding to each distance range is determined based on the wide-angle sensing result corresponding to each first distance range and the narrow-angle sensing result corresponding to each second distance range.
[0035] Here, after obtaining wide-angle sensing results corresponding to each first distance range and narrow-angle sensing results corresponding to each second distance range, the sensing result for each distance range may be set as the first sensing result corresponding to that distance range. For example, the wide-angle sensing result corresponding to each first distance range may be set as the first sensing result corresponding to that first distance range, and the narrow-angle sensing result corresponding to each second distance range may be set as the first sensing result corresponding to that second distance range.
[0036] In this embodiment, for a wide-angle camera, at least one first distance range sensing task model is used to cover the object sensing reproduction in the wide-angle effective sensing range, and for a narrow-angle camera, at least one second distance range sensing task model is used to effectively sense distant objects, supplementing the wide-angle sensing results and improving the sensing reproduction rate of distant objects. Meanwhile, multiple sensing task models are used to obtain object sensing in different distance ranges, and there is no need to rely on multi-scale feature extraction for high-resolution images, which can greatly reduce the complexity of the network structure of the sensing task model, reduce the computational requirements of the model, and make it easier to deploy the model on a terminal.
[0037] FIG. 4 is a flowchart of a method for determining a sensing result according to yet another exemplary embodiment of the present disclosure.
[0038] In some optional embodiments, in addition to the embodiment shown in FIG. 3 above, as shown in FIG. 4, step 2021 of determining wide-angle sensing results corresponding to each first distance range based on a sensing task model of a first image and at least one first distance range corresponding to the first image may include steps 20211 and 20212.
[0039] In step 20211, based on the first image, third images at multiple scales corresponding to the first image are determined.
[0040] Here, the multiple scales may refer to multiple different resolutions. That is, different scales correspond to different resolutions. The third images of the multiple scales may include at least two of the first image at the original scale (which may be expressed as 1 / 1 scale), a 1 / 2 scale image of the first image, a 1 / 4 scale image of the first image, and a 1 / 8 scale image of the first image. The 1 / 2 scale image of the first image means an image whose height and width are both 1 / 2 of the first image, the 1 / 4 scale image of the first image means an image whose height and width are both 1 / 4 of the first image, and the 1 / 8 scale image of the first image means an image whose height and width are both 1 / 8 of the first image. Taking a 1 / 4 scale image as an example, if the resolution of the first image is expressed as H*W, the 1 / 4 scale image of the first image is (H / 4)*(W / 4). Illustratively, the resolution of the first image is 2160*3840, and the resolution of the 1 / 4 scale image of the first image is 540*960.
[0041] In some optional embodiments, the third images at scales other than the image at the original scale among the multiple third images at scales may be obtained by performing a downsampling process or any other form of processing on the first image, and are not particularly limited thereto.
[0042] In step 20212, based on the third images at each scale and the sensing task models respectively corresponding to each first distance range, wide-angle sensing results respectively corresponding to each first distance range are determined.
[0043] Here, sensing task models corresponding to different first distance ranges may use third images of different scales. Images of the same image at different scales may contain different levels of object features, and images of the same scale may have different sensing effects for objects of different sizes. For example, the larger the scale, the greater the resolution, and the clearer the object in the image, or conversely, the blurrier the object. Based on this, the scale corresponding to each sensing task model can be set by combining the computational complexity requirements of the model and the sensing conditions of objects within the distance range that the model needs to cover. For example, for a sensing task model for a short-distance range, objects within the short-distance range generally occupy a relatively large area in the image, and corresponding objects can be effectively reproduced using images of a relatively small scale. Therefore, considering the computational complexity requirements of the sensing task model, for the sensing task model for a short-distance range, images of a relatively small scale may be sampled as input images to reduce the computational complexity of the model.
[0044] In this embodiment, object detection for each first distance range is realized using images of multiple scales corresponding to the first image and sensing task models of different first distance ranges, and the sensing task models of different first distance ranges may use images of different scales as input images, which reduces the scale of the input image of the sensing task model, thereby reducing the calculation amount of the sensing task model and further improving the sensing efficiency of the model.
[0045] In some alternative embodiments, the step 20212 of determining wide-angle sensing results corresponding to each first distance range based on the third images at each scale and sensing task models corresponding to each first distance range may include: The method may include the steps of: determining a target scale and cropping parameters corresponding to the target distance range using any one of the first distance ranges as a target distance range, where different cropping parameters correspond to different distance ranges; determining a first target image from each third image based on the target scale; determining a fourth image based on the cropping parameters and the first target image; and performing sensing processing on the fourth image based on a sensing task model corresponding to the target distance range to obtain a wide-angle sensing result corresponding to the target distance range.
[0046] Here, the target scale (which may be referred to as the first target distance range to distinguish it from the target distance range in the narrow-angle case) corresponding to the target distance range refers to the scale of the third image required for the input image of the sensing task model corresponding to the target distance range. That is, the third image of the target scale among the multiple scales is used to determine the input image of the sensing task model for the target distance range. For example, the third image of the target scale is used as the input image of the sensing task model for the target distance range, or the third image of the target scale is cropped and the cropped area is used as the input image. The cropping parameters corresponding to the target distance range are parameters required to crop the third image of the target scale. The cropping parameters corresponding to the target distance range may include parameters for determining the cropping area, for example, parameters for determining the boundary of the cropping area. Different cropping parameters correspond to different distance ranges. A specific correspondence relationship may be preset and stored according to the size of the area occupied by the actual distance range in the image. The wide-angle sensing result corresponding to the target distance range is called the wide-angle sensing result because it is a sensing result obtained by sensing based on a wide-angle image.
[0047] In some selectable embodiments, a target scale corresponding to each first distance range may be preset for that first distance range, and a correspondence relationship between each first distance range and the target scale may be stored. During use, the target scale corresponding to the target distance range may be determined based on the correspondence relationship. For example, the target scale corresponding to first distance range a may be 1 / 4 scale, the target scale corresponding to first distance range b may be 1 / 2 scale, and the target scale corresponding to first distance range c may be 1 / 1 scale, etc.
[0048] In some alternative embodiments, after determining a target scale and cropping parameters corresponding to the target distance range, a first target image can be determined from each third image based on the target scale, i.e., among the third images at multiple scales, a third image at the same scale as the target scale is set as the first target image. For example, if the target scale is 1 / 2 scale, the third image at 1 / 2 scale is set as the first target image.
[0049] In some alternative embodiments, the first target image may be cropped based on the cropping parameters to obtain a fourth image.
[0050] In some alternative embodiments, the sensing task model corresponding to the target distance range may be a multi-task sensing model, a single-task sensing model, or the like, and the specific sensing task model is not limited. The multi-task sensing model is a model that can simultaneously sense multiple types of objects and / or achieve multiple types of sensing results. The multiple types of objects may include at least two of objects such as curbs, lane markings, other vehicles, bicyclists, pedestrians, traffic lights, and traffic cones. The multiple types of sensing results may include, for example, target detection results and semantic segmentation results.
[0051] In this embodiment, the input image of the sensing task model is obtained by scaling down and cropping the first image, which can significantly reduce the data amount of the input image, thereby effectively reducing the calculation amount of the model, further improving the sensing processing efficiency and more contributing to the deployment of the model on the terminal.
[0052] In some optional embodiments, in addition to the embodiment shown in Figure 3 above, as shown in Figure 4, step 2022 of determining narrow-angle sensing results corresponding to each second distance range based on a sensing task model of a second image and at least one second distance range corresponding to the second image may include steps 20221 and 20222.
[0053] In step 20221, based on the second image, fifth images at multiple scales corresponding to the second image are determined.
[0054] Here, the fifth images at multiple scales may include at least two of the second image at the original scale, a 1 / 2 scale image of the second image, a 1 / 4 scale image of the second image, and a 1 / 8 scale image of the second image, etc. The specific operation of determining the fifth images at multiple scales corresponding to the second image is similar to that of the third image in the above embodiment, and therefore will not be described here.
[0055] In step 20222, narrow-angle sensing results corresponding to each second distance range are determined based on the fifth image at each scale and the sensing task model corresponding to each second distance range.
[0056] Here, the sensing task models corresponding to the different second distance ranges may use fifth images of different scales. Specifically, the scales corresponding to the different second distance ranges may be determined in combination with the computational requirements of the sensing task model and the effective sensing distance range that the sensing task model needs to cover, depending on the sensing situation of the narrow-angle camera for the object in the different second distance ranges. The specific operations for determining the narrow-angle sensing results corresponding to the different second distance ranges are similar to those for the wide-angle sensing results in the above-described embodiment, and therefore will not be described here.
[0057] In this embodiment, object detection for each second distance range is realized by using images of multiple scales corresponding to the second image and sensing task models of different second distance ranges, so that the sensing task models of different second distance ranges can use images of different scales as input images, which effectively reduces the scale of the input image of the sensing task model, thereby reducing the calculation amount of the sensing task model and further improving the model sensing efficiency.
[0058] In some alternative embodiments, step 20222 of determining narrow-angle sensing results corresponding to each second distance range based on a fifth image of each scale and a sensing task model corresponding to each second distance range includes the steps of determining a target scale and cropping parameters corresponding to the target distance range using any of the second distance ranges as a target distance range, where different cropping parameters correspond to different distance ranges; determining a second target image from each fifth image based on the target scale; determining a sixth image based on the cropping parameters and the second target image; and performing sensing processing on the sixth image based on a sensing task model corresponding to the target distance range to obtain narrow-angle sensing results corresponding to the target distance range.
[0059] Here, the target scale (which may be referred to as the second target scale) corresponding to the target distance range (which may be referred to as the second target distance range) refers to the scale of the fifth image required for the input image of the sensing task model corresponding to the target distance range. That is, the fifth image of the target scale among the multiple scales is used to determine the input image of the sensing task model for the target distance range. For example, the fifth image of the target scale is used as the input image of the sensing task model for the target distance range, or the fifth image of the target scale is cropped and the cropped area is used as the input image. The cropping parameters corresponding to the target distance range are parameters required to crop the fifth image of the target scale. The cropping parameters corresponding to the target distance range may include parameters for determining the cropping area, for example, parameters for determining the boundaries of the cropping area. Different cropping parameters correspond to different second distance ranges. A specific correspondence may be preset and stored according to the size of the area occupied by the actual second distance range in the image. The narrow-angle sensing result corresponding to the target distance range is called a narrow-angle sensing result because it is a sensing result obtained by sensing based on a narrow-angle image.
[0060] In some selectable embodiments, a target scale corresponding to each second distance range may be preset for that second distance range, and a correspondence relationship between each second distance range and the target scale may be stored. During use, the target scale corresponding to the target distance range may be determined based on the correspondence relationship. For example, the target scale corresponding to second distance range g may be 1 / 4 scale, the target scale corresponding to second distance range h may be 1 / 2 scale, and the target scale corresponding to second distance range t may be 1 / 1 scale, etc.
[0061] In some alternative embodiments, after determining a target scale and cropping parameters corresponding to the target distance range, a second target image can be determined from each fifth image based on the target scale, i.e., among the fifth images at multiple scales, a fifth image at the same scale as the target scale is set as the second target image. For example, if the target scale is 1 / 2 scale, the fifth image at 1 / 2 scale is set as the second target image.
[0062] In some alternative embodiments, the second target image may be cropped based on the cropping parameters to obtain a sixth image.
[0063] In some alternative embodiments, the sensing task model corresponding to the target distance range may be a multi-task sensing model, a single-task sensing model, etc., and the specific sensing task model is not limited.
[0064] In this embodiment, the input image of the sensing task model is obtained by scaling down and cropping the second image, which can significantly reduce the data amount of the input image, thereby effectively reducing the calculation amount of the model, further improving the sensing processing efficiency and more contributing to the deployment of the model on the terminal.
[0065] FIG. 5 is a flowchart of a method for determining a sensing result according to yet another example embodiment of the present disclosure.
[0066] In one alternative embodiment, as shown in FIG. 5, the method of the embodiment of the present disclosure may further include steps 301 to 303.
[0067] In step 301, a target fifth image of a first preset scale is determined from the fifth images of each scale.
[0068] Here, the first preset scale may be preset according to the sensing requirements of the preset object type. The preset object type may be any object type. For example, the preset object type may be a traffic light, a lane marking, etc. Different preset object types may have the same or different first preset scales. For example, the preset object type may be a traffic light, and the first preset scale may be a half scale. Alternatively, the preset object type may be a lane marking, and the first preset scale may be a quarter scale. This is merely an example and is not intended to limit the embodiments of the present disclosure. Among the fifth images at each scale, the fifth image at the first preset scale is set as the target fifth image.
[0069] In step 302, the target fifth image is cropped based on the preset cropping parameters of the preset object type to obtain a third target image.
[0070] Here, the preset object type may be set according to actual detection requirements. For example, the preset object type may be a traffic light, a lane marking, etc. The preset cropping parameters of the preset object type may be set according to the detection distance range requirements of the preset object type. For example, if the detection purpose of the preset object type is to improve the recall rate of long-distance objects, the preset cropping parameters of the preset object type may be configured to crop an area corresponding to the long-distance range in the fifth target image. If the detection purpose of the preset object type is to address the problem of missing short-distance lane marking detection in the second distance range detection task model, the preset cropping parameters may be configured to crop an area corresponding to the short-distance range in the fifth target image. The corresponding area is cropped from the fifth target image as a third target image based on the preset cropping parameters of the preset object type.
[0071] In step 303, a sensing process is performed on the third target image according to the sensing task model corresponding to the preset object type, and a narrow-angle sensing result corresponding to the preset object type is obtained.
[0072] Here, the sensing task model corresponding to the preset object type is a model dedicated to detecting objects of the preset object type. For example, the sensing task model for the preset object type may be a narrow-angle lane marking detection model for detecting close-range lane markings, a narrow-angle traffic light detection model for detecting distant traffic lights, etc., but is not limited to these. A third target image is inferred based on the sensing task model, and a narrow-angle sensing result corresponding to the preset object type is obtained based on the inference result. The narrow-angle sensing result corresponding to the preset object type may include at least one of sensing results, such as target detection results and semantic segmentation results, for each detected object of the preset object type.
[0073] In some alternative embodiments, the number of preset object types may be one or more. For example, two preset object types, a traffic light and a lane marking, may be set, and each preset object type has a corresponding sensing task model for realizing the sensing task of the corresponding preset object type.
[0074] The step 203 of determining a target sensing result based on the first sensing results corresponding to each distance range includes: The method may include a step 2031 of determining a target sensing result based on the first sensing result corresponding to each distance range and the narrow-angle sensing result corresponding to the preset object type.
[0075] Here, the first sensing results corresponding to each distance range and the narrow-angle sensing results corresponding to the preset object types can be combined to obtain the target sensing results.
[0076] In some alternative embodiments, the first sensing results for each distance range and the narrow-angle sensing results corresponding to a preset object type can be fused to obtain a target sensing result. The fusion process may include coordinate system transformation, deduplication of sensing results for repeatedly sensed objects, aggregation of non-duplicate objects, and splicing fusion of objects across multiple distance ranges (i.e., a single object (e.g., lane markings, road edges, etc.) may exist in multiple distance ranges due to its large size). Here, the purpose of coordinate system transformation is to integrate the wide-angle sensing results and the narrow-angle sensing results into the same coordinate system, thereby determining whether the object matches and whether the object is repeatedly sensed. For a repeatedly sensed object, if the entire object is in the narrow-angle sensing result, the narrow-angle sensing result is used as the target sensing result for that object, and the wide-angle sensing result for that object is deleted.
[0077] In some alternative embodiments, the sensing results of sensing task models with different distance ranges are already clearly divided by distance ranges, so overlap elimination may not be necessary, and overlap elimination processing is performed on the sensing results of sensing task models with overlapping distance ranges.
[0078] In some alternative embodiments, if the sensing objects of different sensing task models are different, deduplication is not required. For example, if one sensing task model senses other vehicles, cyclists, and pedestrians, and another sensing task model senses curbs and lane markings, deduplication is not required for the sensing results of the two sensing task models. The specific fusion method may be set according to actual situations, and the embodiments of the present disclosure are not limited thereto.
[0079] This embodiment can detect objects of a specific object type, and can set specific cropping parameters for the specific object type to supplement the narrow-angle detection results in the second distance range, compensate for the deficiencies of the narrow-angle detection results, and improve the validity and reliability of the detection results. For example, a narrow-angle traffic light detection task model is used to perform detection processing on a narrow-angle cropped image in the long distance range, thereby improving the recall rate of long-distance traffic lights. A narrow-angle lane marking detection task model is used to perform detection processing on a narrow-angle cropped image in the short distance range, thereby compensating for the deficiency that the narrow-angle detection results in the second distance range tend to miss detecting close-distance lane markings, and improving the recall rate of narrow-angle detection of close-distance lane markings.
[0080] In some alternative embodiments, step 302 of cropping the target fifth image based on preset cropping parameters of a preset object type and obtaining a third target image may include determining vanishing point pixel coordinates in the target fifth image, and cropping the target fifth image based on the preset cropping parameters, centered on the vanishing point pixel coordinates, and obtaining a third target image.
[0081] Here, the vanishing point pixel coordinates in the fifth target image may be determined based on the calibration parameters of the narrow-angle camera. For example, the calibration parameters of the narrow-angle camera include the vanishing point pixel coordinates in the second image of the narrow-angle camera, and the vanishing point pixel coordinates in the fifth target image are determined based on the pixel mapping relationship between the fifth target image and the second image. Alternatively, the vanishing point pixel coordinates in the fifth image at each scale may be calculated and stored in advance based on the pixel mapping relationship between the second image and the fifth image at each scale and the calibrated vanishing point pixel coordinates, and the vanishing point pixel coordinates in the fifth target image can be directly obtained from the storage area when used. The preset cropping parameters may include cropping amounts in the negative x direction, negative y direction, positive x direction, and positive y direction based on the vanishing point pixel coordinates. For example, the preset cropping parameters may be expressed as [s1, s2, s3, s4], and the cropping area may be expressed as [xo-s1, yo-s2, xo+s3, yo+s4]. Alternatively, the preset cropping parameters may be marked to indicate a direction. For example, s1=-450, and in this case the cropping area may be expressed as [xo+s1, yo+s2, xo+s3, yo+s4]. The specific display form of the preset cropping parameters is not limited.
[0082] In this embodiment, the image is cropped based on the vanishing point in the image and preset cropping parameters to obtain the corresponding input image for the sensing task model. Because the vanishing point represents the visual intersection point of parallel lines, the distance range that the cropping area can cover can be effectively controlled based on the vanishing point and the cropping parameters, so that the cropped image can better meet the sensing distance range requirements of the sensing task model, improving the effectiveness of the input image for the sensing task model, and thereby improving the accuracy of the sensing results of the sensing task model and the recall rate for objects in the corresponding distance range.
[0083] FIG. 6 is a flowchart of a method for determining a sensing result according to yet another exemplary embodiment of the present disclosure.
[0084] In some alternative embodiments, in the embodiment shown in FIG. 2, the step 203 of determining the target sensing result based on the first sensing result corresponding to each distance range can be: Step 203a: fusing the first sensing results respectively corresponding to each distance range to obtain a fusion result; and determining 203b a target sensing result based on the fusion result.
[0085] Here, the specific operation principle of step 203a and step 203b may refer to the aforementioned step 2031, except that in this embodiment, the narrow-angle sensing result of the preset object type may not be included.
[0086] In some alternative embodiments, FIG. 7 is a flowchart for determining sensing results according to one exemplary embodiment of the present disclosure. As shown in FIG. 7, the wide-angle camera is a front-view wide-angle camera 41, where *1 indicates that there is one wide-angle camera, and the narrow-angle camera is a front-view narrow-angle camera 42, which also has one narrow-angle camera. The front-view wide-angle camera 41 collects and acquires a first image. The first image is subsampled to acquire third images at multiple scales, including a half-scale image and a quarter-scale image. The third images at multiple scales may further include the first image at the original scale. Sensing task model A, sensing task model B, and sensing task model C are sensing task models corresponding to distance range a, distance range b, and distance range c, respectively. Sensing task model A, sensing task model B, and sensing task model C may all be multi-task models, i.e., capable of simultaneously sensing at least one sensing result for each of multiple different objects. The different objects may be, for example, the rear of a vehicle, a vehicle, a pedestrian, a cyclist, a sign, a traffic light, or a road marking. The at least one sensing result may include, for example, a target detection result or a semantic segmentation result. A corresponding region (i.e., a fourth image) is cropped from the 1 / 4-scale image based on the cropping parameters 43, and sensing processing is performed on the cropped region using sensing task model A to obtain a wide-angle sensing result corresponding to distance range a. A corresponding region (i.e., a fourth image) is cropped from the 1 / 2-scale image based on the cropping parameters 44, and sensing processing is performed on the cropped region using sensing task model B to obtain a wide-angle sensing result corresponding to distance range b. A corresponding region (i.e., a fourth image) is cropped from the first image at the original scale based on the cropping parameters 45, and sensing processing is performed on the cropped region using sensing task model C to obtain a wide-angle sensing result corresponding to distance range c. Wide-angle sensing fusion is performed on the wide-angle sensing results for each distance range to obtain a wide-angle fusion result.Alternatively, for the sensing task model A, the 1 / 4 scale image may not be cropped, and the sensing task model A may directly perform sensing processing on the 1 / 4 scale image to obtain a wide-angle sensing result corresponding to the distance range a.
[0087] The forward-view narrow-angle camera 42 collects and acquires a second image, performs subsampling on the second image, and acquires fifth images at multiple scales, including a half-scale image and a quarter-scale image. A corresponding region (i.e., a sixth image) is cropped from the half-scale image based on cropping parameters 46, and sensing processing is performed on the cropped region using sensing task model D to acquire a narrow-angle sensing result corresponding to distance range d. Sensing task model D is a sensing task model corresponding to distance range d. Sensing task model D may be a multi-task model. Preset object types include traffic lights and lane markings. Traffic lights correspond to sensing task model E, and lane markings correspond to sensing task model F. That is, sensing task model E is a traffic light single-task sensing model, and sensing task model F is a lane marking single-task sensing model. A corresponding region is cropped from the 1 / 4-scale image based on the cropping parameters 47, and traffic light detection processing is performed on the cropped region using the sensing task model E to obtain a narrow-angle traffic light detection result. A corresponding region is cropped from the 1 / 4-scale image based on the cropping parameters 48, and lane marking detection processing is performed on the cropped region using the sensing task model F to obtain a narrow-angle lane marking detection result. The wide-angle fusion result, narrow-angle detection result, narrow-angle traffic light detection result, and narrow-angle lane marking detection result are fused to obtain a final target detection result. As shown in the figure, the target detection result includes a target detection result and a semantic segmentation result for each object. The target detection result may include a two-dimensional detection frame for the object (e.g., a 2D box for each object in the figure). The semantic segmentation result may include a probability that each pixel in the image belongs to the object, a pixel point set in the image that belongs to the object, or a classification label for each pixel in the image that belongs to each object. The specific content of the sensing result is not limited.The segmentation LabelMap is a semantic segmentation result represented as a pixel classification label map, where each pixel may correspond to one classification label representing the object type to which the pixel belongs, i.e., the LabelMap includes semantic segmentation results corresponding to each detected object. The forward view lane marking LabelMap represents the semantic segmentation result of lane markings.
[0088] Here, the cropping parameters corresponding to different sensing task models may be different. That is, parameters 43 to 48 may be different parameters, so that the cropping region can cover different distance ranges and better meet the sensing requirements of the sensing task model, such as satisfying the sensing characteristics of traffic lights and lane markings, thereby improving the accuracy and precision of the sensing results of the sensing task model. FIG. 7 is merely an exemplary embodiment of the present disclosure, and actual applications are not limited to the embodiment of FIG. 7. For example, in FIG. 7, only one distance range (i.e., distance range d) of sensing task model D is shown for narrow-angle sensing. In actual applications, multiple distance ranges may be set for narrow-angle sensing. Also, for example, the scale of the image on which each sensing task model is based is not limited to the scale shown in FIG. 7.
[0089] In some alternative embodiments, the narrow-angle sensing task model D may be a further enhancement of the forward-view wide-angle sensing model and may include 2D detection multitasks and semantic segmentation tasks for five types of objects: full vehicle, rear vehicle, pedestrian, bicyclist, and traffic cone, as well as a lane marking semantic segmentation task.
[0090] In some alternative embodiments, the distance ranges corresponding to the sensing task models for each distance range are shown in Table 1 below. [Table 1] Here, the actual physical size refers to the actual size of the object. Based on the actual physical size of the object and the minimum number of pixels required for sensing, the effective sensing distance range for each sensing task model for different objects can be estimated. Because the actual sizes of different objects may differ and the minimum number of pixels required for sensing different objects may differ, the effective sensing distance range for different objects of the same sensing task model may differ. Table 1 is merely a description of exemplary ranges, and in actual application, the distance range for different objects of the sensing task model is not limited to the specific ranges in Table 1.
[0091] In some alternative embodiments, to reduce model training costs, the sensing task model D, the sensing task model E, and the sensing task model F may use the same network structure. For example, a vargnet may be used as the backbone network, and a U-shaped network (unet) may be used as the neck network. Of course, this is merely a possible embodiment and does not limit the methods of the embodiments of the present disclosure. In actual application, each sensing task model may use a different network structure. The specific network structure is not limited to the vargnet and unet described above.
[0092] In some alternative embodiments, based on the embodiment shown in FIG. 7, FIG. 8 is a schematic diagram of the cropping principle of an input image corresponding to a sensing task model D according to one exemplary embodiment of the present disclosure. As shown in FIG. 8, the fifth image (i.e., image pyramid) of multiple scales corresponding to the second image includes an image at the original scale (1 / 1 scale) (in the figure, H*W=2160*3840 is taken as an example, where H represents the image height and W represents the image width), an image at 1 / 2 scale (i.e., the height and width are 1 / 2 of the original scale), an image at 1 / 4 scale (i.e., the height and width are 1 / 4 of the original scale), and an image at 1 / 8 scale (i.e., the height and width are 1 / 8 of the original scale). The sensing task model D is a narrow-angle multi-task sensing model. The target scale corresponding to the distance range d is 1 / 2 scale, xoy represents the pixel coordinate system of the image, and the pixel coordinates of the vanishing point P are expressed as (xo, yo). The half-scale image (i.e., the second target image) is cropped based on the vanishing point pixel coordinates (xo, yo) and the cropping parameters [s1, s2, s3, s4] = [-704, -284, 540, 228] corresponding to the distance range d, to obtain a sixth image of 512*1344, expressed as [xo-704, yo-284, xo+540, yo+228]. Sensing processing is performed on the sixth image based on the sensing task model D, and a narrow-angle sensing result corresponding to the distance range d is obtained. As a further supplement to the wide-angle model, the narrow-angle multi-task sensing task model D has a longer sensing distance, which can effectively improve the recall rate of distant objects.
[0093] In some alternative embodiments, based on the embodiment shown in FIG. 7 , FIG. 9 is a schematic diagram of the cropping principle of an input image corresponding to a sensing task model E according to one exemplary embodiment of the present disclosure. As shown in FIG. 9 , the sensing task model E is a narrow-angle traffic light single-task sensing model. According to the spatial domain sensing range of the traffic light, in combination with the computational complexity requirements and sensing distance requirements of the sensing task model, the first preset scale corresponding to the traffic light is set to 1 / 2 scale, and the corresponding preset cropping parameters are set to [s1, s2, s3, s4] = [704, 540, 540, 164]. The specific cropping principle is similar to that shown in FIG. 8 , and therefore will not be described here. The narrow-angle traffic light single-task sensing can significantly improve the detection recall rate of distant traffic lights.
[0094] In some alternative embodiments, based on the embodiment shown in FIG. 7 , FIG. 10 is a schematic diagram of the cropping principle of an input image corresponding to sensing task model F according to one exemplary embodiment of the present disclosure. As shown in FIG. 10 , sensing task model F is a lane marking single-task sensing model. The specific cropping principle is similar to that shown in FIG. 9 above, except that the first preset scale corresponding to the lane markings is 1 / 4 scale and the preset cropping parameters are [s1, s2, s3, s4]=[458, 4, 438, 188], and a detailed description thereof will be omitted. The narrow-angle lane marking single-task sensing model (sensing task model F) supplements the narrow-angle multi-task sensing model (sensing task model D), thereby compensating for the shortcoming of the narrow-angle multi-task sensing model, which tends to miss nearby lane markings, and improving the completeness and accuracy of lane marking detection results.
[0095] 8 to 10 are merely illustrative examples, and in actual applications, the trimming parameters may be set according to the actual situation and are not limited to the specific parameter values in the figures.
[0096] The above-mentioned embodiments of the present disclosure may be implemented alone or in any combination if there is no conflict, and may be specifically set according to actual needs, and are not limited by the present disclosure.
[0097] Any of the methods for determining sensing results according to the embodiments of the present disclosure may be executed by any suitable device having data processing capabilities, including, but not limited to, a terminal device, a server, etc. Alternatively, any of the methods for determining sensing results according to the embodiments of the present disclosure may be executed by a processor, for example, the processor executes any of the methods for determining sensing results according to the embodiments of the present disclosure by calling corresponding instructions stored in a memory, the description of which will be omitted below.
[0098] Exemplary Apparatus 11 is a structural schematic diagram of an apparatus for determining a sensing result according to an exemplary embodiment of the present disclosure. The apparatus of this embodiment can be used to realize an embodiment of a method corresponding to the present disclosure, and the apparatus shown in FIG. 11 may include a first processing module 51, a second processing module 52, and a third processing module 53.
[0099] The first processing module 51 is used to determine a first image collected by a wide-angle camera and a second image collected by a narrow-angle camera, the field of view of the narrow-angle camera being smaller than the field of view of the wide-angle camera.
[0100] The second processing module 52 is used to determine first sensing results corresponding to each of the distance ranges based on the first image, the second image and a sensing task model corresponding to at least one distance range.
[0101] The third processing module 53 is used for determining a target sensing result based on the first sensing result corresponding to each of the distance ranges respectively.
[0102] FIG. 12 is a structural schematic diagram of an apparatus for determining a sensing result according to another exemplary embodiment of the present disclosure.
[0103] In some alternative embodiments, in addition to the embodiment shown in FIG. 11, the second processing module 52 may include a first processing unit 521, a second processing unit 522, and a third processing unit 523, as shown in FIG. 12.
[0104] The first processing unit 521 is used to determine wide-angle sensing results corresponding to each first distance range based on a sensing task model of a first image and at least one first distance range corresponding to the first image.
[0105] The second processing unit 522 is used to determine narrow-angle sensing results corresponding to each second distance range based on the second image and a sensing task model of at least one second distance range corresponding to the second image.
[0106] The third processing unit 523 is used to determine a first sensing result corresponding to each distance range based on the wide-angle sensing result corresponding to each first distance range and the narrow-angle sensing result corresponding to each second distance range.
[0107] In some alternative embodiments, the first processing unit 521 is specifically used to determine, based on the first image, third images of multiple scales corresponding to the first image, and to determine wide-angle sensing results corresponding to each first distance range based on the third images of each scale and sensing task models corresponding to each first distance range.
[0108] In some alternative embodiments, the first processing unit 521 specifically takes any first distance range as a target distance range, determines a target scale and cropping parameters corresponding to the target distance range, where different cropping parameters correspond to different distance ranges, determines a first target image from each third image based on the target scale, determines a fourth image based on the cropping parameters and the first target image, and performs sensing processing on the fourth image based on a sensing task model corresponding to the target distance range, to obtain a wide-angle sensing result corresponding to the target distance range.
[0109] In some alternative embodiments, the second processing unit 522 is specifically used to determine, based on the second image, fifth images of multiple scales corresponding to the second image, and to determine narrow-angle sensing results corresponding to each second distance range based on the fifth images of each scale and sensing task models corresponding to each second distance range.
[0110] In some alternative embodiments, the second processing unit 522 specifically determines a target scale and cropping parameters corresponding to the target distance range, with any second distance range as the target distance range, where different cropping parameters correspond to different distance ranges, determines a second target image from each fifth image based on the target scale, determines a sixth image based on the cropping parameters and the second target image, and performs sensing processing on the sixth image based on a sensing task model corresponding to the target distance range, to obtain a narrow-angle sensing result corresponding to the target distance range.
[0111] In some alternative embodiments, as shown in FIG. 12, the apparatus of the present disclosure may further include: a fourth processing module 54 for determining a target fifth image at the first preset scale from the fifth image at each scale; a fifth processing module 55 for cropping the target fifth image based on preset cropping parameters of the preset object type to obtain a third target image; and a sixth processing module 56 for performing sensing processing on the third target image based on a sensing task model corresponding to the preset object type, and obtaining a narrow-angle sensing result corresponding to the preset object type.
[0112] The third processing module 53 may include a fourth processing unit 531 for determining a target detection result based on the first detection result corresponding to each distance range and the narrow-angle detection result corresponding to the preset object type.
[0113] In some optional embodiments, the fifth processing module 55 is specifically used to determine the vanishing point pixel coordinates in the target fifth image, crop the target fifth image based on the preset cropping parameters around the vanishing point pixel coordinates, and obtain the third target image.
[0114] FIG. 13 is a structural schematic diagram of an apparatus for determining a sensing result according to yet another exemplary embodiment of the present disclosure.
[0115] In some alternative embodiments, on the embodiment shown in FIG. 11, the third processing module 53 a fusion processing unit 53a for fusing the first sensing results respectively corresponding to each distance range to obtain a fusion result; and a determining unit 53b for determining a target sensing result based on the fusion result.
[0116] The beneficial technical effects corresponding to the exemplary embodiments of the present apparatus may refer to the beneficial technical effects corresponding to the exemplary method part above, and will not be described here.
[0117] Exemplary Electronic Devices FIG. 14 is a structural diagram of an electronic device according to an embodiment of the present disclosure, where an electronic device 90 includes at least one processor 91 and a memory 92.
[0118] The processor 91 may be a central processing unit (CPU) or other form of processing unit having data processing and / or instruction execution capabilities, and may control other components in the electronic device 90 to perform desired functions.
[0119] The memory 92 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 91 may execute the one or more computer program instructions to implement the methods and / or other desired functions of each embodiment of the present disclosure described above.
[0120] In one example, electronic device 90 may further include input devices 93 and output devices 94, with these components interconnected via a bus system and / or other form of connection (not shown).
[0121] The input device 93 may include a keyboard, a mouse, and the like.
[0122] The output device 94 is capable of outputting various types of information to the outside, and may include a display, a speaker, a printer, a communication network, and remote output devices connected thereto.
[0123] 14 shows only some of the components related to the present disclosure in the electronic device 90, and omits components such as a bus and an input / output interface, for the sake of simplicity of explanation. Furthermore, the electronic device 90 may further include any other appropriate components depending on the specific application situation.
[0124] Exemplary Computer Program Products and Computer-Readable Storage Media In addition to the methods and apparatus described above, embodiments of the present disclosure may further provide a computer program product including computer program instructions that, when executed by a processor, cause the processor to perform the steps of the methods of various embodiments of the present disclosure described in the "Exemplary Methods" section above.
[0125] The computer program product may have program code for carrying out operations of embodiments of the present disclosure written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and traditional procedural programming languages such as "C" or similar programming languages. The program code may execute entirely on a user's computing device, partially on a user's device, as separate software packages, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0126] An embodiment of the present disclosure may also be a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, cause the processor to perform the steps of the methods of various embodiments of the present disclosure described in the "Exemplary Methods" section above.
[0127] A computer-readable storage medium may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but not limited to, an electric, magnetic, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0128] Although the basic principles of the present disclosure have been described above with reference to specific embodiments, the benefits, advantages, effects, etc. mentioned in the present disclosure are not limited but merely illustrative, and these benefits, advantages, effects, etc. do not necessarily exist in each embodiment of the present disclosure. Furthermore, the specific details disclosed above are not limited but merely serve to serve as examples and to facilitate understanding, and the above details do not necessarily limit the present disclosure to be realized by the above specific details.
[0129] Those skilled in the art can make various modifications and variations to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure also intends to include these modifications and variations.
Claims
1. determining a first image collected with a wide-angle camera and a second image collected with a narrow-angle camera, the field of view of the narrow-angle camera being smaller than the field of view of the wide-angle camera; determining a first sensing result corresponding to each of the distance ranges based on the first image, the second image, and a sensing task model corresponding to at least one distance range; determining a target sensing result based on the first sensing results corresponding to each of the distance ranges.
2. determining a first sensing result corresponding to each of the distance ranges based on the first image, the second image, and a sensing task model corresponding to at least one distance range, determining wide-angle sensing results corresponding to each of the first distance ranges based on the sensing task model of the first image and at least one first distance range corresponding to the first image; determining narrow-angle sensing results corresponding to each of the second distance ranges based on the sensing task model of the second image and at least one second distance range corresponding to the second image; and determining the first sensing results corresponding to each of the distance ranges based on the wide-angle sensing results corresponding to each of the first distance ranges and the narrow-angle sensing results corresponding to each of the second distance ranges.
3. determining a wide-angle sensing result corresponding to each of the first distance ranges based on the sensing task model of the first image and at least one first distance range corresponding to the first image, determining third images at multiple scales corresponding to the first image based on the first image; and determining the wide-angle sensing results corresponding to each of the first distance ranges based on the third image at each of the scales and the sensing task model corresponding to each of the first distance ranges.
4. determining the wide-angle sensing results corresponding to each of the first distance ranges based on the third images at each of the scales and the sensing task models corresponding to each of the first distance ranges, determining a target scale and a cropping parameter corresponding to the target distance range by setting any one of the first distance ranges as a target distance range, wherein different cropping parameters correspond to different distance ranges; determining a first target image from each of the third images based on the target scale; determining a fourth image based on the cropping parameters and the first target image; performing sensing processing on the fourth image based on the sensing task model corresponding to the target distance range, and obtaining the wide-angle sensing result corresponding to the target distance range.
5. determining narrow-angle sensing results corresponding to each of the second distance ranges based on the sensing task model of the second image and at least one second distance range corresponding to the second image, determining, based on the second image, fifth images at a plurality of scales corresponding to the second image; and determining the narrow-angle sensing results corresponding to each of the second distance ranges based on the fifth image at each of the scales and the sensing task model corresponding to each of the second distance ranges.
6. determining the narrow-angle sensing results corresponding to each of the second distance ranges based on the fifth images at each of the scales and the sensing task models corresponding to each of the second distance ranges, determining a target scale and a target trimming parameter corresponding to the target distance range by setting any one of the second distance ranges as a target distance range, wherein different trimming parameters correspond to different distance ranges; determining a second target image from each of the fifth images based on the target scale; determining a sixth image based on the cropping parameters and the second target image; and performing sensing processing on the sixth image based on the sensing task model corresponding to the target distance range to obtain the narrow-angle sensing result corresponding to the target distance range.
7. determining a target fifth image at a first preset scale from the fifth image at each of the scales; cropping the target fifth image based on preset cropping parameters of a preset object type to obtain a third target image; performing a sensing process on the third target image based on a sensing task model corresponding to the preset object type, and obtaining a narrow-angle sensing result corresponding to the preset object type; The step of determining a target sensing result based on the first sensing results corresponding to each of the distance ranges includes: The method for determining a sensing result according to claim 5, further comprising: determining the target sensing result based on the first sensing result corresponding to each of the distance ranges and the narrow-angle sensing result corresponding to the preset object type.
8. Cropping the target fifth image based on preset cropping parameters of a preset object type to obtain a third target image includes: determining vanishing point pixel coordinates in the target fifth image; The method for determining the sensing result of claim 7, further comprising: cropping the target fifth image based on the preset cropping parameters around the vanishing point pixel coordinates to obtain the third target image.
9. The step of determining a target sensing result based on the first sensing results corresponding to each of the distance ranges includes: fusing the first sensing results corresponding to each of the distance ranges to obtain a fusion result; and determining the target sensing result based on the fusion result.
10. a first processing module for determining a first image collected by a wide-angle camera and a second image collected by a narrow-angle camera, the field of view of the narrow-angle camera being smaller than the field of view of the wide-angle camera; a second processing module for determining a first sensing result corresponding to each of the distance ranges based on the first image, the second image, and a sensing task model corresponding to at least one distance range; a third processing module for determining a target sensing result based on the first sensing results corresponding to each of the distance ranges.
11. A computer-readable storage medium storing a computer program for executing the method for determining a sensing result according to any one of claims 1 to 8.
12. a processor; a memory for storing instructions executable by the processor; The processor is adapted to read and execute the executable instructions from the memory to implement the method for determining a sensing result according to any one of claims 1 to 8.
Citation Information
Patent Citations
Rider support system and method
JP2021192303A