Obstacle detection method, electronic equipment and storage medium

By using multi-scale multi-view training and image size adaptability internal parameter matrix adjustment in the obstacle detection model, the poor detection effect caused by the image size changes in the vehicle environment is solved, and efficient obstacle detection for images of any size is achieved.

CN120356175APending Publication Date: 2025-07-22BEIJING MAICHI ZHIXING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311705243.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, the image size of the vehicle environment is not robust, resulting in poor obstacle detection effect.

Method used

By acquiring environmental images of different view angles of the target vehicle, an obstacle detection model trained based on multi-scale multi-view sample environmental image is used, and the internal reference matrix is adjusted according to the image size adaptability to perform obstacle detection.

Benefits of technology

The robustness of the obstacle detection model for vehicle environmental images of different sizes is improved, effective obstacle detection for images of any size is realized, and the detection effect is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356175A_ABST
    Figure CN120356175A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an obstacle detection method, electronic equipment and a storage medium. The method comprises the following steps: acquiring environment images of a target vehicle at different visual angles; wherein the environment images are collected by first vehicle-mounted cameras deployed in different directions of a vehicle body of the target vehicle; determining a first internal reference matrix for obstacle detection according to the first image size of the environment image and an original internal reference matrix of the first vehicle-mounted camera; inputting the environment image and the first internal reference matrix into an obstacle detection model, and performing obstacle detection according to the environment image and the first internal reference matrix through the obstacle detection model to obtain an obstacle detection result corresponding to the target vehicle; wherein the obstacle detection model is obtained based on multi-scale multi-view sample environment image training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent driving technology, and particularly relates to an obstacle detection method, an electronic device, and a storage medium. Background Art

[0002] Intelligent driving technology is a technology that enables a vehicle to drive autonomously without human control. It is based on advanced sensors, computer vision, artificial intelligence, machine learning, and other technologies, allowing the vehicle to perceive the surrounding environment, make decisions, and execute corresponding actions. As an important service of intelligent driving technology, the obstacle detection service uses environmental images from different perspectives of the vehicle to detect obstacles near the vehicle in the Bird’s Eye View (BEV) in real time, thus avoiding collisions during the planning and control process of the vehicle.

[0003] In related technologies, when performing obstacle detection, the requirements for the size of the vehicle's environmental image are relatively strict, and it is not robust to changes in the size of the vehicle's environmental image, resulting in poor detection effects when the vehicle's environmental image does not meet the size requirements. Summary of the Invention

[0004] Embodiments of this application provide an obstacle detection method, an electronic device, and a storage medium to solve the technical problem of poor detection effects caused by the non-robustness of the size change of the vehicle's environmental image in related technologies.

[0005] According to the first aspect of this application, an obstacle detection method is disclosed. The method includes:

[0006] Obtain environmental images of a target vehicle from different perspectives; the environmental images are collected by first in-vehicle cameras deployed at different positions on the body of the target vehicle;

[0007] Determine a first internal parameter matrix for obstacle detection according to the first image size of the environmental image and the original internal parameter matrix of the first in-vehicle camera;

[0008] Input the environmental image and the first internal parameter matrix into an obstacle detection model, and perform obstacle detection on the environmental image and the first internal parameter matrix through the obstacle detection model to obtain an obstacle detection result corresponding to the target vehicle; the obstacle detection model is trained based on multi-scale multi-perspective sample environmental images.

[0009] According to the second aspect of this application, an electronic device is disclosed, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the obstacle detection method as in the first aspect.

[0010] According to a third aspect of the present application, a computer-readable storage medium is disclosed, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the obstacle detection method in the first aspect is implemented.

[0011] According to a fourth aspect of the present application, a computer program product is disclosed, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the obstacle detection method in the first aspect is implemented.

[0012] In the embodiments of the present application, for a target vehicle that needs to perform obstacle detection, environmental images of different perspectives of the target vehicle and an internal parameter matrix corresponding to the image size of the environmental image can be input into an obstacle detection model for processing to obtain an obstacle detection result corresponding to the target vehicle. On the one hand, since the obstacle detection model is trained based on multi-scale multi-perspective sample environmental images, through multi-scale input, the training input of each obstacle size can be expanded, enriching the distribution of detection targets. On the other hand, since the internal parameter matrix for obstacle detection is adaptively adjusted according to the image size of the vehicle environmental image, the learning target can be fixed when the input image size changes, reducing the learning difficulty. Therefore, the robustness of the obstacle detection model for vehicle environmental images of different sizes can be improved, and thus obstacle detection for vehicle environmental images of any size can be realized, improving the detection effect. Description of the Drawings

[0013] Figure 1 is a flowchart of an obstacle detection method provided by an embodiment of the present application;

[0014] Figure 2 is an example diagram of an obstacle detection method provided by an embodiment of the present application;

[0015] Figure 3 is a flowchart of the processing process of the obstacle detection model provided by an embodiment of the present application;

[0016] Figure 4 is a structural diagram of the obstacle detection model provided by an embodiment of the present application;

[0017] Figure 5 is a flowchart of the training process of the obstacle detection model provided by an embodiment of the present application;

[0018] Figure 6 is a schematic structural diagram of an obstacle detection device provided by an embodiment of the present application;

[0019] Figure 7 is a structural block diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0020] To make the above objects, features, and advantages of the present application more apparent and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] It should be noted that for method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of the present application are not limited by the described action sequences, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present application.

[0022] In recent years, important progress has been made in the research of technologies such as computer vision, deep learning, machine learning, image processing, and image recognition based on artificial intelligence. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies, and application systems for simulating and extending human intelligence. The discipline of artificial intelligence is a comprehensive discipline, involving many technical categories such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, and neural networks. As an important branch of artificial intelligence, computer vision specifically enables machines to recognize the world. Computer vision technology usually includes face recognition, liveness detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, object detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping, computational photography, robot navigation and positioning, and other technologies. With the research and progress of artificial intelligence technology, this technology has been applied in many fields, such as security prevention, urban management, traffic management, building management, park management, face access, face attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile imaging, cloud services, smart home, wearable devices, driverless, autonomous driving, intelligent healthcare, face payment, face unlocking, fingerprint unlocking, person-certificate verification, smart screen, smart TV, cameras, mobile Internet, webcasting, beauty, makeup, medical beauty, intelligent temperature measurement, and other fields.

[0023] It should be noted that all the data obtained in the present application are accessed, collected, stored, and applied to subsequent analysis and processing after clearly informing the user or the relevant data owner of information such as the content of data collection, data usage, and processing methods, and with the consent and authorization of the user or the relevant data owner. Moreover, a way to access, correct, and delete the data, as well as a method to revoke consent and authorization, can be provided to the user or the relevant data owner.

[0024] The goal of bird's-eye view three-dimensional obstacle detection is to detect driving obstacles using environmental images from different perspectives of the vehicle, which is an important part of intelligent driving. The main difference between bird's-eye view three-dimensional obstacle detection and traditional two-dimensional detection is that it involves a depth estimation problem, that is, decoupling the three-dimensional spatial position into a two-dimensional detection and a depth estimation problem.

[0025] A representative method of the related technology is LSS. Its main idea is to first input environmental images from different perspectives of the vehicle, perform feature extraction, and explicitly perform depth estimation to construct a stereo frustum feature for each perspective. Then, according to the calibrated internal and external parameter matrices, map the image features to the bird's-eye view coordinate system for conventional two-dimensional detection. It can be seen that there are conversions between two spaces, namely the image space and the bird's-eye view space, during the learning process, and the corresponding features are strongly coupled. In this mode, the depth estimation problem depends on the style of the input image, such as the size of the image. If the size of the input image changes, it will greatly affect the subsequent feature distribution, making the existing methods basically ineffective. To solve the above technical problems, the embodiments of the present application provide an obstacle detection method, an electronic device, and a storage medium.

[0026] For ease of understanding, some concepts involved in the embodiments of the present application are introduced first below.

[0027] Camera Intrinsics: It is fixed after the camera leaves the factory and will not change during use, including focal lengths (fx, fy), principal point coordinates (cx, cy), distortion parameters, etc. Its function is to convert coordinates from the camera coordinate system to the image coordinate system.

[0028] Inverse Perspective Mapping (IPM): A physical transformation process that converts coordinates in the image coordinate system to the vehicle coordinate system through the internal and external parameters of the camera sensor.

[0029] Next, a bird's-eye view three-dimensional obstacle detection method provided by the embodiments of the present application will be introduced with reference to the accompanying drawings.

[0030] It should be noted that the obstacle detection method provided by the embodiments of the present application is applicable to electronic devices. In practical applications, the electronic device can be an in-vehicle device, a computer device, or a server, etc. The embodiments of the present application do not limit this.

[0031] Figure 1 It is a flowchart of an obstacle detection method provided by the embodiments of the present application. As Figure 1 shown, the method may include the following steps: Step 101, Step 102, and Step 103;

[0032] In step 101, environmental images from different perspectives of the target vehicle are obtained; wherein, the environmental images are captured by first vehicle-mounted cameras deployed at different positions on the body of the target vehicle.

[0033] In the embodiments of the present application, the target vehicle is a vehicle for intelligent driving. For example, it is a vehicle for autonomous driving or a vehicle for assisted driving.

[0034] In the embodiments of the present application, multiple first vehicle-mounted cameras can be deployed at different positions on the body of the target vehicle. During the driving process of the target vehicle, the environmental images around the target vehicle, that is, environmental images from multiple different perspectives, are captured in real time by the multiple first vehicle-mounted cameras.

[0035] In the embodiments of the present application, the environmental images from different perspectives of the target vehicle can be directly obtained from each of the first vehicle-mounted cameras deployed on the vehicle body.

[0036] In some embodiments, considering cost issues, the first vehicle-mounted cameras deployed on the vehicle body can be ordinary cameras for capturing normal images.

[0037] In some embodiments, considering that the fisheye camera has a larger field of view and the captured information is more comprehensive, the first vehicle-mounted cameras deployed on the vehicle body can be fisheye cameras.

[0038] In one example, 6 first vehicle-mounted cameras are deployed on the body of the target vehicle, which are respectively used to capture the environmental image in front of the target vehicle, the environmental image in the front left of the target vehicle, the environmental image in the front right of the target vehicle, the environmental image in the rear left of the target vehicle, the environmental image in the rear right of the target vehicle, and the environmental image behind the target vehicle.

[0039] It should be noted that in the embodiments of the present application, the deployment positions and quantities of the first vehicle-mounted cameras are not limited, as long as the environmental images around the vehicle can be captured.

[0040] In the embodiments of the present application, since the environmental images of each perspective are captured by first vehicle-mounted cameras of the same model, the environmental images from different perspectives of the target vehicle are single-scale images, that is, the sizes of all environmental images are the same.

[0041] In step 102, according to the first image size of the environmental image and the original internal parameter matrix of the first vehicle-mounted camera, a first internal parameter matrix for obstacle detection is determined.

[0042] In some embodiments, the above step 102 may include the following steps: step 1021 and step 1022;

[0043] In step 1021, according to the mapping relationship between the image size and the fine-tuning coefficient, determine the first fine-tuning coefficient corresponding to the first image size; wherein, the image size is positively correlated with the fine-tuning coefficient.

[0044] In the embodiments of the present application, the mapping relationship between the image size and the fine-tuning coefficient can be preset. The mapping relationship records the values of the fine-tuning coefficients corresponding to different image sizes. The larger the image size, the larger the value of the corresponding fine-tuning coefficient, and the smaller the image size, the smaller the value of the corresponding fine-tuning coefficient.

[0045] In the embodiments of the present application, the fine-tuning coefficient corresponding to the first image size can be found from the mapping relationship between the image size and the fine-tuning coefficient, and the found fine-tuning coefficient is determined as the first fine-tuning coefficient.

[0046] In step 1022, perform a multiplication operation on the original internal parameter matrix of the first vehicle-mounted camera and the first fine-tuning coefficient to obtain the first internal parameter matrix for obstacle detection.

[0047] In the embodiments of the present application, the first internal parameter matrix = the original internal parameter matrix of the first vehicle-mounted camera * the first fine-tuning coefficient.

[0048] It can be seen that in the embodiments of the present application, considering that the main difference between the bird's-eye view three-dimensional obstacle detection and the traditional two-dimensional detection is that it contains a depth estimation problem, that is, the three-dimensional space position is decoupled into a two-dimensional detection and a depth estimation problem, and the depth estimation problem depends on the size of the input image. If the size of the input image changes, it will greatly affect the subsequent feature distribution. Therefore, in order to reduce the difficulty of the model depth estimation, the adaptability of the camera internal parameters to the image size is considered in the depth estimation, and different internal parameter matrix inputs corresponding to different sizes of image inputs are proposed, so that the depth estimation learning target can be fixed when the input image size changes, and the learning difficulty can be reduced.

[0049] In step 103, input the environmental image and the first internal parameter matrix into the obstacle detection model, and perform obstacle detection on the environmental image and the first internal parameter matrix through the obstacle detection model to obtain the obstacle detection result corresponding to the target vehicle; wherein, the obstacle detection model is trained based on multi-scale multi-view sample environmental images.

[0050] In the embodiments of the present application, the obstacle detection result corresponding to the target vehicle includes: whether there are obstacles around the body of the target vehicle. When there are obstacles around the body of the target vehicle, the obstacle detection result further includes: the three-dimensional space coordinates of the obstacles.

[0051] In an example, such as Figure 2As shown, the input of the obstacle detection model is: environmental images of different perspectives of the target vehicle and the first internal parameter matrix, and the output of the obstacle detection model is: the obstacle detection result corresponding to the target vehicle.

[0052] As can be seen from the above embodiments, in this embodiment, for the target vehicle that needs to perform obstacle detection, the environmental images of different perspectives of the target vehicle and the internal parameter matrix corresponding to the image size of the environmental image can be input into the obstacle detection model for processing to obtain the obstacle detection result corresponding to the target vehicle. On the one hand, since the obstacle detection model is trained based on multi-scale multi-perspective sample environmental images, through multi-scale input, the training input of each obstacle size can be expanded, enriching the distribution of detection targets. On the other hand, since the internal parameter matrix for obstacle detection is adaptively adjusted according to the image size of the vehicle environmental image, the learning target can be fixed when the input image size changes, reducing the learning difficulty. Therefore, the robustness of the obstacle detection model for vehicle environmental images of different sizes can be improved, and then obstacle detection can be realized for vehicle environmental images of any size, improving the detection effect.

[0053] In some embodiments provided by the present application, as Figure 3 shown, the above step 103 may include the following steps: step 301, step 302, step 303, step 304, and step 305;

[0054] In step 301, visual image features of each environmental image are extracted through the obstacle detection model.

[0055] In the embodiments of the present application, first, visual image features of each environmental image in the image coordinate system are extracted

[0056] In some embodiments, when the environmental image is a fisheye environmental image, the above step 301 may include the following steps: perform distortion removal processing on each fisheye environmental image; extract visual image features of each image after the distortion removal processing.

[0057] In the embodiments of the present application, on the one hand, considering that currently in the business, the obstacle detection model is usually trained based on normal images, performing distortion removal processing on the fisheye environmental image, the obtained image will be closer to the domain of the pre-training dataset, which can accelerate the convergence of the model and obtain a better convergence effect; on the other hand, considering that the content of the image center region is more concerned in the business, when performing distortion removal processing on the fisheye environmental image, the center region of the fisheye environmental image is mainly intercepted.

[0058] In some embodiments, for each environmental image, multi-scale features of the environmental image can be extracted first, and then the multi-scale features are fused. Finally, the visual image features of the environmental image are obtained. Since multi-scale feature extraction and fusion can capture feature information at different scales, thereby improving feature representation, compared with single-scale feature extraction, multi-scale feature extraction and fusion can generate more discriminative and robust features, thus improving the performance of classification, detection, and recognition.

[0059] In step 302, based on each visual image feature and the first intrinsic matrix, a first depth estimate and semantic features are generated.

[0060] In the embodiments of the present application, after obtaining each visual image feature in the image coordinate system, based on each visual image feature and the first intrinsic matrix, depth estimation is performed, and the position information of the visual image feature from the image coordinate system is converted to the world coordinate system to generate a first depth estimate and semantic features.

[0061] In step 303, based on the first depth estimate and semantic features, a first bird's-eye view is generated.

[0062] In the embodiments of the present application, based on the extrinsic matrix of the camera, the position information of the bird's-eye view from the vehicle coordinate system can be converted to the world coordinate system, and a coordinate alignment operation is performed in the world coordinate system. Based on the corresponding relationship of the aligned position information, the first depth estimate and semantic features are converted into a first bird's-eye view.

[0063] It can be seen that in the embodiments of the present application, geometric priors from the physical world can be added to the generation process of the bird's-eye view, which can improve the feature quality and accelerate the model convergence speed.

[0064] In step 304, the first bird's-eye view feature of the first bird's-eye view is extracted.

[0065] In the embodiments of the present application, the first bird's-eye view feature of the first bird's-eye view is extracted to further refine the features.

[0066] In step 305, based on the first bird's-eye view feature, an obstacle detection result corresponding to the target vehicle is predicted.

[0067] In some embodiments, in order to obtain more accurate detection results, the first bird's-eye view feature can be subjected to feature enhancement processing; based on the bird's-eye view feature after feature enhancement processing, an obstacle detection result corresponding to the target vehicle is predicted.

[0068] In the embodiments of the present application, various feature enhancement methods are adopted, such as convolutional neural networks, gradient operators, etc., to perform feature enhancement processing on the bird's-eye view feature.

[0069] It can be seen that in the embodiments of the present application, only the environmental images of different perspectives of the target vehicle and the corresponding intrinsic matrix need to be input into the obstacle detection model, and the obstacle detection model processes the input images step by step according to the hierarchy of image-semantics-obstacle prediction, so that the obstacles in the environment where the target vehicle is located can be detected. Since the obstacle detection model is trained based on multi-scale multi-perspective sample environmental images, through multi-scale input, the training input of each obstacle size can be expanded, and the distribution of detection targets can be enriched. Also, because the intrinsic matrix for obstacle detection is adaptively adjusted according to the image size of the vehicle environmental image, the learning target can be fixed when the input image size changes, reducing the learning difficulty. Therefore, the robustness of the obstacle detection model for vehicle environmental images of different sizes can be improved, and then obstacle detection for vehicle environmental images of any size can be realized, improving the detection effect.

[0070] In some embodiments provided by the present application, corresponding to Figure 3 the processing process of the obstacle detection model shown in Figure 4 as shown, the obstacle detection model 40 may include: a target image encoding network 41, a target depth estimation network 42, a target view transformation network 43, a target bird's-eye view encoding network 44, and a target detection network 45; wherein, the target depth estimation network 42 is connected after the target image encoding network 41, the target view transformation network 43 is connected after the target depth estimation network 42, the target bird's-eye view encoding network 44 is connected after the target view transformation network 43, and the target detection network 45 is connected after the target bird's-eye view encoding network 44;

[0071] Extract the visual image features of each environmental image through the target image encoding network 41; wherein, in practical applications, the target image encoding network may adopt a Resnet network.

[0072] Generate the first depth estimation and semantic features through the target depth estimation network 42 according to each visual image feature and the first intrinsic matrix; wherein, in practical applications, the target depth estimation network may be a network based on the U-Net architecture.

[0073] Generate the first bird's-eye view through the target view transformation network 43 according to the first depth estimation and semantic features; wherein, the target view transformation network is used to convert the camera perspective view to the bird's-eye view. In practical applications, the target view transformation network may adopt a projection method represented by LSS, a sampling method represented by BEVFormer, a position encoding method represented by PETR, or directly use a transformer to automatically learn the mapping relationship between the two spaces.

[0074] Extract the first bird's-eye view feature of the first bird's-eye view through the target bird's-eye view encoding network 44; where, in practical applications, the target bird's-eye view encoding network can be a BEV network.

[0075] Based on the first bird's-eye view feature, predict the obstacle detection result corresponding to the target vehicle through the target detection network 45; where, in practical applications, the target detection network can be a region convolutional neural network.

[0076] In the embodiments of the present application, considering that there are two decoupled spaces in 3D object detection: the image space and the bird's-eye view space, the target image encoding network, the target depth estimation network, the target view transformation network, the target bird's-eye view encoding network, and the target detection network are introduced to decouple the 3D spatial position into a 2D detection and depth estimation problem, thereby realizing 3D obstacle detection from a bird's-eye view.

[0077] In the embodiments of the present application, the target image encoding network is a trained image encoding network, the target depth estimation network is a trained depth estimation network, the target view transformation network is a trained view transformation network, the target bird's-eye view encoding network is a trained bird's-eye view encoding network, and the target detection network is a trained detection network.

[0078] It can be seen that in the embodiments of the present application, a complete obstacle detection model for obstacle detection in the intelligent driving scenario is built through the target image encoding network, the target depth estimation network, the target view transformation network, the target bird's-eye view encoding network, and the target detection network, expanding the application scenario of obstacle detection, enabling obstacle detection for vehicle environment images of any size, and improving the detection effect.

[0079] In some embodiments provided by the present application, as Figure 5 shown, the training process of the obstacle detection model may include the following steps: Step 501, Step 502, Step 503, Step 504, Step 505, Step 506, Step 507, Step 508, and Step 509;

[0080] In Step 501, obtain the initial image encoding network, the initial depth estimation network, the initial view transformation network, the initial bird's-eye view encoding network, and the initial detection network.

[0081] In the embodiments of the present application, the initial image encoding network is an untrained image encoding network. For example, the initial image encoding network is an open-source Resnet network.

[0082] In the embodiments of the present application, the initial depth estimation network is an untrained depth estimation network. For example, the initial depth estimation network is an open-source U-Net network.

[0083] In the embodiments of the present application, the initial view transformation network is an untrained view transformation network. For example, the initial view transformation network is the open-source MVCNN network.

[0084] In the embodiments of the present application, the initial bird's-eye view encoding network is an untrained bird's-eye view encoding network. For example, the initial bird's-eye view encoding network is the open-source BEV network.

[0085] In the embodiments of the present application, the initial detection network is an untrained detection network. For example, the initial detection network is the open-source region convolutional neural network.

[0086] In step 502, a training set is obtained; wherein, the training set includes: multiple sample image groups and coordinate annotation information of sample obstacles corresponding to each sample image group, and each sample image group includes: sample environment images from multiple perspectives.

[0087] In the embodiments of the present application, the sample environment images in the training set can be noisy images to train a more robust and better-performing obstacle detection model.

[0088] In the embodiments of the present application, the sample environment images in the training set are collected by the second vehicle-mounted camera. In practical applications, the models of the first vehicle-mounted camera and the second vehicle-mounted camera can be the same or different.

[0089] In the embodiments of the present application, to ensure the comprehensiveness of information in the training samples, each sample image group can include: multiple sample environment images collected by the second vehicle-mounted camera within a target time period.

[0090] In the embodiments of the present application, the image sizes of the sample environment images in each sample image group are the same.

[0091] In step 503, for each sample image group, according to a preset first size list, the sample environment images in the sample image group are randomly sampled to a second image size, and the sample environment images of the second image size are input into the initial image encoding network for processing to obtain corresponding visual image features; wherein, the preset first size list includes multiple different image sizes.

[0092] For example, the preset first size list can include: the original image size, one-half of the original image size, one-quarter of the original image size, one-eighth of the original image size, one-sixteenth of the original image size, and so on.

[0093] In the embodiments of the present application, based on each sample image group, according to a preset first size list, the sample environment images in each sample image group can be randomly subjected to multi-scale transformation to obtain sample environment images of different sizes. For example, for the first sample image group, the image size of the first sample image group is respectively scaled to one-half of the original size and to one-fourth of the original size; for the second sample image group, the image size of the second sample image group is respectively scaled to one-half of the original size and the image size of the first sample image group is scaled to one-eighth of the original size, etc. The sample environment images of different sizes are used as training data for model training to expand the training input of each object size, enrich the distribution of samples, and enhance the robustness of the model to image changes through multi-scale training.

[0094] In step 504, according to the second image size and the original intrinsic matrix of the second vehicle-mounted camera used to collect the sample environment image, the second intrinsic matrix is determined.

[0095] In the embodiments of the present application, in order to reduce the difficulty of model depth estimation, for input sample environment images of different sizes, the intrinsic matrix corresponding to the input size is used for depth estimation.

[0096] In the embodiments of the present application, the second fine-tuning coefficient corresponding to the second image size can be determined according to the mapping relationship between the image size and the fine-tuning coefficient; the product operation is performed on the original intrinsic matrix of the second vehicle-mounted camera and the second fine-tuning coefficient to obtain the second intrinsic matrix.

[0097] In step 505, the visual image features of the sample environment images of the second image size and the second intrinsic matrix are input into the initial depth estimation network for processing to obtain the second depth estimation and semantic features.

[0098] In step 506, the second depth estimation and semantic features are input into the initial view transformation network for processing to obtain the second bird's-eye view.

[0099] In some embodiments, while performing multi-scale training on the model in the sample environment image space, multi-scale training is also performed on the model in the bird's-eye view space to further enhance the robustness of the model to image changes. Correspondingly, the above step 506 may include the following steps: step 5061 and step 5062;

[0100] In step 5061, an image size is randomly selected from a preset second size list as the third image size of the bird's-eye view output by the initial view transformation network; wherein, the preset second size list includes multiple different image sizes.

[0101] For example, the preset second size list may include: the original image size, one-half of the original image size, one-fourth of the original image size, one-eighth of the original image size, one-sixteenth of the original image size, and so on.

[0102] In the embodiments of the present application, the size of the second bird's-eye view can be controlled by controlling the output size of the initial view conversion network.

[0103] In step 5062, the second depth estimate and the semantic features are input into the initial view conversion network for processing to obtain a second bird's-eye view of the third image size.

[0104] It can be seen that in the embodiments of the present application, bird's-eye views of different sizes can be used as intermediate training data for model training to expand the training input of each object size, enrich the distribution of samples, and enhance the robustness of the model to image changes through multi-scale training.

[0105] In step 507, the second bird's-eye view is input into the initial bird's-eye view encoding network for processing to obtain a second bird's-eye view feature.

[0106] In step 508, the second bird's-eye view feature is input into the initial detection network for processing to obtain coordinate prediction information of the sample obstacle.

[0107] In step 509, based on the coordinate annotation information and the coordinate prediction information of the sample obstacle, a loss value is calculated, and the network parameters of the initial image encoding network, the initial depth estimation network, the initial view conversion network, the initial bird's-eye view encoding network, and the initial detection network are adjusted according to the loss value; the above training process is repeated until the model converges to obtain an obstacle detection model.

[0108] In some embodiments, when performing multi-scale training on the model in the bird's-eye view space, the above step 508 may include the following steps: step 5081 and step 5082;

[0109] In step 5081, according to the third image size and the coordinate annotation information of the sample obstacle, the coordinate true information of the sample obstacle corresponding to the third image size is determined.

[0110] In step 5082, a loss value is calculated according to the coordinate true information and the coordinate prediction information of the sample obstacle.

[0111] In the embodiments of the present application, considering that the coordinate annotation information of the sample obstacle is marked based on the original image size, when the third image size of the second bird's-eye view is different from the original image size, it is necessary to proportionally adjust the coordinate annotation information of the sample obstacle to obtain the coordinate true information of the sample obstacle that matches the current bird's-eye view size.

[0112] In some embodiments, the loss value during model training can be calculated based on the L1 loss function, and the parameters of each network in the model can be updated by backpropagation. Alternatively, the loss value during model training can also be calculated based on other loss functions, and the parameters of each network in the model can be updated by backpropagation. The embodiments of the present application do not limit this.

[0113] In the embodiments of the present application, during the model training process, the initial model needs to be trained multiple times to converge and obtain an obstacle detection model. For any one of the trainings, a training data can be obtained through step 503 above, and steps 504 to 509 above are executed on the training data to complete an update of the model parameters. Then, the next training is performed. Another training data is obtained through step 503 above, and steps 504 to 509 above are executed on the training data to complete another update of the model parameters, and so on, until the model converges and the training stops to obtain the obstacle detection model.

[0114] In the embodiments of the present application, each time during training, the training data used is different and has a certain degree of randomness. Exemplarily, the training data for two adjacent trainings can be from the same sample image group. For example, the training data for the previous training is a sample environment image of the original size of the sample image group, and the training data for the next training is a sample environment image of half of the original size of the sample image group; or, the training data for two adjacent trainings can be from two sample image groups. For example, the training data for the previous training is a sample environment image of the original size of one sample image group, and the training data for the next training is a sample environment image of the original size of another sample image group, or the training data for the previous training is a sample environment image of the original size of one sample image group, and the training data for the next training is a sample environment image of half of the original size of another sample image group, or the training data for the previous training is a sample environment image of half of the original size of one sample image group, and the training data for the next training is a sample environment image of half of the original size of another sample image group, etc.

[0115] It can be seen that in the embodiments of the present application, the robustness of the model to image changes can be enhanced through multi-scale training of the environmental image space input and the bird's-eye view space input; by matching the corresponding internal parameter inputs for the multi-scale inputs, the difficulty of model depth estimation can be reduced, and the robustness of the model can be further improved.

[0116] In some embodiments provided by the present application, in order to improve the accuracy of the detection result, step 101 above may include the following steps: step 1011;

[0117] In step 1011, a plurality of environmental images collected by each first vehicle-mounted camera within the target duration are obtained.

[0118] In the embodiments of the present application, the target duration can be set according to the actual test situation, for example, set to 10 milliseconds.

[0119] In the embodiments of the present application, in order to ensure the comprehensiveness of the information in the training samples, each sample image group also includes: a plurality of sample environmental images collected by the second vehicle-mounted camera within the target duration.

[0120] It can be seen that in the embodiments of the present application, since the panoramic view image for obstacle detection is a multi-view image within a period of time and contains more comprehensive environmental information of the target vehicle, the detection result is more accurate.

[0121] Figure 6 It is a schematic structural diagram of an obstacle detection device provided by the embodiments of the present application, as Figure 6 shown, the obstacle detection device 600 may include: an acquisition module 601, a determination module 602, and a detection module 603;

[0122] The acquisition module 601 is configured to acquire environmental images of different perspectives of the target vehicle; the environmental images are acquired by first vehicle-mounted cameras deployed at different positions on the body of the target vehicle;

[0123] The determination module 602 is configured to determine a first internal parameter matrix for obstacle detection according to the first image size of the environmental image and the original internal parameter matrix of the first vehicle-mounted camera;

[0124] The detection module 603 inputs the environmental image and the first internal parameter matrix into an obstacle detection model, and performs obstacle detection on the environmental image and the first internal parameter matrix through the obstacle detection model to obtain an obstacle detection result corresponding to the target vehicle; the obstacle detection model is trained based on multi-scale multi-view sample environmental images.

[0125] As can be seen from the above embodiments, in this embodiment, for a target vehicle that needs to perform obstacle detection, environmental images of different perspectives of the target vehicle and the internal parameter matrix corresponding to the image size of the environmental image can be input into an obstacle detection model for processing to obtain an obstacle detection result corresponding to the target vehicle. On the one hand, since the obstacle detection model is trained based on multi-scale multi-view sample environmental images, through multi-scale input, the training input of each obstacle size can be expanded, and the distribution of detection targets can be enriched. On the other hand, since the internal parameter matrix for obstacle detection is adaptively adjusted according to the image size of the vehicle environmental image, the learning target can be fixed when the input image size changes, reducing the learning difficulty. Therefore, the robustness of the obstacle detection model for vehicle environmental images of different sizes can be improved, and thus obstacle detection for vehicle environmental images of any size can be realized, improving the detection effect.

[0126] Optionally, as an embodiment, the determining module 602 may include:

[0127] A first determination sub-module, configured to determine a first fine-tuning coefficient corresponding to the first image size according to the mapping relationship between the image size and the fine-tuning coefficient; wherein, the image size is positively correlated with the fine-tuning coefficient;

[0128] A second determination sub-module, configured to perform a multiplication operation on the original internal parameter matrix of the first vehicle-mounted camera and the first fine-tuning coefficient to obtain a first internal parameter matrix for obstacle detection.

[0129] Optionally, as an embodiment, the detection module 603 may include:

[0130] A first extraction sub-module, configured to extract visual image features of each of the environmental images through the obstacle detection model;

[0131] A first generation sub-module, configured to generate first depth estimation and semantic features according to each of the visual image features and the first internal parameter matrix;

[0132] A second generation sub-module, configured to generate a first bird's-eye view according to the first depth estimation and semantic features;

[0133] A second extraction sub-module, configured to extract first bird's-eye view features of the first bird's-eye view;

[0134] A prediction sub-module, configured to predict an obstacle detection result corresponding to the target vehicle according to the first bird's-eye view features.

[0135] Optionally, as an embodiment, the obstacle detection model may include: a target image encoding network, a target depth estimation network, a target view conversion network, a target bird's-eye view encoding network, and a target detection network; the target depth estimation network is connected after the target image encoding network, the target view conversion network is connected after the target depth estimation network, the target bird's-eye view encoding network is connected after the target view conversion network, and the target detection network is connected after the target bird's-eye view encoding network;

[0136] Extract visual image features of each of the environmental images through the target image encoding network;

[0137] Generate first depth estimation and semantic features through the target depth estimation network according to each of the visual image features and the first internal parameter matrix;

[0138] Generate a first bird's-eye view through the target view conversion network according to the first depth estimation and semantic features;

[0139] Extract the first bird's-eye view feature of the first bird's-eye view through the target bird's-eye view encoding network;

[0140] Based on the first bird's-eye view feature, predict the obstacle detection result corresponding to the target vehicle through the target detection network.

[0141] Optionally, as an embodiment, the training process of the obstacle detection model may include:

[0142] Obtain an initial image encoding network, an initial depth estimation network, an initial view transformation network, an initial bird's-eye view encoding network, and an initial detection network;

[0143] Obtain a training set; wherein, the training set includes: multiple groups of sample images and coordinate annotation information of sample obstacles corresponding to each group of sample images, and each group of sample images includes: sample environment images from multiple perspectives;

[0144] For each group of sample images, randomly sample the sample environment images in the group of sample images to a second image size according to a preset first size list, and input the sample environment images of the second image size into the initial image encoding network for processing to obtain corresponding visual image features; wherein, the preset first size list includes multiple different image sizes;

[0145] Determine a second internal parameter matrix according to the second image size and the original internal parameter matrix of the second vehicle-mounted camera used to collect the sample environment images;

[0146] Input the visual image features of the sample environment images of the second image size and the second internal parameter matrix into the initial depth estimation network for processing to obtain a second depth estimation and semantic feature;

[0147] Input the second depth estimation and semantic feature into the initial view transformation network for processing to obtain a second bird's-eye view;

[0148] Input the second bird's-eye view into the initial bird's-eye view encoding network for processing to obtain a second bird's-eye view feature;

[0149] Input the second bird's-eye view feature into the initial detection network for processing to obtain the coordinate prediction information of the sample obstacle;

[0150] Based on the coordinate annotation information and coordinate prediction information of the sample obstacle, calculate a loss value, and adjust the network parameters of the initial image encoding network, initial depth estimation network, initial view transformation network, initial bird's-eye view encoding network, and initial detection network according to the loss value;

[0151] Repeat the above training process until the model converges to obtain the obstacle detection model.

[0152] Optionally, as an embodiment, the training process of the obstacle detection model may include:

[0153] Randomly select an image size from a preset second size list as the third image size of the bird's-eye view output by the initial view transformation network; wherein, the preset second size list includes multiple different image sizes;

[0154] Input the second depth estimation and semantic features into the initial view transformation network for processing to obtain a second bird's-eye view of the third image size.

[0155] Optionally, as an embodiment, the obtaining module 601 may include:

[0156] An obtaining sub-module, configured to obtain a plurality of environmental images collected by each of the first vehicle-mounted cameras within a target duration.

[0157] Any step in the embodiment of the obstacle detection method provided in this application and the specific operations in any step can be completed by the corresponding module in the obstacle detection device. The process of the corresponding operations completed by each module in the obstacle detection device refers to the process of the corresponding operations described in the embodiment of the obstacle detection method.

[0158] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiment.

[0159] Figure 7 It is a block diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device includes a processing component 722, which further includes one or more processors, and memory resources represented by a memory 732 for storing instructions executable by the processing component 722, such as application programs. The application programs stored in the memory 732 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 722 is configured to execute instructions to perform the above method.

[0160] The electronic device may further include a power component 726 configured to perform power management of the electronic device, a wired or wireless network interface 750 configured to connect the electronic device to a network, and an input / output (I / O) interface 758. The electronic device may operate based on an operating system stored in the memory 732, such as Windows ServerTM, MacOS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.

[0161] According to another embodiment of the present application, the present application further provides a computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps in the obstacle detection method described in any of the above embodiments are implemented.

[0162] According to another embodiment of the present application, the present application further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps in the obstacle detection method described in any of the above embodiments are implemented.

[0163] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0164] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0165] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing terminal devices generate a device for implementing the specified functions in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0166] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the specified functions in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0167] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0168] Finally, it should also be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.

[0169] The above has introduced in detail a method for obstacle detection, an electronic device and a storage medium provided by the present application. Specific examples are used in this text to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for obstacle detection, characterized in that, The method includes: Obtaining environmental images of different perspectives of a target vehicle; the environmental images are captured by first vehicle-mounted cameras deployed at different orientations on the body of the target vehicle; Determining a first intrinsic matrix for obstacle detection according to a first image size of the environmental images and an original intrinsic matrix of the first vehicle-mounted camera; Inputting the environmental images and the first intrinsic matrix into an obstacle detection model, and performing obstacle detection according to the environmental images and the first intrinsic matrix through the obstacle detection model to obtain an obstacle detection result corresponding to the target vehicle; the obstacle detection model is trained based on multi-scale multi-perspective sample environmental images.

2. The method according to claim 1, wherein The determining a first intrinsic matrix for obstacle detection according to a first image size of the environmental images and an original intrinsic matrix of the first vehicle-mounted camera includes: Determining a first fine-tuning coefficient corresponding to the first image size according to a mapping relationship between an image size and a fine-tuning coefficient; wherein, the image size is positively correlated with the fine-tuning coefficient; Performing a multiplication operation on the original intrinsic matrix of the first vehicle-mounted camera and the first fine-tuning coefficient to obtain a first intrinsic matrix for obstacle detection.

3. The method according to claim 1 or 2, characterized in that, The performing obstacle detection according to the environmental images and the first intrinsic matrix through the obstacle detection model to obtain an obstacle detection result corresponding to the target vehicle includes: Extracting visual image features of each of the environmental images through the obstacle detection model; Generating a first depth estimation and semantic feature according to each of the visual image features and the first intrinsic matrix; Generating a first bird's-eye view according to the first depth estimation and semantic feature; Extracting a first bird's-eye view feature of the first bird's-eye view; Predicting an obstacle detection result corresponding to the target vehicle according to the first bird's-eye view feature.

4. The method according to claim 3, wherein The obstacle detection model includes: a target image encoding network, a target depth estimation network, a target view transformation network, a target bird's-eye view encoding network, and a target detection network; the target depth estimation network is connected after the target image encoding network, the target view transformation network is connected after the target depth estimation network, the target bird's-eye view encoding network is connected after the target view transformation network, and the target detection network is connected after the target bird's-eye view encoding network; Extracting visual image features of each of the environmental images through the target image encoding network; Generating a first depth estimation and semantic feature according to each of the visual image features and the first intrinsic matrix through the target depth estimation network; Generating a first bird's-eye view according to the first depth estimation and semantic feature through the target view transformation network; Extracting a first bird's-eye view feature of the first bird's-eye view through the target bird's-eye view encoding network; Predicting an obstacle detection result corresponding to the target vehicle according to the first bird's-eye view feature through the target detection network.

5. The method according to claim 4, characterized in that, The training process of the obstacle detection model includes: Obtaining an initial image encoding network, an initial depth estimation network, an initial view transformation network, an initial bird's-eye view encoding network, and an initial detection network; Obtain a training set; wherein, the training set includes: a plurality of sample image groups and coordinate annotation information of sample obstacles corresponding to each sample image group, and each sample image group includes: sample environment images from multiple perspectives; For each sample image group, according to a preset first size list, randomly sample each sample environment image in the sample image group to a second image size, and input each sample environment image of the second image size into the initial image encoding network for processing to obtain corresponding visual image features; wherein, the preset first size list includes a plurality of different image sizes; Determine a second intrinsic matrix according to the second image size and the original intrinsic matrix of the second vehicle-mounted camera used to collect the sample environment images; Input the visual image features of each sample environment image of the second image size and the second intrinsic matrix into the initial depth estimation network for processing to obtain second depth estimation and semantic features; Input the second depth estimation and semantic features into the initial view transformation network for processing to obtain a second bird's-eye view; Input the second bird's-eye view into the initial bird's-eye view encoding network for processing to obtain second bird's-eye features; Input the second bird's-eye features into the initial detection network for processing to obtain coordinate prediction information of the sample obstacles; Based on the coordinate annotation information and coordinate prediction information of the sample obstacles, calculate a loss value, and adjust the network parameters of the initial image encoding network, initial depth estimation network, initial view transformation network, initial bird's-eye view encoding network, and initial detection network according to the loss value; Repeat the above training process until the model converges to obtain the obstacle detection model.

6. The method according to claim 5, wherein The step of inputting the second depth estimation and semantic features into the initial view transformation network for processing to obtain a second bird's-eye view includes: Randomly select an image size from a preset second size list as the third image size of the bird's-eye view output by the initial view transformation network; wherein, the preset second size list includes a plurality of different image sizes; Input the second depth estimation and semantic features into the initial view transformation network for processing to obtain a second bird's-eye view of the third image size.

7. The method according to claim 1, characterized in that The step of obtaining environment images of different perspectives of the target vehicle includes: Obtain a plurality of environment images collected by each of the first vehicle-mounted cameras during a target time period.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the obstacle detection method according to any one of claims 1-7.

9. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the obstacle detection method according to any one of claims 1-7 is implemented.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the obstacle detection method according to any one of claims 1-7 is implemented.