Image depth-based vehicle control method, device, equipment and storage medium

By identifying the depth range of objects in the image of the environment in front of the vehicle, and combining neural network models and control rules, the problems of slow calculation speed and low accuracy in existing technologies are solved, and efficient and safe vehicle control is achieved.

CN116142173BActive Publication Date: 2026-04-28AUTOMOTIVE INTELLIGENCE & CONTROL OF CHINA CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AUTOMOTIVE INTELLIGENCE & CONTROL OF CHINA CO LTD
Filing Date
2023-01-19
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, vehicle control methods require the calculation of specific distance values, resulting in slow calculation speed and low accuracy. This fails to meet the real-time and low-computing-power requirements of autonomous driving, thus affecting driving safety.

Method used

By acquiring images of the environment in front of the vehicle, identifying objects and determining their image depth range rather than specific distance values, a neural network model is used for object recognition and depth range determination, combined with preset control rules to control the vehicle's movement.

Benefits of technology

It improves the precision and efficiency of vehicle control, meets the real-time requirements of autonomous driving, and ensures driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116142173B_ABST
    Figure CN116142173B_ABST
Patent Text Reader

Abstract

The application provides a vehicle control method and device based on image depth, equipment and storage medium. The method is applied to a vehicle and includes: acquiring an environment image in a preset space range in front of the vehicle; identifying an object in the environment image, and determining a depth interval of an image depth corresponding to the object in the environment image; wherein the depth interval is used to represent a distance range between the object and the vehicle; and controlling the vehicle to travel according to the depth interval of the image depth corresponding to the object in the environment image. The application determines the distance range between the vehicle and the front object, controls the vehicle, improves the accuracy, is faster in calculation, and improves the vehicle control efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to autonomous driving technology, and more particularly to a vehicle control method, apparatus, device, and storage medium based on image depth. Background Technology

[0002] With rapid economic development, the number of vehicles on the road is increasing, leading to frequent traffic accidents caused by driver fatigue, drunk driving, or improper driving skills. Road traffic accidents bring enormous economic losses and emotional distress to individuals, families, and society as a whole. Calculating the distance between the vehicle and pedestrians, vehicles, and obstacles in front of it can ensure safe driving.

[0003] In existing technologies, Markov random fields are used to learn the mapping relationship between the features of the input image and the depth of the output. By utilizing multi-scale depth features such as texture and blur in the image, Gaussian Markov random field models and Laplace Markov random field models are constructed to estimate the depth of a single image.

[0004] However, the model established in this method is applicable to very limited scenarios, and requires the calculation of specific distance values ​​of image depth, which is computationally intensive and time-consuming. The accuracy and efficiency of determining image depth are low, resulting in low accuracy and efficiency of vehicle control. Summary of the Invention

[0005] This application provides a vehicle control method, apparatus, device, and storage medium based on image depth to improve the accuracy and efficiency of vehicle control.

[0006] In a first aspect, this application provides a vehicle control method based on image depth, which is applied to a vehicle and includes:

[0007] Acquire environmental images within a preset space area in front of the vehicle;

[0008] Identify objects in the environmental image and determine the depth range of the image depth corresponding to the objects in the environmental image; wherein, the depth range is used to represent the distance range between the objects and the vehicle;

[0009] The vehicle is controlled to move based on the depth range of the image depth corresponding to the objects in the environmental image.

[0010] Secondly, this application provides a vehicle control device based on image depth, which is applied to a vehicle and includes:

[0011] The image acquisition module is used to acquire environmental images within a preset spatial range in front of the vehicle;

[0012] An interval determination module is used to identify objects in the environmental image and determine the depth interval of the image depth corresponding to the objects in the environmental image; wherein, the depth interval is used to represent the distance range between the object and the vehicle;

[0013] The vehicle control module is used to control the vehicle to move based on the depth range of the image depth corresponding to the objects in the environmental image.

[0014] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0015] The memory stores computer-executed instructions;

[0016] The processor executes computer execution instructions stored in the memory to implement the image depth-based vehicle control method as described in the first aspect of this application.

[0017] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the image depth-based vehicle control method as described in the first aspect of this application.

[0018] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the image depth-based vehicle control method as described in the first aspect of this application.

[0019] This application provides a vehicle control method, apparatus, device, and storage medium based on image depth. It acquires environmental images in front of the vehicle using image acquisition equipment installed on the vehicle, identifies objects in the environmental images, and determines the depth range of the objects in the image, i.e., determines the distance range between the objects and the vehicle. Vehicle movement is controlled based on this distance range, improving driving safety. This solves the problems of slow calculation speed or low calculation accuracy caused by the need to calculate specific distance values ​​in existing technologies. By determining the depth range, the accuracy of the calculation is improved, and the calculation speed is faster, meeting the low computing power requirements of the vehicle-side computing platform and improving the accuracy and efficiency of vehicle control. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0021] Figure 1 A schematic flowchart of a vehicle control method based on image depth provided in an embodiment of this application;

[0022] Figure 2This is a schematic diagram of a vehicle equipped with an image acquisition device, provided as an embodiment of this application.

[0023] Figure 3 A schematic flowchart of a vehicle control method based on image depth provided in an embodiment of this application;

[0024] Figure 4 A structural block diagram of a vehicle control device based on image depth provided in an embodiment of this application;

[0025] Figure 5 A structural block diagram of a vehicle control device based on image depth provided in an embodiment of this application;

[0026] Figure 6 A structural block diagram of an electronic device provided in an embodiment of this application;

[0027] Figure 7 This is a structural block diagram of an electronic device provided in an embodiment of this application.

[0028] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0030] It should be understood that the described embodiments are merely some, not all, of the embodiments in this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0031] In the following description, when referring to the accompanying drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0032] In the description of this application, it should be understood that the terms "first," "second," "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0033] It should be noted that, due to space limitations, this application specification does not exhaustively list all possible implementation methods. Those skilled in the art, after reading this application specification, should be able to deduce that, as long as the technical features do not contradict each other, any combination of technical features can constitute an optional implementation method. The following provides a detailed description of each embodiment.

[0034] In recent years, with rapid economic development, the number of vehicles on the road has increased dramatically. Traffic accidents caused by driver fatigue, drunk driving, or improper driving skills are frequent. Statistics show that approximately 1.3 million lives are lost in road traffic accidents worldwide each year, and another 20 to 50 million suffer non-fatal injuries. Road traffic accidents bring enormous economic losses and emotional anguish to individuals and families. Annual losses from road traffic collisions account for 3% of GDP. Autonomous driving technology can significantly reduce the number of road traffic accidents. Through precise algorithms and sophisticated sensor equipment, vehicles can be driven safely. Monocular depth estimation, in particular, can calculate the distance between the vehicle and pedestrians, vehicles, and obstacles in front, ensuring safe driving.

[0035] Traditional monocular image depth estimation algorithms model depth relationships based on Markov random fields or conditional random fields in machine learning. They solve for image depth by minimizing an energy function based on maximum a posteriori probability. Depending on whether the model contains parameters, this method can be further divided into parametric learning methods and non-parametric learning methods. Parametric learning methods assume that the model contains unknown parameters, and the training process is essentially solving for these unknown parameters. Typically, Markov random fields are used to learn the mapping relationship between input image features and output depth. Gaussian Markov random field models and Laplacian Markov random field models are constructed using multi-scale texture and blur depth features in the image to estimate the depth of a single image. However, the models assumed by parametric learning methods are difficult to simulate real-world mapping relationships, limiting their applicability. Non-parametric learning methods infer image depth by using existing datasets for similarity retrieval, without needing to learn parameters. However, they rely on image retrieval, resulting in high computational costs and time consumption, which cannot meet the real-time and low-computing-power performance requirements of autonomous driving.

[0036] Furthermore, existing technologies require the calculation of depth values ​​for distances, which is computationally difficult and prone to errors. This can lead to mistakes in vehicle control, resulting in lower accuracy and efficiency and impacting driving safety.

[0037] This application provides a vehicle control method, apparatus, device, and storage medium based on image depth, which aims to solve the above-mentioned technical problems in the prior art.

[0038] The technical solution of this application and how it solves the above-mentioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0039] Figure 1 This is a flowchart illustrating a vehicle control method based on image depth according to an embodiment of this application. The method is applied to a vehicle and can be executed by a vehicle control device based on image depth. Figure 1 As shown, the method includes the following steps:

[0040] S101. Obtain an environmental image within a preset space range in front of the vehicle.

[0041] For example, an image acquisition device can be installed on the vehicle. This device could be a camera, capable of acquiring images of the environment surrounding the vehicle in real time or at set intervals. A preset image acquisition cycle can be established; for example, a cycle of one minute would allow the device to acquire an environmental image every minute. While the vehicle is in motion, to prevent collisions with pedestrians, other vehicles, or other obstacles, the image acquisition device can capture images of the environment in front of the vehicle. For instance, the image acquisition device could be mounted at the front of the vehicle. Figure 2 A schematic diagram of a vehicle equipped with image acquisition equipment. Figure 2 The image acquisition device is installed on the roof inside the vehicle and can capture environmental images facing forward.

[0042] The spatial range that the image acquisition device can capture can be preset. The image acquisition device captures environmental images within the preset spatial range, and the vehicle obtains environmental images through the image acquisition device. That is, the vehicle can acquire environmental images within the preset spatial range in front of it in real time or at regular intervals.

[0043] S102. Identify objects in the environmental image and determine the depth range of the image depth corresponding to the objects in the environmental image; wherein, the depth range is used to represent the distance range between the object and the vehicle.

[0044] For example, when performing target recognition on an environmental image, a target recognition algorithm can be pre-set to identify the targets as objects in the environmental image. Objects in the environmental image can be vehicles, pedestrians, and other obstacles, and an environmental image can contain one or more objects. For instance, a target recognition network model can be pre-trained to identify vehicles, pedestrians, and other obstacles in the image.

[0045] One or more depth ranges can be predefined, and the range of these ranges can be used to represent the distance between an object and a vehicle. For example, a depth range of 1 to 2 meters indicates that the distance between the object and the vehicle is between 1 and 2 meters. When determining the image depth of each object in an environmental image, the depth range in which the image depth of each object lies can be determined, that is, the range of the image depth of each object can be determined. For example, if there is a pedestrian in the environmental image, and the depth range of the pedestrian's image depth is 3 to 5 meters, then the distance between the pedestrian and the vehicle can be determined to be between 3 and 5 meters.

[0046] An algorithm for determining depth intervals can be pre-set. For example, the depth interval of an object can be determined based on its size, color, and lighting in the environmental image. When determining the depth interval, it is not necessary to determine the specific image depth value, avoiding errors in image depth value calculation. Multiple depth intervals can be pre-set. For example, the area within 100 meters in front of the vehicle can be divided into multiple depth intervals, each with a total length of 100 meters. The depth intervals can be represented as [0, 1], [1, 3], [3, 8], [8, 15], and [15, 30], etc. If the depth interval is determined based on the object's size, the correlation between different object sizes and depth intervals can be pre-set, thereby determining the depth interval where the object is located, reducing the amount of computation and improving computational efficiency.

[0047] S103. Control the vehicle to drive based on the depth range of the image depth corresponding to the object in the environmental image.

[0048] For example, after determining the depth range of each object in the environmental image, intelligent control of the vehicle can be performed based on the depth range. For instance, if there is a car in front of the vehicle in the environmental image, and the car is located in the depth range [2, 3], meaning the distance between the car and the vehicle is between 2 and 3 meters, a deceleration command can be sent to the vehicle to control it to slow down and avoid a collision.

[0049] Vehicles can also preset depth ranges for issuing warning messages, designated as danger depth ranges. That is, when an object in front of the vehicle is within the preset danger depth range, the vehicle can issue a warning message to the user, reminding them to take timely action to avoid danger. For example, if the danger depth range is [0, 1], then when an object is less than 1 meter away from the vehicle, the vehicle can issue an audible and visual alarm to remind the user to pay attention to driving safety.

[0050] If there are multiple objects in the environmental image, each located in a different depth range, the object closest to the vehicle can be determined based on its depth range. The vehicle's movement is then controlled based on the depth range of the closest object.

[0051] In this embodiment, controlling the vehicle to drive based on the depth range of the image depth corresponding to the object in the environmental image includes: determining a control rule associated with the depth range of the image depth corresponding to the object in the environmental image based on the relationship between a preset depth range and a control rule, which is the current control rule; and controlling the vehicle to drive according to the current control rule.

[0052] Specifically, the vehicle's control actions can differ depending on the depth range of the object in front of it. For example, if the object is too close to the vehicle, the vehicle needs to slow down and sound its horn; if the object is too far away, the vehicle continues to drive normally.

[0053] Multiple control rules for the vehicle can be preset. The vehicle performs driving control according to these rules, which can include driving operations such as stopping, decelerating, accelerating, and honking. Different depth ranges and their associated control rules are preset. After obtaining the depth range of an object, the control rule associated with that depth range is determined as the current control rule. For example, preset control rules include Control Rule 1 and Control Rule 2. The depth range associated with Control Rule 1 is [1, 2], and the depth range associated with Control Rule 2 is [2, 4]. If the object's depth range is [1, 2], then Control Rule 1 can be determined as the current control rule. The current control rule is executed to intelligently control the vehicle. For example, Control Rule 1 could be that the vehicle decelerates at a preset acceleration until it stops.

[0054] The advantage of this setup is that it can quickly determine the control operations required for objects in front of the vehicle based on preset control rules, thereby improving the efficiency of vehicle control, realizing intelligent vehicle operation, and ensuring safe driving.

[0055] This application provides a vehicle control method based on image depth. It acquires environmental images of the area in front of the vehicle using an image acquisition device installed on the vehicle, identifies objects in the environmental images, and determines the depth range of the objects within the image, i.e., determines the distance range between the objects and the vehicle. Vehicle movement is controlled based on this distance range, improving driving safety. This method solves the problems of slow calculation speed or low calculation accuracy caused by the need to calculate specific distance values ​​in existing technologies. By determining the depth range, the accuracy of the calculation is improved, and the calculation speed is faster, meeting the low computing power requirements of the vehicle-side computing platform and improving the accuracy and efficiency of vehicle control.

[0056] Figure 3 This is a flowchart illustrating a vehicle control method based on image depth, which is an optional embodiment based on the above embodiments.

[0057] In this embodiment, identifying objects in an environmental image and determining the depth range of the image depth corresponding to the objects in the environmental image can be refined as follows: inputting the environmental image into a preset neural network model to obtain the objects in the environmental image; determining the probability that the image depth of the objects in the environmental image falls within each pre-divided depth range; wherein, at least one depth range is pre-divided within a preset distance in front of the vehicle; and determining the depth range of the image depth corresponding to the objects in the environmental image based on the probability that the objects in the environmental image fall within each pre-divided depth range.

[0058] like Figure 3 The method includes the following steps:

[0059] S301. Acquire an environmental image within a preset space range in front of the vehicle.

[0060] For example, this step can refer to step S101 above, and will not be repeated here.

[0061] S302. Input the environmental image into the preset neural network model to obtain the objects in the environmental image.

[0062] For example, a neural network model can be pre-trained, such as a deep learning neural network model, to determine objects in an environmental image and the depth range of each object. The environmental image is input into the neural network model, which is then pre-trained to recognize pedestrians, vehicles, trees, animals, and other obstacles in the image. The neural network model extracts features from the input environmental image and outputs the objects within the environmental image.

[0063] In this embodiment, the environmental image is input into a preset neural network model to obtain objects in the environmental image, including: inputting the environmental image into a U-shaped network in the preset neural network model; extracting features from the environmental image based on the U-shaped network to obtain feature vectors of the environmental image; determining the foreground part of the environmental image based on the feature vectors, and identifying objects in the foreground part as objects in the environmental image.

[0064] Specifically, the U-net (U-shaped network) structure can be used as the backbone network of the neural network. The environmental image is input into the U-net network, which extracts features from the image, such as through convolution, downsampling, and deconvolution. After feature extraction, the U-net network outputs a feature vector of the environmental image.

[0065] Based on the feature vectors of the environmental image, the foreground portion is segmented from the image. The foreground portion refers to the image parts of pedestrians, vehicles, and obstacles, while the background portion can refer to the image parts of the sky, road, trees on both sides of the road, and signs hanging on the road. After obtaining the foreground portion of the environmental image, objects in the foreground portion are identified and determined as objects in the environmental image. That is, when determining objects in the environmental image, objects such as signs in the road background can be disregarded.

[0066] The advantage of this setup is that by pre-training the neural network model, it can adapt to various road scenarios. By using the U-net network and segmenting the foreground, objects can be accurately identified, improving object recognition accuracy and thus enhancing vehicle control accuracy.

[0067] In this embodiment, determining the foreground portion of an environmental image based on feature vectors includes: segmenting the foreground and background portions of the environmental image based on feature vectors and a preset foreground and background segmentation algorithm; adding a preset first identifier to the foreground portion and a preset second identifier to the background portion; and determining the foreground portion of the environmental image based on the first identifier.

[0068] Specifically, a foreground and background segmentation algorithm is pre-set, which can be used to segment the foreground and background parts of the environment image. In this embodiment, the foreground and background segmentation algorithm is not specifically limited.

[0069] Using a foreground and background segmentation algorithm, the foreground and background portions of the environment image are segmented based on the feature vectors of the environment image. A first identifier and a second identifier are pre-defined, for example, the first identifier is 0 and the second identifier is 1. The first identifier is added to the foreground portion, and the second identifier is added to the background portion. By obtaining the identifiers on the segmented image, the portion with the first identifier is determined to be the foreground portion.

[0070] The advantage of this setup is that by segmenting the foreground and background, objects in the environmental image can be accurately identified, avoiding object recognition errors and thus improving the accuracy of vehicle control.

[0071] In this embodiment, before inputting the environmental image into the preset neural network model, the method further includes: acquiring a pre-collected image to be trained, determining the actual objects in the image to be trained and the actual depth range of each actual object; inputting the image to be trained into the neural network model to be trained, and outputting the target objects in the image to be trained and the target depth ranges corresponding to each target object; and training the neural network model based on a preset ordered regression loss function according to the actual objects in the image to be trained, the actual depth ranges of each actual object, the target objects in the image to be trained, and the target depth ranges corresponding to each target object, to obtain the trained neural network model.

[0072] Specifically, the constructed neural network model is pre-trained. This model can be a DORN (Deep Ordinal Regression Network) model. The original DORN network model's structure involves the input image passing through a convolutional network. This convolutional network differs from other feature extraction backbone networks in that it removes some pooling layers, resulting in a denser feature extractor with a relatively large feature map size. The feature map then passes through five channels, including a full-image encoder, a convolution operation, and three dilated convolutions. The results from these five channels are concatenated to form a scene understanding modularity. Finally, a 1×1 convolution integrates the information from the five channels and is fed into the ordinal regression part for output. In this embodiment, a U-net network is used instead of the part before ordinal regression. That is, during training, the input environment image is directly processed by the U-net network to obtain feature vectors. Based on these feature vectors, the foreground portion of the environment image is determined, and objects within the foreground are identified as objects in the environment image. The depth range corresponding to the objects in the environment image is determined, and the results of identifying the objects and depth ranges are input into the ordinal regression part for training.

[0073] A large number of training images are pre-collected and labeled. The objects in the training images and their depth ranges are manually determined. These manually determined objects are then considered real objects, and their actual depth ranges are defined as the actual depth ranges. In other words, the actual objects in the training images and their actual depth ranges are pre-determined. The actual depth range is the depth range corresponding to the actual distance between the camera and the objects in the image at the time of image capture.

[0074] The training image is input into the neural network model, and the output is the target objects in the training image and the corresponding depth ranges of each target object. The target objects are the objects in the environment image calculated by the neural network model, and the target depth ranges are the depth ranges of the target objects calculated by the neural network model.

[0075] Backpropagation is performed based on a preset ordered regression loss function. The neural network model is trained using the actual objects in the training image, the actual depth ranges of each actual object, the target objects in the training image, and the target depth ranges corresponding to each target object, based on the preset ordered regression loss function. This process determines whether the actual objects and target objects are consistent, and whether the actual depth range of an object is consistent with its corresponding target depth range. If they are consistent, the neural network model training is considered complete; otherwise, training can continue.

[0076] The advantages of this setup are that it replaces the five channels and scene understanding module of the original DORN network model with the U-net network, reducing the computation process, resulting in more accurate calculations, faster computation speed, meeting the low computing power requirements of the vehicle-side computing platform, and improving the accuracy and efficiency of vehicle control.

[0077] S303. Determine the probability that the image depth of an object in the environmental image falls within each pre-divided depth interval; wherein, at least one depth interval is pre-divided within a preset distance in front of the vehicle.

[0078] For example, multiple depth intervals can be pre-divided, and the depth intervals can be divided in a discretization manner according to increasing distance. For instance, the closer to the vehicle, the denser the depth intervals; the farther away from the vehicle, the sparser the depth intervals. A range of 100 meters can be divided into 120 intervals. In this embodiment, the distance range and the number of depth intervals are not specifically limited.

[0079] Based on the trained neural network model, the probability of each object being located in each depth interval is determined. For example, if there are 120 depth intervals, then for a single object, 120 probabilities can be calculated. The calculated probabilities can be expressed as the likelihood of the object being in each depth interval.

[0080] In this embodiment, determining the probability that the image depth of an object in the environmental image falls within a pre-divided depth interval includes: calculating the probability that the distance between each object and the vehicle falls within each depth interval based on the feature vector of the object in the environmental image.

[0081] Specifically, the neural network model extracts features from the environmental image to obtain feature vectors for each object in the image. These feature vectors can include vectors of features such as the object's size and pixel values. Based on these feature vectors, the probability of each object being located in each depth interval is calculated. For example, the larger the object, the greater the probability of it being located in a forward depth interval. The probability calculation formula can be pre-set and the model trained; however, in this embodiment, no specific limitations are placed on the probability calculation formula.

[0082] The advantage of this setup is that it determines the probability of an object being located in each depth range based on the image features of each object, making it easier to find the depth range where the object is located and improving the accuracy of vehicle control.

[0083] S304. Based on the probability of objects in the environment image falling within each pre-divided depth interval, determine the depth interval of the image depth corresponding to the objects in the environment image.

[0084] For example, after obtaining the probability of an object in each depth interval, the actual depth interval in which the object is located can be determined without calculating the object's specific image depth. If there are multiple objects in the environment image, the depth interval in which each object is located can be determined.

[0085] In this embodiment, the depth interval of the image depth corresponding to the object in the environment image is determined according to the probability of the object in the environment image in each pre-divided depth interval, including: sorting the probability of the object in the environment image in each pre-divided depth interval from largest to smallest; and determining the depth interval corresponding to the probability ranked first as the depth interval of the image depth corresponding to the object in the environment image.

[0086] Specifically, after obtaining the probabilities of each object in each depth interval, the probabilities of each object are sorted, either from smallest to largest or largest to smallest. That is, each object yields a sorting result. Based on the sorting results, the highest probability among all probabilities of each object is determined. For example, if multiple probabilities of an object are sorted from largest to smallest, the probability ranked first in the sorting result is the object's highest probability. The depth interval corresponding to the highest probability is then determined as the depth interval in which the object is located.

[0087] The advantage of this setup is that it allows us to find the true depth range of an object by comparing its size, thereby determining the distance range between the object and the vehicle. This facilitates vehicle control and ensures driving safety.

[0088] S305. Control the vehicle to drive based on the depth range of the image depth corresponding to the object in the environmental image.

[0089] For example, this step can refer to step S103 above, and will not be repeated here.

[0090] This application provides a vehicle control method based on image depth. It acquires environmental images of the area in front of the vehicle using an image acquisition device installed on the vehicle, identifies objects in the environmental images, and determines the depth range of the objects within the image, i.e., determines the distance range between the objects and the vehicle. Vehicle movement is controlled based on this distance range, improving driving safety. This method solves the problems of slow calculation speed or low calculation accuracy caused by the need to calculate specific distance values ​​in existing technologies. By determining the depth range, the accuracy of the calculation is improved, and the calculation speed is faster, meeting the low computing power requirements of the vehicle-side computing platform and improving the accuracy and efficiency of vehicle control.

[0091] Figure 4 This is a structural block diagram of a vehicle control device based on image depth, provided as an embodiment of this application. For ease of explanation, only the parts relevant to the embodiments of this disclosure are shown. This device is applied to a vehicle, see reference... Figure 4 The device includes: an image acquisition module 401, a range determination module 402, and a vehicle control module 403.

[0092] Image acquisition module 401 is used to acquire environmental images within a preset space range in front of the vehicle;

[0093] The interval determination module 402 is used to identify objects in the environmental image and determine the depth interval of the image depth corresponding to the objects in the environmental image; wherein, the depth interval is used to represent the distance range between the object and the vehicle;

[0094] The vehicle control module 403 is used to control the vehicle to drive based on the depth range of the image depth corresponding to the object in the environmental image.

[0095] Figure 5 This application provides a structural block diagram of a vehicle control device based on image depth, in which... Figure 4 Based on the illustrated embodiments, as Figure 5 As shown, the interval determination module 402 includes an object recognition unit 4021, a probability determination unit 4022, and a depth interval determination unit 4023.

[0096] The object recognition unit 4021 is used to input the environmental image into a preset neural network model to obtain the objects in the environmental image;

[0097] The probability determination unit 4022 is used to determine the probability that the image depth of an object in the environmental image falls within each pre-divided depth interval; wherein, at least one depth interval is pre-divided within a preset distance in front of the vehicle.

[0098] The depth interval determination unit 4023 is used to determine the depth interval of the image depth corresponding to the object in the environmental image based on the probability of the object in the environmental image being in each pre-divided depth interval.

[0099] In one example, the object recognition unit 4021 includes:

[0100] An image input subunit is used to input the environmental image into a U-shaped network in a preset neural network model;

[0101] The feature extraction subunit is used to extract features from the environmental image based on the U-shaped network to obtain the feature vector of the environmental image;

[0102] The foreground determination subunit is used to determine the foreground portion of the environment image based on the feature vector, identify objects in the foreground portion, and identify them as objects in the environment image.

[0103] In one example, the foreground determining sub-unit is specifically used for:

[0104] Based on the feature vector, and using a preset foreground and background segmentation algorithm, the foreground and background portions of the environmental image are segmented.

[0105] A preset first identifier is added to the foreground portion, and a preset second identifier is added to the background portion. Based on the first identifier, the foreground portion of the environmental image is determined.

[0106] In one example, probability determination unit 4022 is specifically used for:

[0107] Based on the feature vectors of objects in the environmental image, calculate the probability that the distance between each object and the vehicle lies within each depth interval.

[0108] In one example, the depth range determination unit 4023 is specifically used for:

[0109] The probabilities of objects in the environmental image within each pre-divided depth interval are sorted from largest to smallest.

[0110] The depth range corresponding to the probability that ranks first is determined as the depth range of the image depth corresponding to the object in the environmental image.

[0111] In one example, the device also includes:

[0112] The model training module is used to acquire pre-collected images to be trained before inputting the environmental images into a preset neural network model, and to determine the actual objects in the images to be trained and the actual depth range of each actual object.

[0113] The image to be trained is input into the neural network model to be trained, and the target objects in the image to be trained and the target depth ranges corresponding to each target object are output.

[0114] Based on the actual objects in the image to be trained, the actual depth range of each actual object, the target objects in the image to be trained, and the target depth range corresponding to each target object, the neural network model is trained using a preset ordered regression loss function to obtain a trained neural network model.

[0115] In one example, vehicle control module 403 is specifically used for:

[0116] Based on the depth range of the image depth corresponding to the object in the environmental image, and based on the association relationship between the preset depth range and the control rule, the control rule associated with the depth range of the image depth corresponding to the object in the environmental image is determined, and this rule is the current control rule.

[0117] The vehicle is controlled to move according to the current control rules.

[0118] Figure 6 A structural block diagram of an electronic device provided in an embodiment of this application, such as... Figure 6 As shown, the electronic device includes: a memory 61 and a processor 62; the memory 61 is a memory used to store instructions executable by the processor 62.

[0119] The processor 62 is configured to perform the methods provided in the above embodiments.

[0120] The electronic device also includes a receiver 63 and a transmitter 64. The receiver 63 is used to receive instructions and data sent by other devices, and the transmitter 64 is used to send instructions and data to external devices.

[0121] Figure 7 This is a structural block diagram of an electronic device according to an exemplary embodiment. The device may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0122] Device 700 may include one or more of the following components: processing component 702, memory 704, power supply component 706, multimedia component 708, audio component 710, input / output (I / O) interface 712, sensor component 714, and communication component 716.

[0123] Processing component 702 typically controls the overall operation of device 700, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 702 may include one or more processors 720 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 702 may include one or more modules to facilitate interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module to facilitate interaction between multimedia component 708 and processing component 702.

[0124] Memory 704 is configured to store various types of data to support the operation of device 700. Examples of this data include instructions for any application or method operating on device 700, contact data, phonebook data, messages, pictures, videos, etc. Memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0125] Power supply component 706 provides power to various components of device 700. Power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 700.

[0126] Multimedia component 708 includes a screen that provides an output interface between the device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When the device 700 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0127] Audio component 710 is configured to output and / or input audio signals. For example, audio component 710 includes a microphone (MIC) configured to receive external audio signals when device 700 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 704 or transmitted via communication component 716. In some embodiments, audio component 710 also includes a speaker for outputting audio signals.

[0128] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0129] Sensor assembly 714 includes one or more sensors for providing state assessments of various aspects of device 700. For example, sensor assembly 714 may detect the on / off state of device 700, the relative positioning of components such as the display and keypad of device 700, changes in the position of device 700 or a component of device 700, the presence or absence of user contact with device 700, the orientation or acceleration / deceleration of device 700, and temperature changes of device 700. Sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 714 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 714 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0130] Communication component 716 is configured to facilitate wired or wireless communication between device 700 and other devices. Device 700 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 716 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0131] In an exemplary embodiment, device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0132] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, which can be executed by a processor 720 of device 700 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0133] A non-transitory computer-readable storage medium, wherein instructions in the storage medium, when executed by a processor of a terminal device, enable the terminal device to perform the aforementioned image depth-based vehicle control method of the terminal device.

[0134] This application also discloses a computer program product, including a computer program that, when executed by a processor, implements the method described in this embodiment.

[0135] Various embodiments of the systems and technologies described above in this application can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0136] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or electronic device.

[0137] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0138] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0139] The systems and technologies described herein can be implemented in computing systems that include back-end components (e.g., as data electronic devices), or computing systems that include middleware components (e.g., application electronic devices), or computing systems that include front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0140] Computer systems can include client and electronic devices. Clients and electronic devices are generally geographically separated and typically interact via communication networks. The client-electronic device relationship is created by computer programs running on the respective computers and having a client-electronic device relationship with each other. The electronic device can be a cloud electronic device, also known as a cloud computing electronic device or cloud host, a host product within the cloud computing service system, addressing the shortcomings of traditional physical hosts and VPS services ("Virtual Private Server," or simply "VPS") in terms of management difficulty and weak business scalability. The electronic device can also be an electronic device in a distributed system or an electronic device incorporating blockchain technology. It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application is achieved, and this is not limited herein.

[0141] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0142] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A vehicle control method based on image depth, characterized in that, The method is applied to a vehicle, and the method includes: Acquire environmental images within a preset space area in front of the vehicle; The environmental image is input into a preset neural network model to obtain objects in the environmental image; the probability that the image depth of the objects in the environmental image falls within each pre-divided depth interval is determined; wherein, at least one depth interval is pre-divided within a preset distance in front of the vehicle; based on the probability that the objects in the environmental image fall within each pre-divided depth interval, the depth interval corresponding to the image depth of the objects in the environmental image is determined; wherein, the depth interval is used to represent the distance range between the object and the vehicle; The vehicle is controlled to move based on the depth range of the image depth corresponding to the objects in the environmental image.

2. The method according to claim 1, characterized in that, The environmental image is input into a preset neural network model to obtain the objects in the environmental image, including: The environmental image is input into a U-shaped network in a preset neural network model; The environmental image is feature extracted using the U-shaped network to obtain the feature vector of the environmental image; Based on the feature vector, the foreground portion of the environmental image is determined, and objects in the foreground portion are identified as objects in the environmental image.

3. The method according to claim 2, characterized in that, Determining the foreground portion of the environment image based on the feature vector includes: Based on the feature vector, and using a preset foreground and background segmentation algorithm, the foreground and background portions of the environmental image are segmented. A preset first identifier is added to the foreground portion, and a preset second identifier is added to the background portion. Based on the first identifier, the foreground portion of the environmental image is determined.

4. The method according to claim 2, characterized in that, Determining the probability that the image depth of an object in the environmental image falls within each pre-divided depth interval includes: Based on the feature vectors of objects in the environmental image, calculate the probability that the distance between each object and the vehicle lies within each depth interval.

5. The method according to claim 1, characterized in that, Based on the probability of objects in the environmental image falling within pre-defined depth intervals, the depth interval corresponding to the image depth of the objects in the environmental image is determined, including: The probabilities of objects in the environmental image within each pre-divided depth interval are sorted from largest to smallest. The depth range corresponding to the probability that ranks first is determined as the depth range of the image depth corresponding to the object in the environmental image.

6. The method according to claim 1, characterized in that, Before inputting the environmental image into the preset neural network model, the method further includes: Acquire pre-collected training images and determine the actual objects in the training images and the actual depth range of each actual object; The image to be trained is input into the neural network model to be trained, and the target objects in the image to be trained and the target depth ranges corresponding to each target object are output. Based on the actual objects in the image to be trained, the actual depth range of each actual object, the target objects in the image to be trained, and the target depth range corresponding to each target object, the neural network model is trained using a preset ordered regression loss function to obtain a trained neural network model.

7. The method according to any one of claims 1-6, characterized in that, Controlling the vehicle to move based on the depth range of the image depth corresponding to objects in the environmental image includes: Based on the depth range of the image depth corresponding to the object in the environmental image, and based on the association relationship between the preset depth range and the control rule, the control rule associated with the depth range of the image depth corresponding to the object in the environmental image is determined, and this rule is the current control rule. The vehicle is controlled to move according to the current control rules.

8. A vehicle control device based on image depth, characterized in that, The device is applied to a vehicle, and the device includes: The image acquisition module is used to acquire environmental images within a preset spatial range in front of the vehicle; An interval determination module is used to identify objects in the environmental image and determine the depth interval of the image depth corresponding to the objects in the environmental image; wherein, the depth interval is used to represent the distance range between the object and the vehicle; The vehicle control module is used to control the vehicle to move based on the depth range of the image depth corresponding to the objects in the environmental image; The interval determination module includes an object recognition unit, a probability determination unit, and a depth interval determination unit; The object recognition unit is used to input the environmental image into a preset neural network model to obtain the objects in the environmental image; The probability determination unit is used to determine the probability that the image depth of an object in the environmental image falls within each pre-divided depth interval; wherein, at least one depth interval is pre-divided within a preset distance in front of the vehicle. The depth interval determination unit is used to determine the depth interval of the image depth corresponding to the object in the environmental image based on the probability of the object in each pre-divided depth interval.

9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the image depth-based vehicle control method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the image depth-based vehicle control method as described in any one of claims 1-7.

11. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the image depth-based vehicle control method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Visual overlay for providing depth perception

    CN114821496A