Method, device, equipment, vehicle and medium for determining object in road

Through the method based on semantic processing and distance calculation in the autonomous driving system, the accuracy of road object recognition in light changes and dynamic environments is solved, and a more robust object matching and recognition effect is achieved.

CN120236259APending Publication Date: 2025-07-01ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311856347.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In autonomous driving and driving assistance systems, changes in light conditions and dynamic environments make it difficult for existing image recognition methods to accurately segment and identify road objects, especially in complex parking scenarios, where there is a risk of matching failure.

Method used

By determining the object of interest in the road image based on semantic processing and calculating the distance between the pixel and other pixels, the distance within a predetermined threshold range is used to include the pixels in the object of interest, and combining normalization and nonlinear optimization techniques to improve the accuracy of object recognition.

Benefits of technology

Accurate identification of road objects in dynamic light and complex environments, providing more robust and accurate matching results, reducing error and noise impact.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236259A_ABST
    Figure CN120236259A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for determining an object in a road image, equipment, a vehicle and a medium. In the method, a controller in the vehicle determines an object of interest in the road image based on semantic processing. A controller in a vehicle determines a distance between a pixel in an object of interest and other pixels. The controller also determines whether the determined distance is within a range of a predetermined threshold. When the distance is within a predetermined threshold range, the controller may include other pixels in the object of interest. Through the method implemented by the invention, the interested object in the image can be identified or determined more accurately, and a more robust and accurate matching result is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to the field of data processing, and more particularly to methods, devices, equipment, vehicles, and media for determining objects in a road. Background Art

[0002] With the rapid development of artificial intelligence, the research on autonomous driving and driver assistance systems has attracted extensive attention in the industry. Technologies such as automatic parking assistance and valet parking have become increasingly important for drivers. Automatic parking assistance systems are dedicated to helping drivers complete the parking operation of vehicles to address the challenges of parking in narrow spaces or complex environments. Such systems use sensors, cameras, and computer vision technologies to detect the surrounding environment and achieve precise and safe parking by automatically controlling the steering, acceleration, and braking of the vehicle.

[0003] In the parking lot scenarios for autonomous driving and driver assistance systems, due to factors such as light conditions, the dynamic changes in the vehicle placement may be relatively large. As matching features, it is necessary to address the issues of feature changes and associations in the time domain, and there is a risk of matching failure under different conditions. Road semantic features (such as arrows, lane lines, parking space lines, etc.) are features that remain stable over a long time and can provide robust and stable matching results. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, device, equipment, vehicle, and medium for determining objects in a road.

[0005] According to a first aspect of the present disclosure, there is provided a method for determining an object in a road, the method including determining an object of interest in a road image based on semantic processing. The method further includes determining the distance between pixels in the object of interest and other pixels, and including the other pixels in the object of interest in response to the distance being within a predetermined threshold range.

[0006] According to a second aspect of the present disclosure, there is provided a device for determining an object in a road, the device including a first determination unit configured to determine an object of interest in a road image based on semantic processing. The device further includes a second determination unit configured to determine the distance between pixels in the object of interest and other pixels, and an inclusion unit configured to include the other pixels in the object of interest in response to the distance being within a predetermined threshold range.

[0007] According to a third aspect of the present disclosure, there is provided a controller. The controller includes at least one processor; and a memory coupled to the at least one processor and having instructions stored thereon, the instructions, when executed by the at least one processor, cause the controller to perform the steps of the method in the first aspect of the present disclosure.

[0008] According to a fourth aspect of the present disclosure, a vehicle is provided, which includes the controller in the third aspect of the present disclosure.

[0009] According to a fifth aspect of the present disclosure, a machine-readable storage medium is provided. Computer-executable instructions are stored on the machine-readable storage medium, and the computer-executable instructions are executed by a processor to implement the steps of the method in the first aspect of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent. In the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.

[0011] Figure 1A A schematic diagram illustrating an example environment in which an apparatus and / or method according to an embodiment of the present disclosure may be implemented;

[0012] Figure 1B A schematic diagram illustrating another example environment in which an apparatus and / or method according to an embodiment of the present disclosure may be implemented;

[0013] Figure 2 A flowchart illustrating a method for determining an object in a road according to an embodiment of the present disclosure;

[0014] Figure 3 A flowchart of a method for determining an object in a road image according to an embodiment of the present disclosure;

[0015] Figure 4 A flowchart illustrating a method for performing distance transformation and normalization on an object identified in a road image according to an embodiment of the present disclosure;

[0016] Figure 5 An effect diagram illustrating that a pose optimization module adjusts a pose based on a cost image according to an embodiment of the present disclosure;

[0017] Figure 6 A comparison diagram illustrating the results of a method implemented according to an embodiment of the present disclosure and a previous method according to an embodiment of the present disclosure;

[0018] Figure 7 A schematic diagram illustrating an apparatus for determining an object in a road image according to an embodiment of the present disclosure; and

[0019] Figure 8 A schematic block diagram illustrating an example device that may be used to implement an embodiment of the present disclosure. DETAILED DESCRIPTION

[0020] The embodiments of the present disclosure described below with reference to the accompanying drawings are only for exemplary purposes and are not intended to limit the protection scope of the present disclosure. Additionally, before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users and user authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0021] In current image recognition, methods based on semantic segmentation or semantic vectorization usually cannot accurately divide the objects in the image due to factors such as large camera glare, reflection, and changing light conditions, resulting in significant errors. In a dynamic environment with changing light, the background and foreground objects may change frequently due to factors such as uneven illumination, shadows and reflections, and changing light intensity, making it difficult for current recognition methods to accurately segment and identify the objects in front of the vehicle. In the scenario of a vehicle moving forward, motion blur may also be caused by the rapid movement of the vehicle or surrounding objects, making it difficult for current methods to accurately capture the boundaries and features of the objects. In addition, different shooting angles may also cause significant changes in the shape of the objects, increasing the complexity of recognition.

[0022] To at least solve the above and other potential problems, embodiments of the present disclosure provide a method for determining objects in a road image. In this method, a controller in the vehicle determines the objects of interest in the road image based on semantic processing. The controller in the vehicle determines the distance between the pixels in the object of interest and other pixels. The controller also determines whether the determined distance is within a predetermined threshold range. When the distance is within the predetermined threshold range, the controller may include the other pixels in the object of interest. Through the method implemented by the present disclosure, the objects of interest in the image can be more accurately identified or determined, providing a more robust and accurate matching result.

[0023] The embodiments of the present disclosure will be described in further detail below with reference to the accompanying drawings, where Figure 1A FIG. 100A shows a schematic diagram of an example environment in which the device and / or method according to an embodiment of the present disclosure may be implemented.

[0024] As Figure 1AAs shown, vehicle 101 equipped with a controller implemented according to an embodiment of the present disclosure can use its own camera to collect images related to road 106, such as parking space information, sidewalks, traffic light signals, etc. Subsequently, the controller can identify these images to generate an image 102 identified based on semantic segmentation and an image 104 identified based on semantic vectorization. According to an embodiment of the present disclosure, the image 102 identified based on semantic segmentation can be a pixel-level semantic detection result. Semantic segmentation can classify each pixel and assign it to a specific category (such as road, vehicle, pedestrian, etc.). There may be no association between pixels. Pixels can be assigned category labels to create an identified image, where different colors or gray levels are typically used to represent different categories. The machine learning model can learn the complex mapping from the original pixels to the category labels. The segmented image usually represents different categories with different colors or gray levels, making it possible to intuitively understand the semantic content of the image.

[0025] In some embodiments, the image 104 identified based on semantic vectorization can be a target-level vectorized expression of the semantic detection result. According to an embodiment of the present disclosure, the controller can convert specific targets (such as objects, people, scene elements, etc.) in the image into a vectorized representation form to capture and represent the key semantic features of these targets. It focuses on the overall targets or objects in the image. It identifies the key entities in the image and converts each entity into a vector representation form, usually a feature vector. The vector contains key information describing the target, such as features like shape, size, position, texture, etc., thus providing a deep semantic understanding of the image content.

[0026] Figure 1B FIG. 100B shows a schematic diagram of another example environment in which an apparatus and / or method according to an embodiment of the present disclosure can be implemented. As Figure 1B shown, vehicle 101 equipped with a controller implemented according to an embodiment of the present disclosure can use its own camera to collect road-related images, thereby generating a (plural) road object map 103 including elements such as road markings, curbs, traffic lights, etc. and a self-vehicle road boundary map 105 including information such as lane lines.

[0027] The distance transformation module 107 in the controller can perform distance transformation on the collected road object map 103 and the self-vehicle road boundary map 105, for example, to determine the minimum distance from each point in the image to the nearest obstacle or feature. In some embodiments, the distance transformation calculates the distance from each pixel in the image to the nearest object pixel. Different metrics can be used to calculate this distance, such as methods like Euclidean distance, Manhattan distance, or chessboard distance. This method helps the navigation system of vehicle 101 understand the surrounding space and make decisions regarding safe and efficient paths.

[0028] According to an embodiment of the present disclosure, the distance transformation module 107 may then generate a distance-transformed (plural) image 109, where the darker regions in the sub-images 109-1 and 109-2 in the (plural) image 109 indicate a smaller distance from the vehicle 101, and the lighter regions indicate a larger distance from the vehicle 101.

[0029] As described above in connection with Figure 1A and Figure 1B is a schematic diagram for determining an object in a road image in which an embodiment of the present disclosure can be implemented. The following describes a flowchart of a method 200 for determining an object in a road image according to an embodiment of the present disclosure in connection with Figure 2 The method 200 may be executed at a controller of the vehicle 101 in Figure 1A and any suitable computing device.

[0030] Figure 2 FIG. 200 is a flowchart of a method for determining an object in a road image according to an embodiment of the present disclosure. As Figure 2 shown, at block 202, an object of interest in the road image is determined based on semantic processing. According to an embodiment of the present disclosure, a controller in the vehicle may first perform semantic processing on an image acquired by a camera to identify one or more regions of interest in the image. According to an embodiment of the present disclosure, the method of semantic processing may include one or more of semantic segmentation or semantic vectorization processing, where semantic segmentation focuses on semantic detection at the pixel level, and semantic vectorization processing focuses on semantic detection of vectorized expressions at the target level. Therefore, depending on the semantic processing method adopted, the segmented images may also be different.

[0031] In the context of the present disclosure, a region of interest may be a location, area, object, pedestrian, parking space, etc. that needs to be particularly noted or processed during vehicle travel. For example, when the vehicle is to park, the region of interest may be the parking space where it is to dock, and the vehicle needs to identify and locate the parking space to ensure safe and accurate parking. When the vehicle is in motion, the regions of interest may be zebra crossings, traffic lights, signs, road markings, sidewalks, etc. The vehicle needs to be able to identify these elements and adjust its driving behavior accordingly.

[0032] At block 204, the distances between the pixels in the object of interest and other pixels are determined. According to an embodiment of the present disclosure, the vehicle's controller can determine the distances between each other by calculating the distances between the pixels in the identified object of interest and other pixels at any position in the image, such as by Manhattan distance transform, Euclidean distance, or chessboard distance, etc. For example, for each pixel point in the image, the controller can estimate the actual distance between the two by calculating the minimum straight-line distance from it to any pixel in the object of interest.

[0033] At block 206, in response to the distance being within a predetermined threshold range, the other pixels are included in the object of interest. According to an embodiment of the present disclosure, when the distance determined at block 204 is within the predetermined threshold range, the vehicle's controller can include this pixel point in the object of interest. For example, when the vehicle is parking, when the distance between a determined pixel point in the image and the identified parking space is less than the predetermined threshold, the pixel point can be considered as part of the identified parking space. The vehicle can avoid this pixel point when parking.

[0034] In this way, the vehicle's controller can achieve an accurate division of the region of interest. When a pixel point falls within the predetermined threshold, it can be considered as part of the region of interest, and the pixel points outside the predetermined threshold are "truncated" and not concerned. According to an embodiment of the present disclosure, the predetermined threshold can be set based on prior knowledge and test results. In some embodiments, the predetermined threshold can also be set based on the type of the object of interest, the characteristics of the image, the requirements of the application, and the required processing accuracy.

[0035] In some embodiments, the vehicle's controller can also perform normalization processing on the object of interest to generate normalized data. The normalized data can be further subjected to non-linear optimization, such as one or more of the gradient descent method, Newton's method, and quasi-Newton method to more accurately determine the position and angle of the object of interest.

[0036] Additionally or alternatively, in some embodiments, the pose information of the object of interest determined by the vehicle's controller (e.g., the position and angle of the region of interest, etc.) can also be compared with the prior pose information including prior position and angle information. The prior pose information can come from a high-precision map, GPS data, or information provided by other vehicles and infrastructure. The vehicle's controller can also be retrained or adjusted based on the comparison differences, such as adjusting the parameters of the controller, experimental data, etc. For example, the controller can apply the gradient descent method to linearly optimize the cost image obtained after normalization and non-linear transformation to be in a manner with the smallest difference from the prior position and angle, where the gradient is the derivative between the two and descends fastest along the derivative direction. This process can be iterated until the gradient approaches zero. For example, in some embodiments, non-linear optimization can be performed by adjusting the position and angle starting from the prior position and angle such that the sum of the distance values of the map points projected onto the cost image is minimized.

[0037] Figure 3 FIG. 300 is a flowchart illustrating a method for determining an object in a road image according to an embodiment of the present disclosure. As Figure 3 shown, the road image data collected by the camera 302 can be fed into the perception module 304 in the controller implemented according to the embodiments of the present disclosure. The perception module 304 can preprocess the collected road image data. For example, it can clean and standardize the image data, such as resizing, cropping, or normalizing pixel values. In some embodiments, the perception module 304 can be a deep neural network model based on one or more of a convolutional neural network (CNN) for image recognition and processing, a recurrent neural network (RNN) for processing temporal data, a semantic segmentation network for distinguishing different regions and objects in an image, and an object detection and classification network.

[0038] In some embodiments, the perception module 304 can extract key features of the image. For example, the perception module 304 can perform semantic processing on the road image features based on semantic processing. In some embodiments, the perception module 304 can perform semantic segmentation on the road image. For example, the perception module 304 can identify features in the image, such as edges, shapes, and textures. The perception module 304 can then extract deep features of the image through multi-layer convolution and pooling operations. In some embodiments, the perception module 304 can also control the image resolution and edge details using upsampling and skip connections during the decoding stage of the image.

[0039] According to an embodiment of the present disclosure, the perception module 304 may also perform semantic vectorization on the collected road images. For example, the perception module 304 may extract the features of the collected road images or perform object detection. Subsequently, the perception module 304 may convert the extracted features into a vector form. For example, the perception module 304 may identify key feature points in the image, such as corner points and edges.

[0040] In some embodiments, the perception module 304 may also convert these descriptors into a vector form, which may be a fixed-length vector for calculation and comparison. In some embodiments, the perception module 304 may also perform spatial distribution encoding, such as preserving spatial relationships during the vectorization process to maintain an understanding of the image structure. The perception module 304 may classify the vectorized data. For example, the vectorized data may be mapped to specific semantic categories, such as road boundaries, pedestrians, lane markings, etc. Additionally or alternatively, in some embodiments, the perception module 304 may also output the identified semantic objects based on the identified features.

[0041] For example, the perception module 304 may also determine the classification of the identified semantic objects based on the probability that each pixel belongs to different categories, such as roads, pedestrians, vehicles, etc. In some implementations, the perception module 304 may also perform processing such as smoothing filtering and boundary refinement on the input image to improve the accuracy of segmentation.

[0042] In some embodiments, the perception module 304 may subsequently feed the identified semantic objects into the cost image generation module 306. The cost image generation module 306 may perform a distance transformation on the identified semantic objects to generate a cost image. The generation process of the cost image will be described in detail below with reference to Figure 4 for a specific description. The cost image may be used to determine the distance between different semantic objects.

[0043] For example, in some embodiments, the cost image generation module 306 uses a distance transformation algorithm to calculate the distance from each pixel point in the semantic objects in the image to the nearest semantic object boundary by calculating the Euclidean distance, Manhattan distance, or chessboard distance between (multiple) semantic objects. In some embodiments, the cost image generation module 306 may assign a cost value to each pixel point based on the calculated distance, thereby forming a cost matrix, which may reflect the distance from any point to the nearest semantic object. Additionally or alternatively, the cost image generation module 306 may convert the cost matrix into a cost image, where different color or brightness levels represent different cost values, thereby showing the distance between different semantic objects.

[0044] In some embodiments, the map service module 308 in a vehicle may provide a semantic map based on the map information therein. The semantic map may include not only the geographical information of a normal map but also integrated additional semantic data such as road types, traffic signs, the status and height of traffic lights, pedestrian walkway widths, bike lanes, and so on. The semantic map can help the vehicle identify the surrounding environment faster and more accurately, thus enabling effective parsing and response to complex traffic environments.

[0045] The semantic map and the cost image can then be fed into the pose optimization module 310 of the vehicle to adjust the position or angle of one or more of the identified elements such as traffic signs, traffic lights, parking spaces, etc. In some embodiments, the pose optimization module 310 can determine the position and orientation of the vehicle itself. The pose optimization module 310 can then perform an environmental analysis to determine and analyze the data in the semantic map and the cost image, thereby identifying the current positions and orientations of the respective elements.

[0046] In some embodiments, the pose optimization module 310 can also compare the prior positions and angles of these elements with the expected generated positions and angles in the pose module, and perform correction by calculating the deviation. In some embodiments, the pose optimization module 310 can also make adjustments based on the deviation, so that the position and angle information generated by the next prediction is close to or the same as the prior positions and angles. Additionally or alternatively, according to embodiments of the present disclosure, the controller for determining an object in a road image may include one or more of the perception module 304, the cost image generation module 306, and the pose optimization module 310, and the controller can utilize these modules to perform map matching 312.

[0047] In some embodiments, the pose optimization module 310 can perform adjustment calculations based on the correction result or the deviation to determine the preferred positions or orientations of one or more elements such as traffic signs, traffic lights, parking spaces, etc. Finally, the pose optimization module 310 can output the adjusted position or rotation angle information to guide the vehicle's navigation and decision-making system.

[0048] Figure 4 FIG. 400 is a flowchart illustrating a method for distance transformation and normalization of identified objects in a road image according to an embodiment of the present disclosure. As Figure 4 shown, an image 402 with semantic category detection information from the perception module, such as lane lines, arrows, zebra crossings, etc., can be fed into the cost image generation module 404.

[0049] The image 402 can be generated by the perception module into binary images respectively for each semantic category of interest. As an example, in the lane line layer, only the pixel values of the lane lines are 0, and other objects can be regarded as background pixels with a value of 255. In some embodiments, in the layer for pedestrian detection, the pixel values of the pedestrians of interest can be set to 0, and all other pixels are set to 255. In the layer for traffic sign recognition, the pixel values of the pedestrians of interest can be set to 0, and all other pixels are set to 255, and so on. Additionally or alternatively, in the example of consecutive video frames, binary images can be created to track moving objects (such as walking pedestrians or moving vehicles), where the pixel values of the moving objects are 0 and the other parts are 255. This can make specific categories stand out in the image for further analysis and processing.

[0050] According to an embodiment of the present disclosure, the cost image generation module 404 can be based on a distance transformation kernel. This distance transformation kernel can be used to calculate the distance from each pixel point to the nearest specific feature. This distance is usually the distance to the nearest non-zero pixel point (in the binary image). According to an embodiment of the present disclosure, various different distance measurement methods can be adopted, such as Euclidean distance, Manhattan distance, or Chebyshev distance, etc. As an example, during the process of performing distance transformation on a binary image, for pixel point a, the distance to the nearest semantic pixel b is the distance value of pixel a.

[0051] According to an embodiment of the present disclosure, the cost image generation module 404 can also process the image based on a truncated threshold th_truncated. For example, for a pixel point a in the object of interest, when the distance between another pixel point b and it is within the range of the truncated threshold th_truncated, then pixel point b can be considered to be within the range of the object of interest. Thus, when performing image segmentation processing, pixel point b is regarded as part of the object of interest. For example, in the scenario of determining a parking space, the pixel points within the range of the truncated threshold th_truncated from the parking space can be regarded as part of the parking space.

[0052] Additionally or alternatively, in an autonomous driving system, a truncated threshold can be set to determine whether various elements on the road are relevant to the vehicle. For example, for a parked vehicle by the roadside, if the distance between a pixel point (such as a part of the vehicle) and the pixel point a of the expected driving trajectory of the autonomous vehicle is within the range of the truncated threshold, this pixel point b of the parked vehicle can be regarded as an object of interest associated with the driving trajectory, so that these parked vehicles can be considered during path planning.

[0053] According to embodiments of the present disclosure, the truncation threshold th_truncated can be determined according to the application scenario and specific requirements. In some embodiments, it can be determined based on testing, collecting data, and prior knowledge in an actual or simulated environment. In some embodiments, the truncation threshold th_truncated can also be dynamically adjusted according to different environmental variables, such as factors like lighting and weather. In some embodiments, the truncation threshold th_truncated can be based on the purpose of use, for example, whether it is used for obstacle detection, lane line recognition, etc.

[0054] In some embodiments, for a pixel point a in the object of interest, when the distance between another pixel point b and it is outside the range of the truncation threshold th_truncated, then pixel point b can be considered outside the range of the object of interest. Thus, when performing image segmentation processing, pixel point b is not regarded as part of the object of interest. For example, in an urban scene, if the distance of pixel point b exceeds the set truncation threshold th_truncated (such as for identifying lane lines), then this point will not be regarded as part of the lane line.

[0055] According to embodiments of the present disclosure, the cost image generation module 404 can process multiple objects of interest simultaneously, such as lane lines, crosswalks, traffic signs, etc., and each object has a corresponding truncation threshold to ensure accurate identification and segmentation of different types of road elements. In this way, corresponding segmentation decisions can be made for complex different scenarios.

[0056] In this way, the cost image generation module 404 compares the value obtained by distance transformation with the truncation threshold. If the distance is less than or equal to the threshold, then this pixel point is regarded as part of the object of interest; if the distance is greater than the threshold, it is regarded as a non - interested region. The cost image generation module 404 can then assign a cost value or a distance value to each pixel point based on these judgments. For example, in some embodiments, pixel points within the object of interest can be assigned lower cost values, while pixel points in non - interested regions can be assigned higher cost values, thereby showing the distance between different semantic objects and finally generating the cost image 406.

[0057] In some embodiments, the distance values from 0 to the truncation threshold th_truncated can be linearly normalized to floating - point numbers from 0 to 1 to form the cost image 406. For example, during the normalization process, pixel points closer to the object of interest will be closer to 0, while pixel points far from the object of interest will have values closer to 1. By this method, the cost image 406 can finely represent the relationship between each region in the image and the object of interest with a continuous range of floating - point numbers, thereby ensuring that all data is on the same scale, facilitating comparison and processing, converging faster, and reducing errors and instabilities during numerical calculation processes.

[0058] According to embodiments of the present disclosure, truncation can reduce the change of the distance field caused by noise such as false detection, and normalization removes the dependence on the neighborhood range and image resolution. In some examples, after noise pixels are generated, they will affect the surrounding distance values. However, after truncation and normalization, the noise outside the truncated neighborhood range has no effect on the distance field around the target semantics.

[0059] The generated cost image 406 can be fed into the pose optimization module 408. The pose optimization module 408 can further adjust one or more of factors such as the position and rotation angle of the predicted object of interest based on the cost image 406, and finally output the adjusted pose 410.

[0060] Figure 5 Describe the effect diagram 500 of the pose optimization module based on the cost image to adjust the pose according to the embodiments of the present disclosure. The pose optimization module 408 can increase the numerical transformation of the nonlinear model for the cost values within the neighborhood of the semantic elements in the cost image 408 to satisfy the consistency of the maximum and minimum values while the gradient change is more obvious. For example, in some embodiments, the nonlinear transformation can be f(x) = 1 - (1 - x) 2 . The pose optimization module 408 can then adjust the pose of the object of interest based on the updated distance values. According to embodiments of the present disclosure, the nonlinear optimization can be one or more of the gradient descent method, Newton's method, and quasi-Newton method.

[0061] This process can be represented by the following formula:

[0062] CostMat new = f nonlinear (CostMat old ) (1)

[0063] where CostMat new 502 represents the nonlinear optimization implemented according to the present disclosure, and CostMat old 504 represents the transformation implemented according to the general method. It can also be determined from the effect diagram 500 that, compared with CostMat old , by applying the nonlinear transformation, the change of the gradient of CostMat new 502 can be increased, especially the gradient near the optimal solution. For example, in some embodiments, the descent value of the gradient can be increased, that is, the difference between the minimum value and the sub-minimum value can be increased, so as to converge to a more accurate optimal solution.

[0064] Figure 6Figure 600 compares the results of a method implemented according to an embodiment of the present disclosure with those of a prior method. Among them, box 602 represents the result diagram of identifying parking spaces according to the prior method, where the black area represents parking spaces. Box 604 represents the result diagram of identifying parking spaces using a cost image according to the method implemented according to the present disclosure, where the black area represents parking spaces. Comparing box 602 and 604 can determine that more accurate and clear identification of parking spaces can be achieved according to the present disclosure, and the identified errors are reduced.

[0065] Boxes 606, 608, and 610 represent the residual distributions of the non-linear optimization model implemented according to the prior method in different pose directions. In boxes 606, 608, and 610, it can be found that there are many minima in the residual distribution, which are prone to converge to the minima rather than the minimum value. Near the minimum value (optimal solution), it is approximately elliptical and not conducive to fast convergence. Boxes 612, 614, and 616 represent the residual distributions of the non-linear optimization model implemented according to the present disclosure in different pose directions. It can be found that the residual distribution reduces the minima and is prone to converge to the minimum value. Near the minimum value (optimal solution), it is approximately circular, which is conducive to fast convergence.

[0066] Figure 7 The figure shows a schematic diagram of a device for determining an object in a road image according to an embodiment of the present disclosure. The device 700 can be applied to a vehicle 101, which can include multiple modules for performing corresponding steps in process 200 as Figure 2 discussed. As Figure 7 shown, the device 700 can include: a first determination unit 702 configured to determine an object of interest in the road image based on semantic processing; a second determination unit 704 configured to determine the distance between pixels in the object of interest and other pixels; and an inclusion unit 706 configured to include other pixels in the object of interest in response to the distance being within a predetermined threshold range.

[0067] In some embodiments, the device 700 further includes: a normalization unit configured to perform normalization processing on the object of interest to generate normalized data. In some embodiments, the device 700 further includes: a non-linear transformation unit configured to perform non-linear transformation on the normalized data to determine the position and angle of the object of interest.

[0068] In some embodiments, the device 700 further includes: a comparison unit configured to compare one or more of the determined position and angle of the object of interest with one or more of the prior position and angle; and an adjustment unit configured to adjust the controller for determining the object in the road image based on the comparison difference.

[0069] In some embodiments, the apparatus 700 further includes: a semantic processing unit configured to perform one or more of semantic segmentation or semantic vectorization processing on a road image, where semantic segmentation segments the road image based on pixel information in the road image; or semantic vectorization processing performs object detection and vectorization processing on the road image based on objects of interest in the road image. In some embodiments, the apparatus 700 further includes: a determination unit configured to iteratively adjust the normalized data with respect to the prior position and angle to determine the position and angle of the object of interest.

[0070] In some embodiments, the apparatus 700 further includes: an exclusion unit configured to exclude other pixels from the object of interest in response to the distance being outside a predetermined threshold range; and an adjustment unit configured to adjust the predetermined threshold range based on the type of the object of interest.

[0071] In some embodiments, the apparatus 700 further includes: a determination unit configured to determine an object of interest in the road image including one or more of road markings, signal light signs, parking space information, and sidewalks in the road. In some embodiments, the apparatus 700 further includes: a determination unit configured to determine the distance between the object of interest and other pixels by determining the straight-line distance between pixels.

[0072] Figure 8 A schematic block diagram of an example device 800 that can be used to implement the embodiments of the present disclosure is shown. As shown, the device 800 includes a processor 801, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 802 and loaded into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0073] The various processes and treatments described above, such as method 200 and process 300, can be executed by the processor 801. For example, in some embodiments, method 200 and process 300 can be implemented as a computer software program that is tangibly contained in a machine-readable medium. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802. When the computer program is loaded into the RAM 803 and executed by the processor 801, one or more actions of method 200 and process 300 described above can be executed.

[0074] The present disclosure may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.

[0075] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed to be a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0076] The computer-readable program instructions described herein may be downloaded to respective computing / processing devices from a computer-readable storage medium or may be downloaded to an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0077] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0078] Aspects of the present disclosure are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer-readable program instructions.

[0079] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing apparatus, a device is produced that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable medium storing the instructions includes a manufacture comprising instructions that implement various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0080] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0081] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0082] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of technologies in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A method for determining an object in a road image, comprising: Determining an object of interest in the road image based on semantic processing; Determining the distance between pixels in the object of interest and other pixels; And In response to the distance being within a predetermined threshold range, including the other pixels in the object of interest.

2. The method according to claim 1, further comprising: Normalizing the object of interest to generate normalized data.

3. The method according to claim 2, further comprising: Performing a non-linear transformation on the normalized data to determine the position and angle of the object of interest.

4. The method according to claim 3, further comprising: Comparing one or more of the determined position and angle of the object of interest with one or more of a prior position and angle; And Adjusting a controller for determining the object in the road image based on the difference of the comparison.

5. The method according to claim 1, wherein determining the object of interest in the road image based on the semantic processing further comprises: Performing one or more of semantic segmentation or semantic vectorization processing on the road image, wherein: The semantic segmentation segments the road image based on pixel information in the road image; or The semantic vectorization processing performs object detection and vectorization processing on the road image based on an object of interest in the road image.

6. The method according to claim 4, wherein performing the non-linear transformation on the normalized data comprises: Iteratively adjusting the normalized data with respect to the prior position and angle to determine the position and angle of the object of interest.

7. The method according to claim 1, further comprising: In response to the distance being outside the predetermined threshold range, excluding the other pixels from the object of interest; And Adjusting the predetermined threshold range based on the type of the object of interest.

8. The method according to claim 1, wherein the object of interest in the road image includes one or more of road markings, signal light signs, parking space information, and sidewalks in the road.

9. The method according to claim 1, wherein determining the distance between the object of interest and the other pixels comprises: Determining the distance between the object of interest and the other pixels by determining the straight-line distance between pixels.

10. An apparatus for determining an object in a road image, comprising: A first determination unit configured to determine an object of interest in the road image based on semantic processing; A second determination unit configured to determine the distance between pixels in the object of interest and other pixels; And An inclusion unit configured to include the other pixels in the object of interest in response to the distance being within a predetermined threshold range.

11. A controller, comprising: At least one processor; And A memory, coupled to the at least one processor and having instructions stored thereon, which when executed by the at least one processor cause the controller to perform the method according to any one of claims 1-9.

12. A vehicle, comprising the controller according to claim 11.

13. A machine-readable storage medium having machine-executable instructions stored thereon, wherein the machine-executable instructions are executed by a processor to implement the method according to any one of claims 1 to 9.