Method and system for determining an object structure of an object, control device for a system
By extracting feature data of object structure through neural networks and generating line segments of object structure using edge endpoint and midpoint recognizers, the problem of object recognition in complex scenes is solved, and high-precision object structure recognition and driving assistance functions are realized.
Patent Information
- Application Number
- CN202480024561.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-10
- Filing Date
- 2024-05-06
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies cannot effectively recognize objects of arbitrary shapes, especially in complex urban scenarios where the recognition of lane markings and traffic islands is poor.
The system employs a neural network to identify object structures. It extracts feature data through the backbone network and combines edge endpoint and edge midpoint recognizers to generate seed point maps and topology maps. It then uses graph search criteria to connect contour endpoints and generate line segments of the object structure.
It achieves accurate recognition of objects of arbitrary shapes, supports longitudinal and lateral control of autonomous driving and driver assistance systems, and improves the accuracy and efficiency of object structure recognition.
Smart Images

Figure CN120958495A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for determining the object structure of an object, the method comprising the steps of: providing image data that describes an image of the environment in which the object is located and being received from at least one sensor device; inputting the image data into at least one neural network that is trained to determine feature data, wherein the feature data includes predetermined features regarding the basic geometry and / or color of the object. Background Technology
[0002] Object recognition is a subfield of image processing that aims to identify individual objects in an image. The term "object" can refer, for example, to a traffic island and / or lane markings and / or lane edges and / or other vehicles. Traditional methods (such as those using bounding boxes or line recognition) have limited functionality because they can only identify objects of a specific shape. Therefore, to ensure the most advantageous design possible for such object recognition, it is desirable to identify objects of arbitrary shapes, object structures, or object centerlines. Here, "object structure" refers to the basic shape and / or composition of the object's visual appearance.
[0003] Object recognition systems specifically designed to identify lanes typically only consider highway scenarios, where lane markings and road edges can be described by straight lines. However, such object recognition systems fail in urban scenarios, where objects (e.g., lane edges and / or traffic islands) can have any geometry, such as polylines. A polyline or polygonal shape refers to a continuous sequence of open or closed line segments and / or arcs.
[0004] The 2021 publication "CondLaneNet: a Top-to-down Lane Detection Framework Based on Conditional Convolution" (authors: Liu, Lizhe & Chen, Xiaohao & Zhu, Siyu & Tan, Ping) discloses a method for determining lane lines appearing vertically in an image. Because this method is based on line anchors (Zeilenanker), it cannot identify polylines or objects of arbitrary shapes.
[0005] The 2022 publication “A Keypoint-based Global Association Network for Lane Detection” (authors: Wang, Jinsheng & Ma, Yinchao & Huang, Shaofei & Hui, Tianrui & Wang, Fei & Qian, Chen & Zhang, Tianzhu) discloses a method for determining lane lines, in which lane lines are determined by globally assigning markers to the line endpoints.
[0006] The 2022 publication “RCLane: Relay Chain Prediction for Lane Detection” (authors: Xu, Shenghua & Cai, Xinyue & Zhao, Bin & Zhang, Li & Xu, Hang & Fu, Yanwei & Xue, Xiangyang) discloses a method for determining lane lines, which uses a global distance formula that only works for straight lane lines.
[0007] The object structure of an object of arbitrary shape cannot be determined from existing technologies. Summary of the Invention
[0008] The object of this invention is to identify object structures in an environment, particularly a road traffic environment, based on image data. This object is achieved through the subject matter of the independent claims. Advantageous improvements of the invention are illustrated by the dependent claims, the following description, and the accompanying drawings.
[0009] This invention provides a method for determining the object structure of an object. This method can, for example, be used as part of a computer vision function to identify the object structure of an object in a motor vehicle environment. Based on the identified object structure, a driving trajectory around the object can be calculated, for example. The method includes the following steps: receiving image data describing an image of the environment in which the object is located from at least one sensor device, and providing the image data. The image data is input into at least one neural network trained to determine feature data, wherein the feature data contains predetermined features regarding the basic geometry and / or color of the object. These features may, for example, be horizontal lines and / or vertical lines and / or arcs, to name just a few. The feature data is determined by means of a so-called backbone network, wherein the backbone network is designed as at least one artificial neural network, wherein the backbone network includes, for example, at least one residual neural network and / or densely connected convolutional network pre-trained for predetermined features regarding the basic geometry and / or color of the object. Thus, the backbone network can be additionally trained based on data containing polygonal shapes and / or structures. The feature data determined by the backbone network is included or provided, for example, as feature vectors. In other words, the backbone network is configured to receive image data or visual data from at least one sensor device and determine and / or extract feature data therefrom. The at least one neural network may include a convolutional neural network.
[0010] The backbone network can generate so-called feature maps. The backbone network may include multiple convolutional layers, so that the feature maps contain particularly large receptive fields (e.g., detecting 3×3 pixels to 200×200 pixels), thereby enabling, for example, the overall identification of objects such as traffic islands and / or lane edges.
[0011] According to the present invention, in at least one neural network, two different image processing functions can now be connected downstream of the backbone network, which can process features from the feature map independently of each other.
[0012] On one hand, at least one neural network includes a so-called edge endpoint recognizer configured to determine and / or label the edge endpoints of an object and / or the endpoints of the object's object structure. Here, edge endpoints correspondingly label predetermined object end regions (line ends) of the object. The edge endpoint recognizer performs binary classification on the feature data, where binary classification may include normalizing the feature vectors of the feature data, and normalized values above a predetermined threshold, i.e., classification values, represent edge endpoints where object structures exist. The edge endpoint recognizer can generate a so-called seed point map, where the seed point map has labeled edge endpoints. If, for example, a white lane arrow is depicted on a gray road surface in image data, the two ends (top and opposite ends) of the lane arrow, which are edge endpoints, can be identified as edge endpoints. For this purpose, the edge endpoint recognizer can be trained to recognize shapes that have a size ranging from 20x20cm to 75x75cm and terminate in a spatial direction or do not continue in that direction. Therefore, the object terminates in that direction, and an edge endpoint can be set there. Thus, a one-dimensional point (edge endpoint) is determined by means of so-called seed point detection.
[0013] On the other hand, the at least one neural network includes an edge midpoint recognizer configured to determine and / or mark edge midpoints and / or nodes and / or connection points of the object's object structure based on feature data. Edge midpoints mark the linear geometry of the object, wherein the linear geometry is part of the object's object structure. The linear geometry may be formed by straight lines and / or arcs and / or a combination of both and / or by polygonal faces.
[0014] The feature data is classified into a different type of binary classification than the first type using an edge midpoint recognizer. This second type of binary classification may involve normalizing the feature vectors of the feature data, and the normalized value above a predetermined threshold represents the linear geometry of the object structure. Furthermore, the identified edge midpoints may additionally include so-called transition probabilities (i.e., the probability between corresponding pairs of adjacent edge midpoints). The transition probabilities can be calculated, for example, as follows: For example, for edge midpoint x, the classification probability obtained by the second type of binary classification is 0.9, and for edge midpoint y (also obtained by the second type of binary classification), the classification probability is 0.75. Thus, the transition probability between edge midpoint x and edge midpoint y can be calculated, for example, as follows: Therefore, 0.269 is the transition probability between the midpoints x and y of the edge. Here, the transition probability can also be calculated without subtraction, for example, according to the following formula: This transition probability can then be used in a graph search criterion, which will be described below. Therefore, the edge midpoint recognizer can be designed to calculate (generate) the transition probability. A topological map containing the midpoints of the object's structural edges can be generated using the edge midpoint recognizer. The graph can be marked on the topological map using the marked edge midpoints. Therefore, the feature vector is preferably subjected to a binary classification accordingly, where the feature vector contains parameterizable (numerically represented) attributes of the object's geometry and / or style in vector form. Different features characterizing the style can form different dimensions of the feature vector. Therefore, using feature vectors makes subsequent binary classification easier because it significantly reduces the attributes to be classified (e.g., only a feature vector consisting of, for example, ten numbers needs to be considered, instead of the entire image). For example, in the lane arrow example above, the individual segments or lines of the boundary line formed by the intersection of the white lane arrow and the gray road surface can be identified as edge midpoints accordingly. An edge midpoint recognizer can be trained to mark segments or sections of continuous lines extending in two directions in the environment, with measured lengths ranging from 20cm to 75cm, as edge midpoints.
[0015] The determination of edge endpoints and edge midpoints of the object structure is performed independently of each other. Edge endpoint recognizers and / or edge midpoint recognizers may be included as subnetworks and / or as at least one layer of at least one neural network.
[0016] The seed point map and the topology map can then be combined or merged. This generates a seed point map-image that contains the labeled edge endpoints and the labeled edge midpoints.
[0017] For example, the object structure is determined by means of edge endpoints and edge midpoints by at least one neural network or downstream software. For this purpose, the edge midpoint closest to the corresponding edge endpoint is marked as the contour endpoint of the object structure, and at least two corresponding contour endpoints are connected by corresponding pairs of adjacent edge midpoints.
[0018] This yields the final object structure.
[0019] In other words, edge endpoints are interpreted as the contour endpoints of the object structure. The determined object structure is then provided for computer vision functions. These computer vision functions may, for example, be included in a motor vehicle as a parking assistance system and / or a reversing system and / or a lane recognition system and / or a lane change assistance system. Furthermore, the invention can be used for autonomous driving, for which the determined object structure is used for longitudinal and lateral control of the motor vehicle. An advantage of the invention is that the object structure to be determined can include any geometry because the object structure is not described as a bounding box of the object's faces, but rather consists of edge midpoints arranged in the image data to form an object structure of (arbitrary shape).
[0020] The present invention also includes improvements that bring additional advantages.
[0021] One improved approach specifies that at least two contour endpoints are connected via edge midpoints using a graph search criterion, thereby generating line segments (paths). The graph search criterion can be a graph algorithm, particularly Dijkstra's algorithm. Dijkstra's algorithm is used to calculate and / or mark the shortest and / or most cost-effective line segment between two contour endpoints via pairs of adjacent edge midpoints.
[0022] One improved scheme specifies that line segments include cost values, which are calculated using graph search criteria. The cost value of a line segment is derived from the transformation probabilities of its contour endpoints and edge midpoints, where the cost value is, for example, the product of the two or a scalar product. In other words, the cost value of a line segment is obtained from the transformation probabilities of the contour endpoints and edge midpoints contained in the line segment, wherein the line segment correspondingly includes at least two contour endpoints connected by pairs of adjacent edge midpoints. In one optimization algorithm, for example, a line segment with a cost value of 0.16 is more cost-effective than a line segment with a cost value of 0.8.
[0023] One improved approach specifies that line segments with cost values exceeding a predetermined threshold are deleted using a deletion criterion, such that the retained line segments reveal the object's structure. The advantage of this is that duplicate or copied line segments are removed. The deletion criterion can be, for example, a non-maximum suppression unit.
[0024] One improved scheme specifies that the at least two contour endpoints are connected by pairs of adjacent edge midpoints, wherein the connection is achieved through a corresponding transition, wherein the transition includes, for example, a transition line, and wherein the transition includes a regression value or a predetermined number of distance values (distortion values) with respect to local curvature, which causes positional shift when generating the line segment to avoid overfitting of the line segment to be generated with the contour endpoints and / or edge midpoints. Therefore, the predetermined number of distance values defines the orthogonal distance or compensation amount of the line segment to be generated. Positional shift can, for example, refer to the movement of the line segment in a three-dimensional image, wherein the distance values are provided as X-values, for example, in a coordinate system with three axes (X-axis, Y-axis, and Z-axis), resulting in movement of the line segment along the X-axis. In other words, a predetermined number of distance values can cause, for example, dimensional shift or orthogonal shift in a two-dimensional or three-dimensional image of the line segment to be generated. The dimension or number of distance values is obtained by the following formula (F): c = k × (s + 1), where k is the number of neighborhoods (neighborhood connections) of the contour endpoints and / or edge midpoints, and s is a predetermined number of offset sample values, which in particular provides a variable for the multidimensional scaling of the line segment to be generated, wherein the number of offset sample values indicates how the positions of the contour endpoints and / or edge midpoints move, and thus indicates how the line segment to be generated moves. In other words, c is the number of distance values in the cells or image points of the feature map. The offset sample values themselves can be, for example, a kind of "placeholder" between the actual position of the object structure and the desired position of the line segment, and are learned and inferred by at least one neural network. The predetermined number of offset sample values is a so-called hyperparameter of at least one neural network. Thus, the predetermined number of offset sample values s and the predetermined neighborhood connections k can be used to orthogonally regress the actual object structure using the number of distance values c that can be calculated by formula F, wherein each image point contained in the object structure can include a regression segment (line segment) of the actual object structure. Thus, the number of distance values c can provide corresponding grid points / support bits for the generated line segment. Here, the number of offset sample values s can be parameterized. The midpoints of the edges and / or the endpoints of the contours may contain a predetermined number of distance values c, so that the movement or local curvature of the line segment to be generated can be obtained later when applying the graphic search criteria.
[0025] One improvement specifies the use of self-attention units to determine feature data. Using self-attention units reduces the computational complexity of at least one neural network used to determine feature data.
[0026] One improved approach involves using feature pyramid network units to determine feature data. Using feature pyramid network units allows for a more tightly coupled connection between so-called low-level (local) and high-level (semantic) features.
[0027] An improved approach specifies that at least one neural network is trained to determine the object structure of an object using error feedback (backpropagation), wherein training data containing features (labels) of the object structure is used. Therefore, at least one neural network is trained using supervised learning. In the case of backpropagation, the input, converted into an input vector, such as image data, is propagated through at least one neural network. The output thus obtained by at least one neural network is compared with the desired output. The difference between the two values is considered the error of the neural network, which is now backpropagated through the output layer to the input layer. Here, the weights of the neuron connections in at least one neural network vary according to their influence on the error. This ensures that the output is close to the desired output when the input is applied again. At least one neural network can be corrected (improved) through backpropagation. Therefore, the parameters of at least one neural network can be optimized or improved. With the thus improved parameters of at least one neural network, the at least one neural network is able to determine a meaningful output vector (output) from the input vector (input) in the application phase, which deviates from the input vector of the initially learned training case.
[0028] For application scenarios or conditions that may occur in the method but are not explicitly described herein, it may be specified that, according to the method, error messages and / or requests for user feedback be output, and / or default settings and / or predetermined initial states be set.
[0029] The invention also includes a control device for the system. This control device may have a data processing device or a processor device configured to perform embodiments of the method according to the invention. For this purpose, the processor device may have at least one microprocessor and / or at least one microcontroller and / or at least one FPGA (Field Programmable Gate Array) and / or at least one DSP (Digital Signal Processor). Furthermore, the processor device may have program code having program instructions that, when implemented by the processor device, cause the processor device to perform the method according to any one of the preceding method claims. The program code may be stored in the data memory of the processor device. The processor device may, for example, be based on at least one circuit board and / or at least one SoC (System on Chip).
[0030] The invention also includes a system comprising a control device and at least one sensor device. The system can perform the method according to the invention via the control device.
[0031] The invention also includes motor vehicles comprising the aforementioned systems. The motor vehicle according to the invention is preferably designed as an automobile, particularly a passenger car or commercial vehicle, or as a bus or motorcycle.
[0032] As another solution, the invention also includes a computer-readable storage medium containing program code that, when executed by a computer or computer network, causes the computer or computer network to perform an implementation of the method of the invention. The storage medium may be provided at least partially as non-volatile data storage (e.g., flash memory and / or SSD (solid-state drive)) and / or partially as volatile data storage (e.g., RAM (random access memory)). The storage medium may be arranged in a computer or computer network. However, the storage medium may also operate, for example, as a so-called app store server and / or a cloud server on the Internet. A processor circuit having, for example, at least one microprocessor may be provided via the computer or computer network. The program code may be provided as binary code and / or assembly code and / or source code in a programming language (e.g., C) and / or program scripts (e.g., Python).
[0033] The present invention also covers combinations of features of the illustrated embodiments. Therefore, the present invention also covers implementations that correspondingly have combinations of features of multiple embodiments of the illustrated embodiments, unless these embodiments are described as mutually exclusive. Attached Figure Description
[0034] The embodiments of the present invention are described below. Wherein:
[0035] Figure 1 An exemplary graphical representation is shown, illustrating features associated with the method according to this disclosure;
[0036] Figure 2 A flowchart illustrating an embodiment of the method according to the present invention is shown;
[0037] Figure 3 An exemplary graphical representation of a seed point map generated by an edge endpoint recognizer according to an embodiment of the present invention is shown;
[0038] Figure 4 An exemplary graphical representation of a topology map generated by an edge midpoint identifier according to an embodiment of the present invention is shown;
[0039] Figure 5 An exemplary graphical representation of the midpoint in a square grid according to the present invention is shown, wherein a predetermined number of offset sample values affect the orthogonal distance of the line segment to be generated;
[0040] Figure 6 An exemplary graphical representation of line segment position movement according to an embodiment of the present invention is shown using a certain number of offset sample values;
[0041] Figure 7 An exemplary graphical representation of the most cost-effective path in a graph found by a graph search criterion according to an embodiment of the present invention is shown.
[0042] Figure 8 An exemplary graphical representation of marking line segments before applying deletion criteria according to an embodiment of the present invention is shown;
[0043] Figure 9 An exemplary graphical representation of lane markings and lane edges providing a polyline geometry illustration according to an embodiment of the present invention is shown. Detailed Implementation
[0044] The embodiments described below are preferred embodiments of the present invention. In the embodiments, the described components of the embodiments correspond to various features of the present invention that can be considered independently of each other, and which also independently improve the present invention. Therefore, this disclosure should also cover combinations of features of the embodiments other than those shown. Furthermore, the described embodiments may be supplemented by other features among the already described features of the present invention.
[0045] In the accompanying drawings, the same reference numerals respectively denote elements with the same function.
[0046] According to this embodiment, such as Figure 1 As shown, image data 1 can be acquired using at least one sensor device in a system located within a motor vehicle. Here, image data 1 can, for example, illustrate the driving environment of the motor vehicle. Figure 1 It is shown that at least one neural network image data 1 can be received from at least one sensor device of the system, wherein the image data 1 describes the environment. Therefore, the image data 1 can be input into at least one neural network, wherein the at least one neural network includes a backbone network 2, which can be used to determine feature data. Additionally or alternatively, at least one neural network may provide self-attention units 4 and / or feature pyramid network units 5. The backbone network 2 and / or self-attention units 4 and / or feature pyramid network units 5 can be trained, in particular, to determine edges and / or linear and / or geometric points and / or corner points and / or the ends of lines and / or the ends of object regions. Furthermore, the backbone network 2 and / or self-attention units 4 and / or feature pyramid network units 5 can be trained to determine traffic islands and / or lane edges and / or lane markings. Additionally, so-called features can be generated. Figure 6 Among them, features Figure 6 It can contain feature data.
[0047] Then, the feature data can be binary classified using an edge endpoint recognizer 8, which labels edge endpoints 18 in image data 1. For example, binary classification may include normalizing the feature vectors, and normalized values above a predetermined threshold, i.e., classification values, indicate the presence of edge endpoints 18 with object structures. Furthermore, binary classification can be performed on the feature vectors using an edge midpoint recognizer 9. Feature data containing feature vectors can be normalized accordingly, and feature vectors above a predetermined threshold may be (determined) to contain the linear geometry of object O, where the linear geometry can be labeled using the edge endpoint recognizer 8. A so-called seed point map 11 can be generated using the edge endpoint recognizer 8, where the seed point map 11 has labeled edge endpoints 18. A topology map 12 can be generated using the edge midpoint recognizer 9, which, for example, contains edge midpoints 19 of object structures. The seed point map 11 and the topology map 12 can then be merged or fused. This generates a seed point map-image 14, which may have labeled edge endpoints 18 and labeled edge midpoints 19. Then, a graph search criterion 15 can be applied, which calculates the most cost-effective path between two edge endpoints 18, wherein the edge midpoint 19 closest to the corresponding edge endpoint 18 can be marked as the contour endpoint 20 of the object structure, and at least two corresponding contour endpoints 20 can be connected by a pair of adjacent edge midpoints 19, thereby generating a line segment L.
[0048] After applying graph search criterion 15, a so-called line graph 16 can be generated, wherein the line graph 16 may have line segments L generated using graph search criterion 15. Finally, a so-called deletion criterion 17 can be applied, wherein the deletion criterion 17 may in particular be a non-maximum suppression unit, which can delete one or more line segments L that may have a cost value higher than a predetermined threshold, thereby retaining one or more line segments L that can indicate the object structure of object O. The object structure determined by the method according to the invention can then be used for driving assistance functions of motor vehicles, such as lane recognition and / or parking assistance.
[0049] In one design, the determination of edge endpoints 18 by means of edge endpoint recognizer 8 can be performed using a cost function, such as the dissimilarity or similarity of the grayscale values of pixels (pixels) in image data 1. Here, new pixels can be selected from the neighborhood whose grayscale value is most similar to the grayscale value of the starting point of edge endpoint 18; the neighborhood can specifically refer to an eight-neighborhood. The identification and / or labeling of edge endpoints 18 is interrupted when a defined area of object O is reached or when the similarity of grayscale values no longer meets a threshold. A parameterless interruption criterion can also be applied, for example, assigning at least one classification value with respect to the cost function to each pixel. The arithmetic mean of all costs K1 can be calculated for all edge endpoints 18 to the ends of the previous object O. Then, the arithmetic mean of all costs K2 can be calculated for all neighboring pixels of the previous object O. For each newly acquired pixel, the cost difference |K1-K2| can be considered as a function of the number of acquired pixels. The function can be calculated until a maximum number of pixels is determined. Now, as a cessation criterion, we can compute such a part of the function where the function has its global maximum.
[0050] The edge midpoint recognizer 9 can be designed to generate a so-called topology map 12, wherein the topology map 12 can contain the edge midpoints 19 marked by the edge midpoint recognizer 9 of the input image data 1.
[0051] To generate one or more line segments L using the Dijkstra algorithm, a square grid with eight neighborhoods can be used, where the number of pixels (image points) or cells of the feature map can be given. A cost function can be calculated based on the transition probabilities of the image points. Regarding the cost function, the path with the lowest cost can now be found. Here, it can start from the contour endpoint 20 p and end at a predetermined contour endpoint 20 q. Once the contour endpoint 20 q is reached, the line segment L can be generated. Therefore, the contour endpoints 20 and / or edge midpoints 19 of the image points or markers can respectively include one or more transition probabilities (associations) to the adjacent image points or edge midpoints 19. This means that, according to the previous parameter specifications, the contour endpoints 20 and / or edge midpoints 19 of the image points or markers can include the transition probabilities of image points to the right of the corresponding image point, the transition probabilities of image points above the corresponding image point, etc. For example, it can be parameterized to determine whether the diagonal neighborhood of the corresponding image point (e.g., the upper right of the corresponding image point, the upper left of the corresponding image point, etc.) should also be considered.
[0052] Figure 2In method step S10, image data 1 is shown, illustrating an image of the environment in which object O is located, and can be received from at least one sensor device. Image data 1 can be input into at least one neural network, wherein the at least one neural network can be trained to determine feature data, wherein the feature data may contain predetermined features regarding the basic geometry and / or color of object O. In step S20, the edge endpoints 18 of the object structure of object O can be determined by applying at least one neural network edge endpoint recognizer 8 to the feature data, wherein the edge endpoints 18 respectively mark predetermined object end regions. In step S30, the edge midpoints 19 of the object structure can be determined by applying at least one neural network edge midpoint recognizer 9 to the feature data, wherein the edge midpoints 19 can mark the line geometry of object O, wherein the line geometry may be part of the object structure of object O. In step S40, the object structure can be determined by at least one neural network using the edge endpoints 18 and the edge midpoints 19 in such a way that the edge midpoint 19 closest to the corresponding edge endpoint 18 is marked as the contour endpoint 20 of the object structure, and at least two corresponding contour endpoints 20 can be connected by corresponding pairs of adjacent edge midpoints 19. In step S50, the object structure can be provided for computer vision functions.
[0053] Figure 3 An exemplary graphical representation of a seed point map is shown. Edge endpoints 18 of the object structure can be marked on the seed point map using an edge endpoint recognizer 8. Here, for example, traffic islands or lane edges and the top and / or line ends of lane markings can be interpreted as edge endpoints 18. Exemplarily, two of these edge endpoints 18 are indicated by reference numerals. In this case, at least one neural network can determine and / or mark traffic islands and lane markings, which are marked by reference numeral O, because the neural network is trained only for this purpose.
[0054] Figure 4 The output of the edge midpoint recognizer 9 is shown, which is a so-called topology map 12. The figure shows marked edge midpoints 19 on the corresponding line geometry determined by the edge midpoint recognizer 8, wherein the edge midpoints 19 can specifically mark image regions depicting line segments. For example, the edge midpoints 19 are provided with reference numerals 19. These reference numerals can refer to any identified and / or marked edge midpoint 19. An edge midpoint 19 may have one or more transition probabilities to its adjacent edge midpoints 19 and / or neighborhoods, wherein the transition probabilities can be calculated and / or generated by the edge midpoint recognizer 9. Figure 3 similar, Figure 4 The resulting square grid is shown, where each image point or pixel can be provided, for example, in an eight-neighborhood manner.
[0055] Figure 5 The edge midpoint 19 is shown to contain four-neighborhoods (4-con), eight-neighborhoods (8-con), or sixteen-neighborhoods (16-con). A predetermined number of offset sample values can be determined here, wherein the number of offset sample values can affect the number of distance values c. The edge midpoint 19 and / or the contour endpoint 20 can contain this number, thereby enabling the generation of the movement or local curvature of the line segment L to be generated later when applying the graphic search criterion 15.
[0056] Figure 6 A square grid of size w × h is shown, where w represents the width of the square grid and h represents the length of the square grid. A line segment L is shown, which has been moved positionally by means of the stated number of distance values c. The dimensional or orthogonal movement of the line segment L to be generated can be obtained by the number of distance values c, which can be obtained by the formula Fc = k × (s + 1). This means that a predetermined number of offset sample values s and a predetermined number of neighborhood connections k can be used to revert the orthogonal movement to the actual polyline (actual object structure) using the number of distance values c calculated by formula F, where each pixel contained in the polyline (object structure) can include a regression segment of the actual polyline. Therefore, the number of distance values c can provide grid points / support bits for the generated line segment L accordingly. The number of offset sample values s can be parameterized here. The formula (o = w × h × c = (w × h) × (k × (s + 1))) is also shown, which can be used to calculate the number of dimensions of the square grid.
[0057] Figure 7 The diagram shows a line segment L marked after applying graph search criterion 15, where line segment L follows the most cost-effective path. Contour endpoints 20 are also shown. Contour endpoints 20 can be created by marking the edge midpoint 19 closest to the corresponding edge endpoint 18 as the contour endpoint 20 of the object structure. More than one path between contour endpoints 20 can also be calculated and / or found according to graph search criterion 15. This can occur when graph search criterion 15 attempts to find the most cost-effective path from each contour endpoint 20 to another contour endpoint 20. In other words, graph search criterion 15 determines the optimal path within the graph (at least two contour endpoints 20 connected by a pair of adjacent edge midpoints 19). This optimal path can now be marked along a chain passing through contour endpoints 20 and contour midpoints 19, thereby obtaining a line segment L that can illustrate one or at least one object structure of object O.
[0058] Figure 8The diagram shows marked line segments L (polylines) that illustrate one or more object structures, where repeating line segments D are also marked. Repeating line segments D should be removed and / or deleted so that the driver assistance function receives only one polyline as input for each object O. Therefore, deletion criterion 17 can be applied to the determined line segments L in a final step to delete repeating line segments D with a cost value higher than a predetermined threshold, while retaining line segments L indicating the object structure of the corresponding object O, such as... Figure 9 As shown. Therefore, lane markings and traffic islands or lane edges can be identified and described through polyline geometry.
[0059] This invention provides a method for determining the polyline (object structure) of an object O, and an object recognition algorithm (at least one neural network) based on at least one neural network and / or a convolutional neural network. The input may be a camera image (image data 1). The output may be the identified object O or the object structure of object O in the image, which can be described, for example, by a polyline geometry. At least one neural network may be based on a (single-stage) recognition architecture, suitable for real-time applications (e.g., autonomous driving). At least one neural network may use a dual-unit architecture (two binary classifiers). The first unit (edge endpoint recognizer 8) can identify and / or label key points or edge endpoints 18 (endpoints) of the polyline object (object structure) in the image. The second unit (edge midpoint recognizer 9) can identify and / or label the topological connectivity graph (graphics) of the object in the probe image (image data 1). The key points and connectivity graph can then be used in an efficient post-processing step to search for polylines.
[0060] In summary, these examples demonstrate how systems and methods for object recognition can be provided based on at least one neural network and / or convolutional network, where the input is a camera image (image data 1) and the output is one or more identified objects O and / or one or more defined object structures in the image, which are described by a polyline geometry.
[0061] List of reference numerals
[0062] 1 Image Data
[0063] 2. Backbone Network
[0064] 4 Self-attention units
[0065] 5 Feature Pyramid Network Units
[0066] 6 Feature Map
[0067] 8 Edge endpoint recognizer
[0068] 9. Edge Midpoint Recognizer
[0069] 11 Seed Point Map
[0070] 12 Topology Diagram
[0071] 14 Seed Point Diagram
[0072] 15. Graphical Search Standards
[0073] 16-line chart
[0074] 17 Deletion Criteria
[0075] 18 Edge endpoints
[0076] 19. Midpoint of the edge
[0077] 20 Contour Endpoints
[0078] 4-con = four-neighborhood
[0079] 8-con = eight-neighborhood
[0080] 16-con = sixteen-neighborhood
[0081] L = line segment
[0082] O = object
[0083] c = the number of distance values
[0084] s = Number of offset sample values
[0085] k = number of neighborhood connections
[0086] w = width of the square grid
[0087] h = Length of the square grid
[0088] F = Formula
Claims
1. A method for determining the object structure of an object (O), comprising the following steps: - Provide image data (1) that describes an image of the environment in which the object (O) is located, the image data being received from at least one sensor device; - Input image data (1) into at least one neural network, which is trained to determine feature data, wherein the feature data includes predetermined features about the basic geometry and / or color of the object (O); Its features are, a) By applying the edge endpoint recognizer (8) of the at least one neural network to the feature data, the edge endpoints (18) of the object structure of the object (O) are determined, wherein the edge endpoints (18) mark the predetermined object end regions of the object (O); b) By applying the edge midpoint recognizer (9) of the at least one neural network to the feature data, the edge midpoints (19) of the object structure are determined, wherein the edge midpoints (19) mark the line geometry of the object (O), wherein the line geometry is part of the object structure of the object (O), wherein a) and b) are performed independently of each other; c) The object structure is determined by means of edge endpoints (18) and edge midpoints (19) in such a way that the edge midpoint (19) closest to each edge endpoint (18) is marked as the outline endpoint (20) of the object structure, and at least two outline endpoints (20) are connected by corresponding pairs of adjacent edge midpoints (19). - Provides object structures for computer vision functions.
2. The method according to claim 1, characterized in that, In step c), at least two contour endpoints (20) are connected via edge midpoints (19) by means of a graph search criterion (15), thereby generating a line segment (L).
3. The method according to claim 2, characterized in that, The line segment (L) includes a cost value, and the cost value is calculated using the graphic search criterion (15).
4. The method according to claim 3, characterized in that, By using deletion criteria (17), line segments (L) with cost values higher than a predetermined threshold are deleted, so that the retained line segments (L) can display the object structure.
5. The method according to any one of the preceding claims, characterized in that, The edge midpoint (19) includes a predetermined number of distance values (c) with respect to the local curvature, which cause a positional shift when the line segment (L) is generated.
6. The method according to any one of the preceding claims, characterized in that, The feature data is determined by means of a self-attention unit (4).
7. The method according to any one of the preceding claims, characterized in that, The feature data is determined by means of the feature pyramid network unit (5).
8. The method according to any one of the preceding claims, characterized in that, The at least one neural network is trained with the aid of error feedback to determine the structure of an object, wherein training data having features of the object structure is used.
9. A control device for a system, wherein, The control device has a processor device having program instructions that, when executed by the processor device, cause the processor device to perform the method according to any one of the preceding method claims.
10. A system having a control device according to claim 9 and at least one sensor device.