Method and device for detecting indication sign, controller, vehicle and medium

By detecting the location and categories of key points in the indicator marks, and using neural network models to determine the semantics of the indicator marks, the problems of information redundancy and inaccurate detection in traditional methods are solved, and more efficient indicator mark detection is achieved.

CN120236254APending Publication Date: 2025-07-01ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311871123.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

When detecting indicators, identifying the position of the constituent elements through bounding boxes leads to redundancy of information and inaccurate detection. Especially when there are many constituent elements, the bounding boxes overlap severely and the semantics of the arrow elements cannot be obtained.

Method used

By acquiring the images collected by the on-board camera, detecting the location and categories of key points in the indicator marks, and using neural network models to determine the semantics of the indicator marks, avoiding information redundancy caused by bounding box marks, and improving detection accuracy.

Benefits of technology

It improves the accuracy and speed of indicator sign detection, reduces redundant information, enhances the richness of key point information, and ensures the integrity of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236254A_ABST
    Figure CN120236254A_ABST
Patent Text Reader

Abstract

The invention relates to a method and device for detecting an indication sign, a controller, a vehicle and a medium. The method comprises the following steps: acquiring an image containing an indication sign, wherein the composition elements of the indication sign comprise at least one of an arrow element and a line segment element; the method further includes determining positions and categories of key points constituting the elements based on the image, wherein the positions of the key points are used to identify areas of the indication signs. In addition, the method also includes determining semantics of the indication sign based on the category of the keypoint. Through the mode, the position and classification information of the key points can be obtained from the image, the richness of the information of the key points is improved, the indication marks in the image are represented through the key points, and inaccurate and incomplete results caused by direct classification of the whole indication marks are avoided. According to the embodiment of the invention, the accuracy of indication sign detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to the field of intelligent driving, and more particularly to methods, devices, controllers, vehicles, and media for detecting indication signs. Background Art

[0002] As an important traffic sign guiding vehicles, indication signs are widely used in scenarios such as roads and parking lots. For example, indication signs can predict road conditions and indicate information such as driving directions and distances for vehicles, thus standardizing traffic behaviors in various vehicle usage scenarios. Indication signs generally consist of arrow graphics, and a few indication signs also involve other graphics such as line graphics.

[0003] With the development of autonomous driving technology, whether it is an autonomous driving system or an Advanced Driving Assistance System (ADAS for short), its planning and control algorithms need to obtain road environment information such as road markings on the road through in-vehicle cameras to control the driving process of the vehicle. The finer the road environment information obtained by the planning and control algorithm module, the closer the decision made is to that of a human or more in line with traffic regulations. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, device, controller, vehicle, and medium for detecting indication signs. In the embodiments of the present disclosure, first, an image containing an indication sign collected by a sensing device such as an in-vehicle camera is obtained, where the indication sign may include at least one arrow element, at least one line element, or a combination of at least one arrow element and at least one line element. Then, the image is detected to determine the positions and categories of the key points of each constituent element in the indication sign. Further, when the positions of the key points meet the position conditions, based on the categories of the key points of each constituent element in the indication sign, the semantics of the indication sign are determined (for example, the semantics of the indication sign may be the compositional information of the categories of the key points of each constituent element).

[0005] In this way, the positions and classification information of the key points can be obtained from the image, the richness of the information of the key points can be improved, and the indication sign in the image can be characterized by the key points, avoiding inaccurate and incomplete results caused by directly classifying the entire indication sign. Therefore, the method of the present disclosure can improve the accuracy of indication sign detection.

[0006] In a first aspect of the present disclosure, a method for detecting an indication sign is provided. The method includes obtaining an image including the indication sign, where the constituent elements of the indication sign include at least one of an arrow element and a line segment element. The method further includes determining, based on the image, the positions and categories of the key points of the constituent elements. Additionally, the method further includes determining the semantics of the indication sign based on the categories of the key points in response to the positions of the key points satisfying a position condition.

[0007] In a second aspect of the present disclosure, a device for detecting an indication sign is provided. The device includes an image acquisition module configured to obtain an image including the indication sign, where the constituent elements of the indication sign include at least one of an arrow element and a line segment element. The device further includes a key point determination module configured to determine, based on the image, the positions and categories of the key points of the constituent elements. Additionally, the device further includes an indication sign determination module configured to determine the semantics of the indication sign based on the categories of the key points in response to the positions of the key points satisfying a position condition.

[0008] In a third aspect of the present disclosure, a controller is provided. The controller includes one or more processors; and a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method provided in the first aspect of the present disclosure.

[0009] In a fourth aspect of the present disclosure, a vehicle is provided. The vehicle includes the controller provided in the third aspect of the present disclosure.

[0010] In a fifth aspect of the present disclosure, a machine-readable storage medium is provided. Machine-executable instructions are stored on the machine-readable storage medium, where the machine-executable instructions, when executed by a processor, implement the method provided in the first aspect of the present disclosure.

[0011] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0013] Figure 1 A schematic diagram of an example environment in which some embodiments of the present disclosure can be implemented is shown;

[0014] Figure 2The flowchart of a method for detecting an indication sign according to some embodiments of the present disclosure is shown;

[0015] Figure 3A The schematic diagram of an indication sign including a lower left corner point according to some embodiments of the present disclosure is shown;

[0016] Figure 3B The schematic diagram of an indication sign including a lower right corner point according to some embodiments of the present disclosure is shown;

[0017] Figure 3C The schematic diagram of an indication sign including a vertex according to some embodiments of the present disclosure is shown;

[0018] Figure 3D The schematic diagram of an indication sign including an inflection point according to some embodiments of the present disclosure is shown;

[0019] Figure 3E The schematic diagram of an indication sign including a vertex and a corner point according to some embodiments of the present disclosure is shown;

[0020] Figure 3F The schematic diagram of an indication sign including an end point according to some embodiments of the present disclosure is shown;

[0021] Figure 3G The schematic diagram of an indication sign including multiple key points according to some embodiments of the present disclosure is shown;

[0022] Figure 4 The schematic diagram of a process for detecting key points of an indication sign according to some embodiments of the present disclosure is shown;

[0023] Figure 5 The block diagram of a device for detecting an indication sign according to some embodiments of the present disclosure is shown; and

[0024] Figure 6 The schematic block diagram of an example device according to some embodiments of the present disclosure is shown. Detailed Description of the Embodiments

[0025] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0026] In the description of the embodiments of the present disclosure, the term "including" and its like shall be understood as an open inclusion, that is, "including but not limited to". The term "based on" shall be understood as "at least partially based on". The term "one embodiment" or "the embodiment" shall be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.

[0027] For an autonomous driving system or an advanced driver assistance system, in order to ensure the accuracy and timeliness during the driving process, the implementation of its planning and control algorithms requires an accurate and efficient way to obtain indication signs. Conventionally, the positions of the constituent elements in the indication signs are obtained by means of image detection. However, in the conventional method, the positions of the constituent elements are generally identified by bounding boxes, which requires a large number of parameters when the number of constituent elements is large and is prone to information redundancy. In addition, the semantics of arrow elements cannot be obtained in the conventional method, resulting in inaccurate and incomplete detection results.

[0028] To this end, the embodiments of the present disclosure provide a method for detecting indication signs. First, an image containing indication signs collected by a sensing device such as an in-vehicle camera is obtained, where the indication signs may include at least one arrow element, at least one line segment element, or a combination of at least one arrow element and at least one line segment element. Then, the image is detected to determine the positions and categories of the key points of each constituent element in the indication sign. Further, when the positions of the key points meet the position conditions, based on the categories of the key points of each constituent element in the indication sign, the semantics of the indication sign are determined (for example, the semantics of the indication sign may be the compositional information of the categories of the key points of each constituent element).

[0029] In this way, the positions and categories of each key point of the constituent elements in the indication sign can be directly calculated based on the image. The position of the key point can be directly used to identify the area of the indication sign (the position of the key point can be used to identify the area of the corresponding constituent element. When there is one constituent element, the area of the indication sign can be directly identified based on the position of the key point. When there are multiple constituent elements, the area of the indication sign can be identified based on the combination of the positions of the key points). Through the position of the key point, the area of the indication sign can be quickly determined, and the information redundancy caused by identifying the indication sign by means of a bounding box, etc. can be avoided (especially when there are many constituent elements in the indication sign, the bounding boxes will overlap a lot), and at the same time, the number of parameters in the process of determining the area of the indication sign can be reduced. The category of the key point, as the associated information of the key point, can improve the richness of the information of the key point, and thus the semantics of the indication sign can be directly determined based on the category of the key point, avoiding the inaccurate and incomplete results caused by directly classifying the entire indication sign. Therefore, the method of the present disclosure can improve the accuracy and speed of indication sign detection and reduce redundant information.

[0030] Figure 1 FIG. shows a schematic diagram of an exemplary environment 100 in which some embodiments of the present disclosure can be implemented. Referring to Figure 1 Exemplary environment 100 is a partial parking scenario in a parking lot. Exemplary environment 100 includes a straight lane 102 and a transverse lane 104, both of which are double lanes on the left and right. On both sides of the straight lane 102 and the transverse lane 104, a plurality of parking spaces are arranged in sequence, such as parking space 106, and there are a plurality of parked vehicles in the plurality of parking spaces, such as parked vehicle 108. The host vehicle 110 is traveling in the right lane of the straight lane 102.

[0031] Continuing to refer to Figure 1 On the ground in front of the host vehicle 110, there is an indication sign 112, which includes a straight arrow 114 (which can be called an arrow element) and a right-turn arrow 116 (which can be called an arrow element). Therefore, there are two constituent elements in the indication sign 112. There is also an indication sign 118 in the left lane of the straight lane 102, and there is only one straight arrow in the indication sign 118. Therefore, there is only one constituent element in the indication sign 118.

[0032] Continuing to refer to Figure 1, an electronic device 120 is provided in the vehicle 110. The electronic device 120 can be implemented by a microcontroller unit (MCU for short), a central processing unit (CPU for short), a graphics processing unit (GPU for short), a field programmable gate array (FPGA for short), or other programmable logic devices, application specific integrated circuits, discrete gate or transistor logic devices, discrete hardware components, etc. The electronic device 120 includes an image acquisition unit 122, a key point determination unit 124, and an indication mark determination unit 126.

[0033] The image acquisition unit 122 is configured to receive an image containing the indication mark 112 from a vision sensor such as an in-vehicle camera. The constituent elements of the indication mark 112 include a straight arrow 114 and a right-turn arrow 116. The key point determination unit 124 is configured to analyze the image to obtain the positions and categories of the key points of the straight arrow 114 and the right-turn arrow 116 in the indication mark 112. Among them, the position of the key point can be used to identify the position of the straight arrow 114 or the right-turn arrow 116, and the category of the key point is the semantic information of the key point. The category of the key point is generally related to the corresponding straight arrow 114 or right-turn arrow 116. The indication mark determination unit 126 is configured to determine the semantics of the indication mark 116 according to the category of the key point. For example, the semantics of the indication mark 116 can be a combination of the categories of the key points.

[0034] In this way, the positions and categories of each key point of the constituent elements (the straight arrow 114 and the right-turn arrow 116) in the indication mark 112 can be directly calculated based on the collected image. The position of the key point can be directly used to identify the area of the indication mark 114 (the position of the key point can be used to identify the area of the corresponding constituent element, and the area of the indication mark 112 can be identified based on the combination of the positions of the key points). The area of the indication mark 112 can be quickly determined through the position of the key point, and the information redundancy caused by identifying the indication mark 112 by means of a bounding box, etc. can be avoided (especially when there are many constituent elements in the indication mark 112, the bounding boxes will overlap a lot), and at the same time, the number of parameters in the process of determining the area of the indication mark 112 can be reduced. The category of the key point, as the semantic information of the key point, can improve the richness of the information of the key point, and then the semantics of the indication mark 112 can be directly determined based on the category of the key point, avoiding the inaccurate and incomplete results caused by directly classifying the entire indication mark 112. Therefore, the method of the present disclosure can improve the accuracy and speed of detecting the indication mark 112 and reduce redundant information.

[0035] It should be understood that the architecture and functions in the example environment 100 are described only for exemplary purposes, without implying any limitation on the scope of the present disclosure. Embodiments of the present disclosure can also be applied to other environments with different structures and / or functions.

[0036] The following will be combined with Figures 2 to 6 to describe in detail the process according to the embodiments of the present disclosure. For ease of understanding, the specific data mentioned in the following description are all exemplary and are not used to limit the protection scope of the present disclosure. It can be understood that the embodiments described below may also include additional actions not shown and / or actions shown may be omitted, and the scope of the present disclosure is not limited in this regard.

[0037] Figure 2 A flowchart of a method 200 for detecting an indication sign according to some embodiments of the present disclosure is shown. In some embodiments, in the example environment 100 shown in Figure 1 , the method 200 can be executed by the electronic device 120. It should be understood that the method 200 may also include additional actions not shown and / or actions shown may be omitted, and the scope of the present disclosure is not limited in this regard.

[0038] At step 202, the electronic device 120 acquires an image containing an indication sign, and the constituent elements of the indication sign include at least one of an arrow element and a line segment element. In some embodiments, there are multiple constituent elements in the indication sign, and each constituent element corresponds to one of the vehicle guidance information. The constituent elements can be classified into arrow elements and line segment elements according to the category. The indication sign can be composed of at least one arrow element, or at least one line segment element, or also composed of at least one arrow element and at least one line segment element at the same time. For example, in the indication sign 112, its constituent elements include a straight arrow 114 and a right-turn arrow 116 (both are arrow elements), the straight arrow 114 corresponds to the information for guiding the vehicle to go straight, and the right-turn arrow 116 corresponds to the information for guiding the vehicle to turn right.

[0039] In some embodiments, the camera of the vehicle 110 directly collects visual sensing signals in front of the vehicle or in other directions, and sends the visual sensing signals to the electronic device 120, and the electronic device 120 receives the visual sensing signals and generates an image. Alternatively or additionally, the camera of the vehicle 110 (such as an intelligent camera) collects visual sensing signals in front of the vehicle or in other directions, then generates an image based on the visual sensing signals, and sends the generated image to the electronic device 120, and the electronic device 120 receives the image.

[0040] At step 204, the electronic device 120 determines the positions and categories of the key points of the constituent elements based on the image, and the positions of the key points are used to identify the area of the indication sign. In some embodiments, the electronic device 120 analyzes the image to obtain the positions and categories of the key points of the constituent elements in the indication sign. Among them, the key point of the constituent element refers to the point that identifies the constituent element, such as the vertex of the straight arrow 114.

[0041] In some embodiments, the positions of the key points can be used to identify the areas of the corresponding constituent elements. Therefore, when there is one constituent element, the area of the indication sign can be directly identified based on the position of the key point. When there are multiple constituent elements, the area of the indication sign can be identified based on the combination of the positions of the key points. The category of the key point, as the semantic information of the key point, can be used to identify the category of the corresponding constituent element. For example, the semantics corresponding to the vertex of the straight arrow 114 can be "the vertex of the straight arrow". The following will be combined with Figures 3A to 3G to illustrate the positions and categories of the key points.

[0042] In some embodiments, the image can be detected by a neural network model to determine the positions and categories of the key points of the constituent elements in the indication sign. Compared with traditional image processing schemes such as designing handcrafted features, in this way, the limitation of only being able to detect the objects corresponding to the handcrafted features can be avoided, and the influence of interference items of other signs on the road on the detection result can be avoided, thereby improving the accuracy of the indication sign detection. The following will be combined with Figure 4 to illustrate the process of detecting the key points based on the neural network model.

[0043] In some embodiments, the categories of the key points include visible points and invisible points. Among them, visible points include directly observable points such as the vertex of the straight arrow and the corner point of the left-turn arrow, and invisible points are the occluded key points. After the electronic device 120 obtains the image, it detects whether there is an occlusion defect in the constituent elements in the indication sign. If it detects that there is an occlusion area in the constituent elements in the indication sign, it further determines whether there is a key point of the constituent element in the occlusion area. If there is a key point in the occlusion area, it determines the position and category of the key point, where the category is the invisible point. In some embodiments, the key points in the occlusion area can be determined based on inference. By virtual labeling the key points in this way, the comprehensiveness of the key points can be ensured, and thus the complete decoding of the semantics of the key point categories can be ensured.

[0044] At step 206, the electronic device 120 determines whether the position of the key point meets the position condition. Among them, the position condition is used to determine whether the key point belongs to the current indication sign.

[0045] In some embodiments, the electronic device 120 analyzes an image to determine the bounding box of the indication sign and the area of the bounding box in the image. Then, based on the position of the key point and the area of the bounding box, it is determined whether the key point is within the bounding box. If the key point is within the bounding box, it is determined that the key point belongs to the current indication sign, and then the semantics of the indication sign are determined based on the category of the key point. In this way, it can be ensured that each key point belongs to the current indication sign, thereby improving the accuracy of the key points.

[0046] In some embodiments, the bounding box of the indication sign is a rotated rectangular box. By detecting the image, the rotated rectangular box circumscribing the indication sign in the image is obtained. The parameters of the rotated rectangular box at least include the center point position, size parameters, and angle. Based on the parameters of the rotated rectangular box, the area of the rotated rectangular box in the image can be determined. Compared with the conventional non-rotated rectangular box, in this way, the accuracy of the bounding box can be improved and the invalid area can be reduced, thereby improving the accuracy of the key points.

[0047] At step 208, if the position of the key point meets the position condition, the electronic device 120 determines the semantics of the indication sign based on the category of the key point. Wherein, the semantics of the indication sign is the meaning represented by the indication sign. In some embodiments, when there are multiple constituent elements of the indication sign, the semantics of the indication sign can be determined according to the combination of the categories of the key points of the multiple constituent elements. For example, in the indication sign 112, there is a straight arrow 114 and a right-turn arrow 116. The key points of the straight arrow 114 and the right-turn arrow 116 are both vertices. The category of the vertex of the straight arrow 114 is "vertex of the straight arrow", and the category of the vertex of the right-turn arrow 116 is "vertex of the right-turn arrow". Therefore, the semantics of the indication sign 112 can be determined as "straight and right-turn sign".

[0048] Alternatively or additionally, when there is a single constituent element of the indication sign, the semantics of the indication sign can be directly determined based on the category of the key point of the constituent element. For example, in the indication sign 118, there is only one straight arrow. The key point of the straight arrow is the vertex of the straight arrow. The category of the vertex of the straight arrow is "vertex of the straight arrow". Therefore, the semantics of the indication sign 118 can be determined as "straight sign".

[0049] In an embodiment of the present disclosure, first, an image containing an indication sign collected by a sensing device such as an in-vehicle camera is obtained. The indication sign may include at least one arrow element, at least one line segment element, or a combination of at least one arrow element and at least one line segment element. Then, the image is detected to determine the positions and categories of the key points of each constituent element in the indication sign. Further, when the positions of the key points meet the position conditions, based on the categories of the key points of each constituent element in the indication sign, the semantics of the indication sign are determined (for example, the semantics of the indication sign may be the compositional information of the categories of the key points of each constituent element).

[0050] In this way, the positions and categories of each key point of the constituent elements in the indication sign can be directly calculated based on the image. The position of the key point can be directly used to identify the area of the indication sign (the position of the key point can be used to identify the area of the corresponding constituent element. When there is one constituent element, the area of the indication sign can be directly identified based on the position of the key point. When there are multiple constituent elements, the area of the indication sign can be identified based on the combination of the positions of the key points), and the area of the indication sign can be quickly determined through the position of the key point, avoiding information redundancy caused by identifying the indication sign by means of a bounding box, etc. (especially when there are many constituent elements in the indication sign, the bounding boxes will overlap a lot), and at the same time, the number of parameters in the process of determining the area of the indication sign can be reduced. The category of the key point, as the semantic information of the key point, can improve the richness of the information of the key point, and then the semantics of the indication sign can be directly determined based on the category of the key point, avoiding inaccurate and incomplete results caused by directly classifying the entire indication sign. Therefore, the method of the present disclosure can improve the accuracy and speed of indication sign detection and reduce redundant information.

[0051] Figure 3A FIG. shows a schematic diagram of an indication sign 300A including a lower left corner point according to some embodiments of the present disclosure. Among them, the indication sign 300A includes a straight arrow sign, a left turn arrow sign, a right turn arrow sign, a straight and left turn arrow sign, a straight and right turn arrow sign, a U-turn arrow sign, a straight and U-turn arrow sign, and a left turn and U-turn arrow sign. Refer to Figure 3A, 302A shows a straight arrow sign including the bottom left corner point, 304A shows a left turn arrow sign including the bottom left corner point, 306A shows a schematic diagram of a right turn arrow sign including the bottom left corner point, Figure 308A shows a straight and left turn arrow sign including the bottom left corner point (this bottom left corner point serves as the bottom left corner point of both the straight arrow and the left turn arrow), 310A shows a straight and right turn arrow sign including the bottom left corner point (this bottom left corner point serves as the bottom left corner point of both the straight arrow and the right turn arrow), 312A shows a U-turn arrow sign including the bottom left corner point, 314A shows a straight and U-turn arrow sign including the bottom left corner point (this bottom left corner point serves as the bottom left corner point of both the straight arrow and the U-turn arrow), 316A shows a left turn and U-turn arrow sign including the bottom left corner point (this bottom left corner point serves as the bottom left corner point of both the left turn arrow and the U-turn arrow).

[0052] Figure 3B shows a schematic diagram of an indication sign 300B including the bottom right corner point of some embodiments of the present disclosure. Among them, the indication sign 300B includes a straight arrow sign, a left turn arrow sign, a right turn arrow sign, a straight and left turn arrow sign, a straight and right turn arrow sign, a U-turn arrow sign, a straight and U-turn arrow sign, and a left turn and U-turn arrow sign. Refer to Figure 3B , 302B shows a straight arrow sign including the bottom right corner point, 304B shows a left turn arrow sign including the bottom right corner point, 306B shows a right turn arrow sign including the bottom right corner point, 308B shows a straight and left turn arrow sign including the bottom right corner point (this bottom right corner point serves as the bottom right corner point of both the straight arrow and the left turn arrow), 310B shows a straight and right turn arrow sign including the bottom right corner point (this bottom right corner point serves as the bottom right corner point of both the straight arrow and the right turn arrow), 312B shows a U-turn arrow sign including the bottom right corner point, Figure 314B shows a straight and U-turn arrow sign including the bottom right corner point (this bottom right corner point serves as the bottom right corner point of both the straight arrow and the U-turn arrow), 316B shows a left turn and U-turn arrow sign including the bottom right corner point (this bottom right corner point serves as the bottom right corner point of both the left turn arrow and the U-turn arrow).

[0053] Figure 3C shows a schematic diagram of an indication sign 300C including the vertex of some embodiments of the present disclosure. Among them, the indication sign 300C includes a straight arrow sign, a left turn arrow sign, a right turn arrow sign, a straight and left turn arrow sign, a straight and right turn arrow sign, a U-turn arrow sign, a straight and U-turn arrow sign, and a left turn and U-turn arrow sign. Refer to Figure 3C, 302C shows a straight arrow sign containing a vertex, 304C shows a left-turn arrow sign containing a vertex, 306C shows a right-turn arrow sign containing a vertex, 308C shows a straight and left-turn arrow sign containing a vertex, 310C shows a straight and right-turn arrow sign containing a vertex, 312C shows a U-turn arrow sign containing a vertex, 314C shows a straight and U-turn arrow sign containing a vertex, 316C shows a left-turn and U-turn arrow sign containing a vertex.

[0054] Figure 3D Shows a schematic diagram of the indication sign 300D including inflection points according to some embodiments of the present disclosure. The indication sign 300D includes a U-turn arrow sign, a straight and U-turn sign, and a left-turn and U-turn sign. Fig. 302D shows a schematic diagram of a U-turn arrow sign including an inflection point, Fig. 304D shows a schematic diagram of a straight and U-turn arrow sign including an inflection point (the inflection point is the inflection point of the U-turn arrow), and Fig. 306D shows a schematic diagram of a left-turn and U-turn arrow sign including an inflection point (the inflection point is the inflection point of the U-turn arrow).

[0055] Figure 3E Shows a schematic diagram of the indication sign 300E including vertices and corner points according to some embodiments of the present disclosure. The indication sign 300E includes a merge-right arrow sign and a merge-left arrow sign. Refer to Figure 3E , 302E shows a merge-right arrow sign including a vertex, a lower-left corner point, and a lower-right corner point, and 304E shows a merge-left arrow sign including a vertex, a lower-left corner point, and a lower-right corner point.

[0056] Figure 3F Shows a schematic diagram of the indication sign 300F including end points according to some embodiments of the present disclosure. Refer to Figure 3F , the indication sign 300F includes a dotted line sign and a prohibition sign. 302F shows a line segment in the dotted line sign containing two end points, and 304F shows a prohibition sign containing four end points.

[0057] Figure 3G Shows a schematic diagram of the indication sign 300G including multiple key points according to some embodiments of the present disclosure. Refer to Figure 3G, the indication sign 300G is a straight-ahead and left-turn arrow sign, including two constituent elements: a straight-ahead arrow and a left-turn arrow. The electronic device 120 detects the image containing the indication sign 300G, thereby obtaining 4 key points, namely the vertex of the straight-ahead arrow, the vertex of the left-turn arrow, the lower left corner point of the straight-ahead arrow / left-turn arrow, and the lower right corner point of the straight-ahead arrow / left-turn arrow. The positions of the above 4 key points can be used to identify the area of the indication sign 300G (thereby replacing the bounding box). The categories of the above four key points are respectively "the vertex of the straight-ahead arrow", "the vertex of the left-turn arrow", "the lower left corner point of the straight-ahead arrow / left-turn arrow", and "the lower right corner point of the straight-ahead arrow / left-turn arrow". By combining the category information of the key points, the semantics of the indication sign 300G can be obtained as "straight-ahead and left-turn sign".

[0058] Figure 4 FIG. shows a schematic diagram of a process 400 for detecting key points in an indication sign according to some embodiments of the present disclosure. In some embodiments, in Figure 1 the example environment 100 shown, the process 400 can be executed by the electronic device 120. It should be understood that the process 400 may further include additional actions not shown and / or the actions shown may be omitted, and the scope of the present disclosure is not limited in this regard.

[0059] In some embodiments, referring to Figure 4 , the key points in the indication sign are detected based on a neural network model. Among them, the neural network model includes a backbone network 404, a neck network 406, a head network 410 (which can be referred to as the second head network), and a head network 418 (which can be referred to as the first head network). Among them, the backbone network 404 and the neck network 406 are used to extract the features of the image, the head network 410 is used to predict the parameters of the rotated rectangular box circumscribing the indication sign, and the head network 418 is used to predict the positions and categories of the key points.

[0060] In some embodiments, the size of the image 402 is 600*600*3. The backbone network 404 contains four convolutional layers. After the image 402 passes through the four convolutional layers of the backbone network 404, the sizes of the obtained image features are successively 135*240*32, 68*120*64, 34*60*160, and 17*30*384. The neck network 406 also contains four convolutional layers. After the 17*30*384 image features output by the backbone network 404 successively pass through the four convolutional layers of the neck network 406, the sizes of the obtained image features are successively 17*30*384, 34*60*256, 68*120*128, and 135*240*64, where the 135*240*64 image features are the feature map 408.

[0061] In some embodiments, the electronic device 120 acquires the image 402, then inputs the image 402 into the backbone network 404 of the neural network model to obtain the encoded image features, and then inputs the image features into the neck network 406 for decoding to obtain the feature map 408. After obtaining the feature map 408, the electronic device 120 can process the feature map 408 based on the head network 410 to obtain the parameters of the rotated rectangular box. At the same time, the electronic device 120 can also process the feature map 408 based on the head network 418 to obtain the positions and categories of the key points.

[0062] In some embodiments, the electronic device 120 processes the feature map 408 based on the head network 410 to obtain the parameters of the rotated rectangular box (for determining the region of the rotated rectangular box in the image), where the parameters of the rotated rectangular box include the center point position, size parameters, and angle. Then, according to the positions of the key points and the parameters of the rotated rectangular box, it is determined whether the key points are within the rotated rectangular box. If the key points are within the rotated rectangular box, it indicates that the key points belong to the current indication sign, so the semantics of the indication sign can be determined based on the categories of the key points.

[0063] In some embodiments, during the process of training the neural network model, the offset of the center point position of the rotated rectangular box (which can be called the second offset) can be determined based on the head network 410, and the offset of the key point position (which can be called the first offset) can be determined based on the head network 418, where the offset of the center point position and the offset of the key point position are used to learn the loss during the encoding process of the image 402. Then, based on the offset of the center point position and the offset of the key point position, the parameters in the neural network model are adjusted so that the offset of the center point position and the offset of the key point position meet the convergence condition. In this way, the accuracy of the center point position and the key point position can be improved.

[0064] In some embodiments, the head network 410 includes four branches. The first branch includes the convolutional layer 412-1 and the convolutional layer 414-1, with sizes of 3*3*64 and 1*1*2 respectively, and the output of the first branch is the center point position 416-1. The second branch includes the convolutional layer 412-2 and the convolutional layer 414-2, with sizes of 3*3*64 and 1*1*2 respectively, and the output of the second branch is the width and height parameters 416-2. The third branch includes the convolutional layer 412-3 and the convolutional layer 414-3, with sizes of 3*3*64 and 1*1*2 respectively, and the output of the third branch is the offset of the center point position 416-3. The fourth branch includes the convolutional layer 412-4 and the convolutional layer 414-4, with sizes of 3*3*64 and 1*1*2 respectively, and the output of the fourth branch is the angle 416-4 of the rotated rectangular box.

[0065] In the current embodiment, during the training process of the neural network model, multiple loss functions can be set to train the neural network model. In the first branch, the loss of the center point position can be determined based on the Gaussian Focal Loss function. In the second branch, the loss of the width and height parameters can be determined based on the L1 Loss function. In the third branch, the loss of the offset of the center point position can be determined based on the L1 Loss function. In the fourth branch, the loss of the angle of the rotated rectangular box can be determined based on the L1 Loss function. The loss of the neural network model during the training process is determined by the above loss functions, and then the parameters in the neural network model are adjusted to make the loss converge.

[0066] In some embodiments, the head network 418 includes two branches. The first branch includes a convolutional layer 420-1 and a convolutional layer 422-1, with sizes of 3*3*64 and 1*1*11 respectively. After passing through the two convolutional layers, the result 424-1 is obtained, and the result 424-1 includes the position and category of the key points. Then, the electronic device 120 executes step 426, where the key points are matched with the rotated rectangular box to determine whether the key points are within the area of the rotated rectangular box, so as to obtain the matching result 428. If the matching result 428 is within the area of the rotated rectangular box, step 430 is executed. When the key point is a visible point, the position and category of the key point are output, and when the key point is an invisible point, the category of the key point is output. If the matching result 428 is not within the area of the rotated rectangular box, step 432 is executed to determine not to output the position and category of the key points. The second branch includes a convolutional layer 420-2 and a convolutional layer 422-2, with sizes of 3*3*64 and 1*1*2 respectively. The output of the second branch is the offset 424-2 of the position of the key points.

[0067] In the current embodiment, during the training process of the neural network model, multiple loss functions can be set to train the neural network model. In the first branch, the loss of the position of the key points is determined based on the Gaussian Focal Loss function. In the second branch, the loss of the offset of the position of the key points can be determined based on the L1 Loss function. The loss of the neural network model during the training process is determined by the above loss functions, and then the parameters in the neural network model are adjusted to make the loss converge.

[0068] By means of a neural network model, the accuracy of key points and rotated rectangular boxes can be improved, thereby enhancing the accuracy of detecting indicating signs. Compared with the traditional manual feature method, the neural network model can be trained based on a large amount of diverse training data of indicating signs, ensuring the accuracy of indicating sign prediction and the adaptability to different scenarios, expanding the scope of indicating sign detection, and improving the efficiency of indicating sign detection. Especially in scenarios where indicating signs are not standardized, such as parking lots, the method based on the neural network model in this embodiment has high robustness.

[0069] Figure 5 FIG. shows a block diagram of a device 500 for detecting indicating signs according to some embodiments of the present disclosure. Referring to Figure 5 , the device 500 includes an image acquisition module 502 configured to acquire an image containing an indicating sign, and the constituent elements of the indicating sign include at least one of an arrow element and a line segment element. The device 500 further includes a key point determination module 504 configured to determine the position and category of the key points of the constituent elements based on the image. In addition, the device 500 further includes an indicating sign determination module 506 configured to determine the semantics of the indicating sign based on the category of the key points in response to the position of the key points satisfying the position condition.

[0070] In some embodiments, the indicating sign determination module 506 is further configured to determine the region of the bounding box of the indicating sign based on the image; determine whether the key points are within the bounding box based on the position of the key points and the region of the bounding box; and determine the semantics of the indicating sign based on the category of the key points in response to the key points being within the bounding box.

[0071] In some embodiments, the bounding box is a rotated rectangular box, and the indicating sign determination module 506 is further configured to: determine the parameters of the rotated rectangular box based on the image, where the parameters of the rotated rectangular box at least include the center point position, size parameters, and angle; and determine the region of the rotated rectangular box based on the parameters of the rotated rectangular box.

[0072] In some embodiments, the category of the key points includes at least one of the following: the lower left corner point of a straight arrow, the lower right corner point of a straight arrow, the vertex of a straight arrow; the lower left corner point of a turning arrow, the lower right corner point of a turning arrow, the vertex of a turning arrow; the lower left corner point of a U-turn arrow, the lower right corner point of a U-turn arrow, the vertex of a U-turn arrow, the inflection point of a U-turn arrow; and the end points of a line segment element.

[0073] In some embodiments, the category of key points further includes invisible points, and the key point determination module 504 is further configured to: determine whether there is an occlusion area in the constituent element; in response to the existence of an occlusion area in the constituent element, determine whether there is a key point in the occlusion area; and in response to the existence of a key point in the occlusion area, determine the position and category of the key point based on the image, and the category of the key point is an invisible point.

[0074] In some embodiments, the key point determination module 504 is configured to: determine a corresponding feature map based on the image by the backbone network and the neck network of the trained neural network model; and determine the position and category of the key point based on the feature map by the first head network of the neural network model.

[0075] In some embodiments, the indication flag determination module 506 is further configured to: determine the parameters of the rotated rectangular box circumscribing the indication flag based on the feature map by the second head network of the neural network model, and the parameters of the rotated rectangular box at least include the center point position, the size parameter, and the angle; determine whether the key point is within the rotated rectangular box based on the position of the key point and the parameters of the rotated rectangular box; and in response to the key point being within the rotated rectangular box, determine the semantics of the indication flag based on the category of the key point.

[0076] In some embodiments, it further includes a neural network model training module, which is configured to: determine a first offset corresponding to the position of the key point based on the first head network; determine a second offset corresponding to the center point position based on the second head network; and adjust the parameters in the neural network model based on the first offset and the second offset so that the first offset and the second offset meet the convergence condition.

[0077] It can be understood that the device 500 of the present disclosure can achieve at least one of the many advantages that the methods or processes described above can achieve. For example, the device 500 can directly calculate the position and category of each key point of the constituent elements in the indication sign based on the image. The position of the key point can be directly used to identify the area of the indication sign (the position of the key point can be used to identify the area of the corresponding constituent element. When there is one constituent element, the area of the indication sign can be directly identified based on the position of the key point. When there are multiple constituent elements, the area of the indication sign can be identified based on the combination of the positions of the key points). The area of the indication sign can be quickly determined through the position of the key point, and information redundancy caused by identifying the indication sign by means of a bounding box, etc. can be avoided (especially when there are many constituent elements in the indication sign, the bounding boxes will overlap a lot), and at the same time, the number of parameters in the process of determining the area of the indication sign can be reduced. The category of the key point, as the semantic information of the key point, can improve the richness of the information of the key point, and thus the semantics of the indication sign can be directly determined based on the category of the key point, avoiding inaccurate and incomplete results caused by directly classifying the entire indication sign. Therefore, the method of the present disclosure can improve the accuracy and speed of indication sign detection and reduce redundant information.

[0078] Figure 6 FIG. shows a schematic block diagram of an exemplary device 600 that can be used to implement embodiments of the present disclosure. As Figure 6 shown, the device 600 includes a processor 601, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 602 and loaded into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0079] Each of the processes and processes described above, such as the method 200, can be executed by the processor 601. For example, in some embodiments, the method 200 can be implemented as a computer software program, which is tangibly included in a machine-readable medium. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602. When the computer program is loaded into the RAM 603 and executed by the processor 601, one or more actions of the method 200 described above can be executed.

[0080] The present disclosure can be a method, a device, a system, and / or a computer program product. The computer program product can include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.

[0081] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0082] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0083] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0084] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.

[0085] These computer - readable program instructions can be provided to a processing unit of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data - processing apparatus, a device is created that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0086] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0087] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0088] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for detecting an indication sign, comprising: Obtaining an image containing the indication sign, wherein the constituent elements of the indication sign include at least one of an arrow element and a line segment element; Based on the image, determining the positions and categories of the key points of the constituent elements; And In response to the position of the key points satisfying a position condition, based on the categories of the key points, determining the semantics of the indication sign.

2. The method according to claim 1, wherein determining the semantics of the indication sign based on the categories of the key points includes: Based on the image, determining the area of the bounding box of the indication sign; Based on the positions of the key points and the area of the bounding box, determining whether the key points are within the bounding box; And In response to the key points being within the bounding box, based on the categories of the key points, determining the semantics of the indication sign.

3. The method according to claim 2, wherein the bounding box is a rotated rectangular box, and determining the area of the bounding box of the indication sign based on the image includes: Based on the image, determining the parameters of the rotated rectangular box, wherein the parameters of the rotated rectangular box at least include the center point position, the size parameters, and the angle; And Based on the parameters of the rotated rectangular box, determining the area of the rotated rectangular box.

4. The method according to claim 1, wherein the categories of the key points include at least one of the following: The lower left corner point of a straight arrow, the lower right corner point of the straight arrow, the vertex of the straight arrow; The lower left corner point of a turning arrow, the lower right corner point of the turning arrow, the vertex of the turning arrow; The lower left corner point of a U-turn arrow, the lower right corner point of the U-turn arrow, the vertex of the U-turn arrow, the inflection point of the U-turn arrow; And The end points of a line segment element.

5. The method according to claim 4, wherein the categories of the key points further include invisible points, and determining the positions and categories of the key points of the constituent elements based on the image includes: Determining whether there is an occlusion area in the constituent element; In response to the constituent element having the occlusion area, determining whether there is a key point in the occlusion area; And In response to there being a key point in the occlusion area, based on the image, determining the position and the category of the key point, and the category of the key point is the invisible point.

6. The method according to claim 1, wherein determining the positions and categories of the key points of the constituent elements based on the image includes: Based on the image, the backbone network and the neck network of a trained neural network model determining a corresponding feature map; And Based on the feature map, the first head network of the neural network model determining the positions and the categories of the key points.

7. The method according to claim 6, wherein determining the semantics of the indication sign based on the categories of the key points includes: Based on the feature map, the second head network of the neural network model determining the parameters of a rotated rectangular box circumscribing the indication sign, wherein the parameters of the rotated rectangular box at least include the center point position, the size parameters, and the angle; Determine whether the key point is within the rotated rectangular box based on the position of the key point and the parameters of the rotated rectangular box; and In response to the key point being within the rotated rectangular box, determine the semantics of the indication flag based on the category of the key point.

8. The method according to claim 7, wherein the training method of the neural network model comprises: Determine a first offset corresponding to the position of the key point based on the first head network; Determine a second offset corresponding to the center point position based on the second head network; and Based on the first offset and the second offset, adjust the parameters in the neural network model so that the first offset and the second offset meet the convergence condition.

9. A device for detecting an indication flag, comprising: An image acquisition module configured to acquire an image containing the indication flag, the constituent elements of the indication flag including at least one of an arrow element and a line segment element; A key point determination module configured to determine the position and category of the key points of the constituent elements based on the image; and An indication flag determination module configured to determine the semantics of the indication flag based on the category of the key point in response to the position of the key point satisfying a position condition.

10. A controller, comprising: At least one processor; and A memory coupled to the at least one processor and having instructions stored thereon, the instructions, when executed by the at least one processor, cause the controller to execute the method according to any one of claims 1-8.

11. A vehicle comprising the controller according to claim 10.

12. A machine-readable storage medium having machine-executable instructions stored thereon, wherein the machine-executable instructions are executed by a processor to implement the method according to any one of claims 1 to 8.