Semantic positioning method and device, electronic equipment and storage medium

By loading a local semantic map onto the vehicle and constructing a global perspective image using a vehicle surround-view camera, the problems of semantic map update timeliness and memory overflow are solved, thereby improving vehicle positioning accuracy and update efficiency.

CN121453075APending Publication Date: 2026-02-03CHINA FAW CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511788243.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

The existing semantic map update method relies on external communication, which results in poor timeliness in weak signal environments, affecting vehicle positioning accuracy. Furthermore, loading large-size maps may cause memory overflow in the vehicle's infotainment system.

Method used

A local semantic map is loaded based on a global semantic map, and local view images are obtained through vehicle surround view cameras to construct a global view image. The map is updated based on the pixel coordinates of the vehicle's predicted pose and semantic features, and the position and type of semantic elements are optimized.

Benefits of technology

It enables real-time updates of the global semantic map, avoiding memory overflow and improving vehicle positioning accuracy and update timeliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121453075A_ABST
    Figure CN121453075A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic positioning method and device, electronic equipment and a storage medium, and relates to the field of automobile driving and environment recognized.The method comprises the steps that a local semantic map is loaded based on a global semantic map, and pixel coordinates of all semantic elements in the local semantic map are converted into map coordinates; constructing a global visual angle image, and obtaining pixel coordinates of each semantic feature through the global visual angle image; acquiring map coordinates of each semantic feature according to the predicted pose of the vehicle at the current moment and the pixel coordinates of each semantic feature; and updating the global semantic map according to the map coordinates of the semantic features and the map coordinates of the semantic elements. According to the technical scheme provided by the embodiment of the invention, the memory overflow phenomenon of the vehicle machine equipment is avoided, the semantic element updating of the global semantic map is realized based on the semantic features extracted from the global view image, the updating timeliness of the global semantic map is ensured, and the positioning precision of the vehicle is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automobile driving and environmental recognition, and more particularly to a semantic localization method, device, electronic device, and storage medium. Background Technology

[0002] Semantic maps are intelligent navigation maps that overlay semantic information such as lane markings on top of traditional maps. With the continuous development of driver assistance technology, semantic positioning based on semantic maps has also come into people's view.

[0003] Since semantic localization uses stable semantic features for matching and localization, it has the characteristics of high localization stability and high accuracy, and is therefore suitable for vehicle assisted driving in closed environments (e.g., parking lots). In the prior art, the vehicle obtains a semantic map by communicating with a backend server and sends the environmental information detected by the vehicle camera to the backend server so that the semantic map can be updated through the backend server.

[0004] However, parking lot coverage areas vary in size, and loading large-sized semantic maps often leads to memory overflow in vehicle-mounted devices. At the same time, existing semantic map update methods rely on external communication, which often results in poor timeliness of semantic map updates in weak signal environments, thus affecting the vehicle's positioning accuracy. Summary of the Invention

[0005] This invention provides a semantic localization method, apparatus, electronic device, and storage medium to solve the problem of poor timeliness in semantic map updates.

[0006] According to another aspect of the present invention, a semantic localization method is provided, comprising:

[0007] A local semantic map is loaded based on a global semantic map, and the pixel coordinates of each semantic element in the local semantic map are converted into map coordinates;

[0008] A global view image is constructed based on the local view images obtained by the vehicle surround view camera, and the pixel coordinates of each semantic feature are obtained through the global view image.

[0009] Based on the vehicle's predicted pose at the current moment and the pixel coordinates of each semantic feature, obtain the map coordinates of each semantic feature;

[0010] The global semantic map is semantically updated based on the map coordinates of each semantic feature and the map coordinates of each semantic element.

[0011] The process of loading a local semantic map based on a global semantic map includes: determining the local area range based on vehicle speed and memory usage, and loading the local semantic map based on the global semantic map according to the local area range.

[0012] The step of updating the global semantic map based on the map coordinates of each semantic feature and the map coordinates of each semantic element includes: if a first semantic element of the same semantic type as the first semantic feature is detected within a preset interval distance of the first semantic feature, then the first semantic feature is determined to be a valid semantic feature; if only a second semantic element of a different semantic type than the second semantic feature is detected within a preset interval distance of the second semantic feature, then the second semantic feature is determined to be a conflicting semantic feature; if no semantic element is detected within a preset interval distance of the third semantic feature, then the third semantic feature is determined to be a newly added semantic feature.

[0013] The step of semantically updating the global semantic map based on the map coordinates of each semantic feature and the map coordinates of each semantic element further includes: constructing an error equation based on the map coordinates of each effective semantic feature and the map coordinates of the effective semantic elements that match each effective semantic feature, and optimizing the error equation to obtain the optimized vehicle pose at the current moment.

[0014] After determining that the second semantic feature is a conflict semantic feature, the method further includes: incrementing the number of semantic type detections corresponding to the second semantic element according to the semantic type of the second semantic feature; taking the first semantic type with the most detections corresponding to the second semantic element as the semantic type of the second semantic element; wherein the number of detections of the first semantic type is greater than or equal to the first detection threshold; and / or after determining that the third semantic feature is a new semantic feature, the method further includes: obtaining a target pixel that matches the third semantic feature, and incrementing the number of semantic type detections corresponding to the target pixel according to the semantic type of the third semantic feature; taking the second semantic type with the most detections corresponding to the target pixel as the semantic type of the target pixel; wherein the number of detections of the second semantic type is greater than or equal to the second detection threshold.

[0015] The semantic localization method further includes: if the detected distance between the vehicle location and any boundary of the current local semantic map is less than or equal to a preset distance threshold, reloading the local semantic map based on the global semantic map.

[0016] According to another aspect of the present invention, a semantic positioning device is provided, comprising:

[0017] The local semantic map acquisition module is used to load a local semantic map based on the global semantic map and convert the pixel coordinates of each semantic element in the local semantic map into map coordinates.

[0018] The global view image acquisition module is used to construct a global view image based on the local view image acquired by the vehicle surround view camera, and to obtain the pixel coordinates of each semantic feature through the global view image;

[0019] The map coordinate acquisition module is used to acquire the map coordinates of each semantic feature based on the vehicle's predicted pose at the current time and the pixel coordinates of each semantic feature.

[0020] The map update execution module is used to perform semantic updates on the global semantic map based on the map coordinates of each semantic feature and the map coordinates of each semantic element.

[0021] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the semantic localization method according to any embodiment of the present invention.

[0022] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the semantic localization method described in any embodiment of the present invention.

[0023] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the semantic localization method described in any embodiment of the present invention.

[0024] The technical solution of this invention loads a local semantic map based on a global semantic map, and converts the pixel coordinates of each semantic element in the local semantic map into map coordinates; constructs a global perspective image based on local view images obtained from the vehicle surround view camera, and obtains the pixel coordinates of each semantic feature through the global perspective image; obtains the map coordinates of each semantic feature based on the vehicle's predicted pose at the current moment and the pixel coordinates of each semantic feature; and updates the global semantic map based on the map coordinates of each semantic feature and the map coordinates of each semantic element. Thus, by loading the local semantic map, memory overflow in the vehicle's infotainment system is avoided; and by updating the semantic elements of the global semantic map based on the semantic features extracted from the global perspective image, the timeliness of the global semantic map update is ensured, and the vehicle's positioning accuracy is improved.

[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart of a semantic localization method provided in Embodiment 1 of the present invention;

[0028] Figure 2 This is a flowchart of another semantic localization method provided in Embodiment 2 of the present invention;

[0029] Figure 3 This is a flowchart of another semantic localization method provided in Embodiment 3 of the present invention;

[0030] Figure 4 This is a schematic diagram of the structure of a semantic positioning device according to Embodiment 4 of the present invention;

[0031] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the semantic positioning method of this invention. Detailed Implementation

[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0034] Example 1

[0035] Figure 1 This is a flowchart of a semantic localization method provided in Embodiment 1 of the present invention. This embodiment is applicable to updating semantic elements in a global semantic map through semantic features in a global view image. This method can be executed by a semantic localization device, which can be implemented in hardware and / or software, and can be configured in in-vehicle electronic devices such as vehicle infotainment systems. Figure 1 As shown, the method includes:

[0036] S101. Load a local semantic map based on the global semantic map, and convert the pixel coordinates of each semantic element in the local semantic map into map coordinates.

[0037] A global semantic map refers to a complete semantic map of the current driving scenario. Taking a parking lot as an example, since the coverage area of ​​a parking lot is usually large, the global semantic map is usually a large-sized image. For example, taking a physical resolution of 0.1 meters per pixel (i.e., 1 pixel equals 0.1 meters) as an example, since the map size is proportional to the area of ​​the parking lot, a parking lot of 1000 meters × 1000 meters corresponds to an image of 10000 pixels × 10000 pixels.

[0038] If the vehicle's infotainment system loads the entire global semantic map, issues such as memory overflow and positioning program interruption may occur. Therefore, memory-mapped file technology can be used to dynamically load the semantic map (i.e., a local semantic map) within a surrounding area (e.g., 50m x 50m) centered on the vehicle, and release the old area. This satisfies the current semantic positioning needs of the vehicle while avoiding loading the global semantic map, preventing memory overflow and other problems, and improving positioning efficiency.

[0039] Optionally, in this embodiment of the invention, loading a local semantic map based on a global semantic map includes: determining a local area range based on vehicle speed and memory usage, and loading a local semantic map based on the global semantic map according to the local area range. Vehicle speed can be detected by wheel speed sensors. When the vehicle speed is high, the local area range can be configured to be larger to ensure the local semantic map covers a larger area and reduces the update frequency of the local semantic map; when the vehicle speed is low, the local area range can be configured to be smaller so that the local semantic map only covers a smaller area, improving the display accuracy of the local semantic map.

[0040] When memory usage is high, the local region range can be configured to be smaller so that the local semantic map only covers a smaller area, reducing the memory resource consumption of the local semantic map; when memory usage is low, the local region range can be configured to be larger so that the local semantic map covers a larger area, ensuring the integrity of the semantic features of the local semantic map.

[0041] The top left corner of the global semantic map is defined as the origin (in meters) of the map coordinate system (i.e., the physical coordinate system). The current vehicle's map coordinates are: The current map coordinates of the vehicle, i.e., its current position, can be obtained through the navigation system or by calculating the vehicle's trajectory; the pixel coordinates (i.e., pixel index) of the current vehicle in the global semantic map are... It can be obtained by calculation in the following form:

[0042] ;

[0043] Then, using the current vehicle's pixel coordinates Centered on a local semantic map of the surrounding area, taking a 50m x 50m local semantic map as an example in the above technical solution, the actual pixel range of the local semantic map is 500 pixels x 500 pixels. Therefore, the horizontal coordinate range of the local semantic map is from... From The range of the horizontal axis is from From ;in, and These represent the pixel coordinates of the current vehicle. The x and y coordinates of the rectangle are defined as follows: the pixel coordinates of each pixel in the rectangle are defined as follows: x and y. The pixel coordinates described above are converted into physical coordinates using the following equation. .

[0044] ;

[0045] ;

[0046] Therefore, each pixel in the local semantic map can be converted into map coordinates in the map coordinate system. And save it as the map data structure of the current scene.

[0047] Semantic elements can include lane lines, parking lines, speed bumps, arrows, sidewalks, and grid lines, etc. Each semantic element exists in the form of a pixel, and each pixel includes pixel coordinates and a pixel value. The pixel coordinates reflect the location of the semantic element, and the pixel value reflects the type of the semantic element. Obviously, different types of semantic features have different pixel values. For example, the pixel value of a lane line is 1, the pixel value of a parking line is 2, and the pixel value of a speed bump is 3. Based on this, the pixel coordinates of each semantic element in the local semantic map are converted into map coordinates in the above manner.

[0048] S102. Construct a global perspective image based on the local perspective image obtained by the vehicle surround view camera, and obtain the pixel coordinates of each semantic feature through the global perspective image.

[0049] Multiple fisheye cameras are deployed around the vehicle body, serving as surround-view cameras. Each fisheye camera captures a local image from the current perspective (i.e., a local view image). Then, using image stitching methods, such as perspective transformation methods or deep learning-based panoramic stitching models, the various local images are stitched together to obtain a stitched global image (i.e., a global view image). For example, the global view image can cover an area of ​​10-20 meters around the vehicle body.

[0050] The semantic recognition model can include a convolutional neural network model or a semantic segmentation model. It extracts semantic features from the global view image. The semantic recognition model can be deployed in an in-vehicle embedded platform or in-vehicle device to ensure low inference latency. The semantic recognition model takes the global view image as input and outputs a binary matrix of the same size as the input image. Each pixel in the binary matrix is ​​assigned a label, which also indicates the semantic type of the current semantic feature, thereby achieving pixel-level localization of the region.

[0051] S103. Based on the vehicle's predicted pose at the current moment and the pixel coordinates of each semantic feature, obtain the map coordinates of each semantic feature.

[0052] Vehicle pose consists of position and orientation; where orientation can be in the form of a rotation matrix to represent the vehicle's facing direction; assuming the first... The vehicle position at that moment is , Indicates the first Vehicle location at any given time; Indicates the first Vehicle attitude at any given moment; Vehicle speed at any given time This can be obtained through wheel speed sensors, the first Vehicle angular velocity at time This can be obtained from the vehicle's gyroscope, based on which the first... Vehicle status at any time It can be obtained by calculating using the following equation:

[0053] ;

[0054] ;

[0055] in, Indicates the first Vehicle location at any given time Indicates the first Vehicle posture at any given moment; This represents the sampling interval time, also known as the sampling period. Time and the The time interval between moments; This indicates that the rotational change is calculated using an exponential mapping. Since the vehicle pose at the current sampling moment is predicted by calculating the vehicle pose at the previous sampling moment, as well as the vehicle speed and angular velocity at the current sampling moment, it is actually the predicted vehicle pose.

[0056] After obtaining the vehicle's predicted pose, the pixel coordinates of the semantic features extracted from the global view image of the current frame are... The vehicle's predicted pose is transformed to the map coordinate system using the following equation:

[0057] ;

[0058] in, This indicates the vehicle's current position, which is the predicted position of the vehicle mentioned above. This represents the vehicle's current attitude, i.e., the predicted vehicle attitude mentioned above. Pixel coordinates representing semantic features Map coordinates representing semantic features.

[0059] S104. Based on the map coordinates of each semantic element and the map coordinates of each semantic feature, the global semantic map is semantically updated.

[0060] As described above, each semantic feature in the global view image of the current frame and each semantic element in the currently loaded local semantic map are transformed into the same coordinate system (i.e., map coordinate system) and represented in the form of map coordinates. Based on the map coordinates of the semantic features and semantic elements, if a semantic feature in the global view image does not have any semantic element at the same position in the local semantic map, it indicates that there is a missing semantic element in the local semantic map, which means there is a missing semantic element in the global semantic map. At this time, based on the semantic feature in the global view image, a new semantic element can be added at the same position in the global semantic map, and the speech type of the new semantic element is the same as that semantic feature.

[0061] Based on the map coordinates of semantic features and semantic elements, if a semantic feature exists in the global view image, and a semantic element exists at the same location in the local semantic map, but the semantic type of the semantic element does not match the current semantic feature (for example, the semantic feature is a lane line, and the semantic element is a parking space line), then the semantic type of the current semantic element is modified according to the semantic type of the semantic feature. That is, the semantic type of the current semantic element is replaced with the semantic type of the semantic feature, thereby completing the semantic type update of the semantic elements in the global semantic map.

[0062] Optionally, in this embodiment of the invention, the semantic localization method further includes: if the detected distance between the vehicle position and any boundary of the current local semantic map is less than or equal to a preset distance threshold, reloading the local semantic map based on the global semantic map. During vehicle operation, the distance between the vehicle position and each boundary of the current local semantic map is monitored in real time. If any distance is less than or equal to a preset distance threshold (e.g., 15 meters), it indicates that the vehicle may soon leave the area covered by the current local semantic map. In this case, the local semantic map needs to be reloaded based on the global semantic map. This avoids the vehicle not leaving the coverage area of ​​the local semantic map, thus providing a usable semantic map for the vehicle, and also avoids the memory overflow phenomenon caused by the excessively fast update frequency of the semantic map.

[0063] The technical solution of this invention loads a local semantic map based on a global semantic map, and converts the pixel coordinates of each semantic element in the local semantic map into map coordinates; constructs a global perspective image based on local view images obtained from the vehicle surround view camera, and obtains the pixel coordinates of each semantic feature through the global perspective image; obtains the map coordinates of each semantic feature based on the vehicle's predicted pose at the current moment and the pixel coordinates of each semantic feature; and updates the global semantic map based on the map coordinates of each semantic feature and the map coordinates of each semantic element. Thus, by loading the local semantic map, memory overflow in the vehicle's infotainment system is avoided; and by updating the semantic elements of the global semantic map based on the semantic features extracted from the global perspective image, the timeliness of the global semantic map update is ensured, and the vehicle's positioning accuracy is improved.

[0064] Example 2

[0065] Figure 2 This is a flowchart of a semantic localization method provided in Embodiment 2 of the present invention. The relationship between this embodiment and the above embodiments is that it determines which of the following semantic features belongs to: valid semantic features, conflicting semantic features, and newly added semantic features. Figure 2 As shown, the method specifically includes:

[0066] S201. Load a local semantic map based on the global semantic map, and convert the pixel coordinates of each semantic element in the local semantic map into map coordinates.

[0067] S202. Construct a global perspective image based on the local perspective image obtained by the vehicle surround view camera, and obtain the pixel coordinates of each semantic feature through the global perspective image.

[0068] S203. Based on the vehicle's predicted pose at the current moment and the pixel coordinates of each semantic feature, obtain the map coordinates of each semantic feature.

[0069] S204. If a first semantic element of the same semantic type as the first semantic feature is detected within a preset interval distance of the first semantic feature, then the first semantic feature is determined to be a valid semantic feature.

[0070] If a semantic feature of the same type is found within a range of intervals less than a preset threshold (e.g., 0.2 meters), it means that the current semantic feature and the semantic element are successfully matched. The current semantic feature is a valid semantic feature, and the current semantic element is a valid semantic element. For example, if the semantic type of semantic feature A is lane line, and a semantic element B that is also a lane line is found within a range of intervals less than the preset threshold, it can be determined that semantic feature A and semantic element B are matched, that is, they represent the same thing.

[0071] For a valid semantic feature, it indicates that the language type of the first language element corresponding to the valid semantic feature in the global semantic map is accurately labeled, and its original semantic type can be maintained. At the same time, since there is a gap between the valid semantic feature and the first semantic element, the position of the first semantic element can be updated to the position of the valid semantic feature, or the position of the first semantic element can be updated to the intermediate position between the first semantic element and the valid semantic feature, thereby completing the position update of the valid semantic element in the global semantic map.

[0072] S205. If, within a preset interval distance of the second semantic feature, only a second semantic element of a different semantic type than the second semantic feature is detected, then the second semantic feature is determined to be a conflicting semantic feature.

[0073] If, based on the current semantic feature, no semantic element of the same type is found within an interval distance less than a preset threshold, but a semantic element of a different type is found, it indicates that there is a semantic type conflict between the current semantic feature and the semantic element. The current semantic feature is the conflicting semantic feature, and the current semantic element is the conflicting semantic element. For example, if the semantic type of semantic feature A is lane line, and only the semantic element C with the semantic type of parking line is found within an interval distance less than a preset threshold, it can be determined that there is a conflict between semantic feature A and semantic element C.

[0074] For conflict semantic features, it indicates that the language type of the second language element corresponding to the conflict semantic feature in the global semantic map may be incorrectly labeled. The language type of the second language element can be updated to the language type of the conflict semantic feature. At the same time, since there is a gap between the conflict semantic feature and the second semantic element, the position of the second semantic element can be updated to the position of the conflict semantic feature, or the position of the second semantic element can be updated to the intermediate position between the second semantic element and the conflict semantic feature, thereby completing the position update of the conflict semantic element in the global semantic map.

[0075] S206. If no semantic element is detected within the preset interval distance of the third semantic feature, then the third semantic feature is determined to be a newly added semantic feature.

[0076] If no semantic element is found within a preset threshold range based on the current semantic feature, then the current semantic feature is considered a newly added semantic feature. For example, if the semantic type of semantic feature A is lane line, and no semantic element is found within a preset threshold range, then semantic feature A can be determined to be a newly added semantic feature. A newly added semantic feature indicates that there may be missing semantic elements in the global semantic map. In this case, based on the newly added semantic feature, a semantic element of the same semantic type is added to the same location in the global semantic map, thereby updating the global semantic map with the newly added semantic element.

[0077] Optionally, in this embodiment of the invention, the step of semantically updating the global semantic map based on the map coordinates of each semantic feature and the map coordinates of each semantic element further includes: constructing an error equation based on the map coordinates of each effective semantic feature and the map coordinates of the effective semantic elements that match each effective semantic feature, and optimizing the error equation to obtain the optimized vehicle pose at the current moment.

[0078] Specifically, the error equation, constructed based on the map coordinates of each effective semantic feature and the map coordinates of the effective semantic elements that match each effective semantic feature, can be expressed in the following form:

[0079] ;

[0080] The optimal vehicle position and optimal vehicle pose are obtained by using iterative optimization algorithms for solving nonlinear least squares problems, such as the Levenberg-Marquardt method or the Gauss-Newton method. The optimal vehicle position and optimal vehicle pose together constitute the optimal vehicle pose.

[0081] Compared to the vehicle prediction pose obtained by the above technical solutions, the calculated optimized vehicle pose improves the accuracy of the current vehicle pose calculation result, providing a location basis for loading and updating the global semantic map. Furthermore, when calculating the vehicle prediction position for the next sampling time based on the current vehicle position, the more accurate optimized vehicle pose can be used directly to replace the current vehicle prediction pose, greatly improving the accuracy of the vehicle position calculation result and avoiding the situation where the calculation error is further amplified by the increase of sampling time.

[0082] The technical solution of this invention updates the position of corresponding semantic elements in the global semantic map by detecting and obtaining effective semantic features in the global perspective image, updating the semantic type and position of corresponding semantic elements in the global semantic map by detecting and obtaining conflict semantic features in the global perspective image, and adding semantic elements to the global semantic map by detecting and obtaining new semantic features in the global perspective image. This achieves real-time updating of the global semantic map and ensures the effectiveness and integrity of semantic elements in the global semantic map.

[0083] Example 3

[0084] Figure 3 This is a flowchart of a semantic localization method provided in Embodiment 3 of the present invention. The relationship between this embodiment and the above embodiments is that the semantic type of the semantic element to be updated is determined based on the number of semantic type detections of conflicting semantic features and the number of semantic type detections of newly added semantic features. Figure 3 As shown, the method specifically includes:

[0085] S301. Load a local semantic map based on the global semantic map, and convert the pixel coordinates of each semantic element in the local semantic map into map coordinates.

[0086] S302. Construct a global perspective image based on the local perspective image obtained by the vehicle surround view camera, and obtain the pixel coordinates of each semantic feature through the global perspective image.

[0087] S303. Based on the vehicle's predicted pose at the current moment and the pixel coordinates of each semantic feature, obtain the map coordinates of each semantic feature.

[0088] S304. If, within a preset interval distance of the second semantic feature, only a second semantic element of a different semantic type than the second semantic feature is detected, then the second semantic feature is determined to be a conflicting semantic feature.

[0089] S305. Based on the semantic type of the second semantic feature, the number of semantic type detections corresponding to the second semantic element is incremented.

[0090] After determining that the second semantic feature is a conflicting semantic feature, it indicates that the second semantic element has an incorrect semantic type labeling problem. However, the result of a single semantic type detection may also have errors. In this case, based on the semantic type of the second semantic feature, the detection count of the second semantic element of the same type is incremented by 1. For example, if the semantic type of the second semantic feature is lane line, while the semantic type of the second semantic element is parking line, the detection count of the lane line of the second semantic element is incremented by 1. In the subsequent localization process, if a certain semantic feature is detected again that matches the position of the second semantic element, but the semantic type is different, the detection count of the second semantic feature of the same type is also incremented by 1 based on the semantic type of the semantic feature.

[0091] S306. The first semantic type that has been detected the most times corresponding to the second semantic element is taken as the semantic type of the second semantic element; wherein, the number of detections of the first semantic type is greater than or equal to the first detection threshold.

[0092] For the second semantic element, when the number of detections of a certain semantic type is greater than or equal to the first threshold, and the number of detections of that semantic type is the highest among all semantic types, that semantic type is the first semantic type. The first semantic type is used as the semantic type of the second semantic element. By comprehensively judging the results of multiple detections, the semantic type of the conflicting semantic element is updated, which greatly improves the accuracy of the semantic type of the semantic element and ensures the semantic accuracy of the global semantic map.

[0093] S307. If no semantic element is detected within the preset interval distance of the third semantic feature, then the third semantic feature is determined to be a newly added semantic feature.

[0094] S308. Obtain the target pixel that matches the third semantic feature, and increment the number of semantic type detections corresponding to the target pixel according to the semantic type of the third semantic feature.

[0095] After determining that the third semantic feature is a newly added semantic feature, it indicates that there may be missing semantic elements in the global semantic map. However, the result of a single semantic type detection may also have errors. At this time, according to the semantic type of the third semantic feature, the detection count of the target pixel with the same semantic type is incremented by 1. For example, if the semantic type of the third semantic feature is lane line, the detection count of lane line of the target pixel is incremented by 1. In the subsequent localization process, when a certain semantic feature is detected to match the position of the target pixel again, the detection count of the target pixel with the same type is also incremented by 1 according to the semantic type of the semantic feature.

[0096] S309. The second semantic type that has been detected the most times is taken as the semantic type of the target pixel; wherein the number of detections of the second semantic type is greater than or equal to the second threshold.

[0097] For a target pixel, if the number of detections of a certain semantic type is greater than or equal to the second threshold, and the number of detections of this semantic type is the highest among all semantic types, this semantic type is the second semantic type. The second semantic type is used as the semantic type of the target pixel. By comprehensively judging the results of multiple detections, the semantic type of the semantic element is updated, which greatly improves the accuracy of the semantic type of the target pixel and ensures the semantic integrity of the global semantic map.

[0098] The technical solution of this invention, based on the semantic type of the second semantic feature, increments the number of semantic type detections corresponding to the second semantic element, and takes the first semantic type with the most detections as the semantic type of the second semantic element; thus, through the comprehensive judgment of multiple detection results, the semantic type of the conflicting semantic element is updated, ensuring the semantic accuracy of the global semantic map; based on the semantic type of the third semantic feature, increments the number of semantic type detections corresponding to the target pixel, and takes the second semantic type with the most detections as the semantic type of the target pixel; thus, through the comprehensive judgment of multiple detection results, the semantic type of the added semantic element is updated, ensuring the semantic integrity of the global semantic map.

[0099] Example 4

[0100] Figure 4 This is a structural block diagram of a semantic positioning device provided in Embodiment 4 of the present invention. The device specifically includes:

[0101] The local semantic map acquisition module 401 is used to load a local semantic map based on the global semantic map and convert the pixel coordinates of each semantic element in the local semantic map into map coordinates.

[0102] The global perspective image acquisition module 402 is used to construct a global perspective image based on the local perspective image acquired by the vehicle surround view camera, and to obtain the pixel coordinates of each semantic feature through the global perspective image;

[0103] The map coordinate acquisition module 403 is used to acquire the map coordinates of each semantic feature based on the vehicle's predicted pose at the current time and the pixel coordinates of each semantic feature.

[0104] The map update execution module 404 is used to perform semantic updates on the global semantic map based on the map coordinates of each semantic feature and the map coordinates of each semantic element.

[0105] The technical solution of this invention loads a local semantic map based on a global semantic map, and converts the pixel coordinates of each semantic element in the local semantic map into map coordinates; constructs a global perspective image based on local view images obtained from the vehicle surround view camera, and obtains the pixel coordinates of each semantic feature through the global perspective image; obtains the map coordinates of each semantic feature based on the vehicle's predicted pose at the current moment and the pixel coordinates of each semantic feature; and updates the global semantic map based on the map coordinates of each semantic feature and the map coordinates of each semantic element. Thus, by loading the local semantic map, memory overflow in the vehicle's infotainment system is avoided; and by updating the semantic elements of the global semantic map based on the semantic features extracted from the global perspective image, the timeliness of the global semantic map update is ensured, and the vehicle's positioning accuracy is improved.

[0106] Optionally, the local semantic map acquisition module 401 is specifically used to determine the local area range based on the vehicle driving speed and memory usage, and load the local semantic map based on the global semantic map according to the local area range.

[0107] Optionally, the map update execution module 404 is specifically configured to determine the first semantic feature as a valid semantic feature if a first semantic element of the same semantic type as the first semantic feature is detected within a preset interval distance of the first semantic feature; determine the second semantic feature as a conflicting semantic feature if only a second semantic element of a different semantic type as the second semantic feature is detected within a preset interval distance of the second semantic feature; and determine the third semantic feature as a newly added semantic feature if no semantic element is detected within a preset interval distance of the third semantic feature.

[0108] Optionally, the map update execution module 404 is further configured to construct an error equation based on the map coordinates of each effective semantic feature and the map coordinates of the effective semantic elements that match each effective semantic feature, and optimize the error equation to obtain the optimized vehicle pose at the current moment.

[0109] Optionally, the map update execution module 404 is further configured to: increment the number of semantic type detections corresponding to the second semantic element according to the semantic type of the second semantic feature; take the first semantic type with the most semantic type detections corresponding to the second semantic element as the semantic type of the second semantic element; wherein the number of detections of the first semantic type is greater than or equal to the first threshold; and / or after determining that the third semantic feature is a newly added semantic feature, further comprising: obtaining a target pixel that matches the third semantic feature, and incrementing the number of semantic type detections corresponding to the target pixel according to the semantic type of the third semantic feature; taking the second semantic type with the most semantic type detections corresponding to the target pixel as the semantic type of the target pixel; wherein the number of detections of the second semantic type is greater than or equal to the second threshold.

[0110] Optionally, the semantic localization device is also used to reload the local semantic map based on the global semantic map if the detected distance between the vehicle position and any boundary of the current local semantic map is less than or equal to a preset distance threshold.

[0111] The above-described apparatus can execute the semantic localization method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the semantic localization method provided in any embodiment of the present invention.

[0112] Example 5

[0113] Figure 5 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, electronic devices, blade electronic devices, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0114] like Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0115] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0116] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as semantic localization methods.

[0117] In some embodiments, the semantic localization method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on a heterogeneous hardware accelerator via ROM and / or a communication unit. When the computer program is loaded into RAM and executed by a processor, one or more steps of the semantic localization method described above may be performed. Alternatively, in other embodiments, the processor may be configured to perform the semantic localization method by any other suitable means (e.g., by means of firmware).

[0118] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0119] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0120] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0121] To provide user interaction, the systems and techniques described herein can be implemented on a heterogeneous hardware accelerator, which includes: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the heterogeneous hardware accelerator. Other types of devices can also be used to provide user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or haptic feedback); and input from the user can be received in any form (including sound input, voice input, or haptic input).

[0122] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0123] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0124] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and no limitation is imposed herein.

[0125] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A semantic localization method, characterized in that, include: A local semantic map is loaded based on a global semantic map, and the pixel coordinates of each semantic element in the local semantic map are converted into map coordinates; A global view image is constructed based on the local view images obtained by the vehicle surround view camera, and the pixel coordinates of each semantic feature are obtained through the global view image. Based on the vehicle's predicted pose at the current moment and the pixel coordinates of each semantic feature, obtain the map coordinates of each semantic feature; The global semantic map is semantically updated based on the map coordinates of each semantic feature and the map coordinates of each semantic element.

2. The semantic localization method according to claim 1, characterized in that, The loading of a local semantic map based on a global semantic map includes: The local area is determined based on vehicle speed and memory usage, and a local semantic map is loaded based on the global semantic map according to the local area.

3. The semantic localization method according to claim 1, characterized in that, The step of semantically updating the global semantic map based on the map coordinates of each semantic feature and the map coordinates of each semantic element includes: If a first semantic element of the same semantic type as the first semantic feature is detected within a preset interval distance of the first semantic feature, then the first semantic feature is determined to be a valid semantic feature. If, within a preset interval distance of the second semantic feature, only a second semantic element of a different semantic type than the second semantic feature is detected, then the second semantic feature is determined to be a conflicting semantic feature. If no semantic element is detected within the preset interval of the third semantic feature, then the third semantic feature is determined to be a newly added semantic feature.

4. The semantic localization method according to claim 3, characterized in that, The step of semantically updating the global semantic map based on the map coordinates of each semantic feature and the map coordinates of each semantic element further includes: An error equation is constructed based on the map coordinates of each effective semantic feature and the map coordinates of the effective semantic elements that match each effective semantic feature. The error equation is then optimized to obtain the optimized vehicle pose at the current moment.

5. The semantic localization method according to claim 3, characterized in that, After determining that the second semantic feature is a conflict semantic feature, the following steps are also included: Based on the semantic type of the second semantic feature, the number of semantic type detections corresponding to the second semantic element is incremented; The semantic type that has the highest number of detections for the second semantic element is taken as the semantic type of the second semantic element; wherein, the number of detections for the first semantic type is greater than or equal to the first detection threshold. And / or after determining that the third semantic feature is a newly added semantic feature, it also includes: Obtain the target pixel that matches the third semantic feature, and increment the number of semantic type detections corresponding to the target pixel according to the semantic type of the third semantic feature; The second semantic type that has the highest number of semantic type detections for the target pixel is taken as the semantic type of the target pixel; wherein the number of detections for the second semantic type is greater than or equal to the second threshold.

6. The semantic localization method according to claim 1, characterized in that, The semantic localization method further includes: If the detected distance between the vehicle's location and any boundary of the current local semantic map is less than or equal to a preset distance threshold, the local semantic map is reloaded based on the global semantic map.

7. A semantic positioning device, characterized in that, include: The local semantic map acquisition module is used to load a local semantic map based on the global semantic map and convert the pixel coordinates of each semantic element in the local semantic map into map coordinates. The global view image acquisition module is used to construct a global view image based on the local view image acquired by the vehicle surround view camera, and to obtain the pixel coordinates of each semantic feature through the global view image; The map coordinate acquisition module is used to acquire the map coordinates of each semantic feature based on the vehicle's predicted pose at the current time and the pixel coordinates of each semantic feature. The map update execution module is used to perform semantic updates on the global semantic map based on the map coordinates of each semantic feature and the map coordinates of each semantic element.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the semantic localization method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the semantic localization method according to any one of claims 1-6.

10. A computer program product comprising a computer program that, when executed by a processor, implements the semantic localization method of any one of claims 1-6.