A method and device for generating labeled data
By constructing and fusing three-dimensional vector maps from overlapping road segments, the method automates the annotation process, addressing inefficiencies in manual image annotation for autonomous driving, thereby improving data acquisition speed and accuracy.
Patent Information
- Application Number
- CN202310012760.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-05
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-01-05
AI Technical Summary
During autonomous driving or assisted driving, the method of obtaining labeled data by manually labeling static objects in the image is less efficient and has a large workload.
By constructing a three-dimensional vector map of the first and second sections, and performing clustering and fusion to generate a three-dimensional vector map of the third section, static objects are automatically recognized based on the map to generate labeled data.
It improves the efficiency of obtaining labeled data, and can quickly and efficiently obtain known images with labeled data, with better applicability.
Smart Images

Figure CN116259035B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of autonomous driving or assisted driving, and particularly to a method and apparatus for generating annotation data. Background Art
[0002] During the process of autonomous driving or assisted driving, a vehicle needs to identify static objects (such as lane lines, traffic lights, road signs, etc.) in the surrounding geographical environment through an automatic recognition model and sensors (such as cameras), so as to plan a correct driving route.
[0003] The automatic recognition model needs to be trained through known images with annotation data. Therefore, a large number of known images with annotation data need to be prepared first. In the prior art, manual annotation is usually used to annotate static objects in images, thereby determining the annotation data of static objects and obtaining known images with annotation data. In this way, it is necessary to manually annotate each frame of image, resulting in a large workload and low efficiency. Summary of the Invention
[0004] Currently, in the process of training an automatic recognition model in the application scenarios of autonomous driving or assisted driving, the method of obtaining known images with annotation data has a large workload and low efficiency.
[0005] To solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a method and apparatus for generating annotation data.
[0006] According to one aspect of the present disclosure, a method for generating annotation data is provided. The method includes:
[0007] Based on a first image set obtained by photographing a first road section at a first moment, constructing a first three-dimensional vector map of the first road section;
[0008] Based on a second image set obtained by photographing a second road section at a second moment, constructing a second three-dimensional vector map of the second road section; the area where the second road section is located overlaps with the area where the first road section is located;
[0009] Based on the first three-dimensional vector map and the second three-dimensional vector map, clustering and fusing to generate a third three-dimensional vector map of a third road section; the area where the third road section is located covers the area where the first road section is located and the area where the second road section is located;
[0010] Based on the third three-dimensional vector map, generating annotation data of static objects in the third road section.
[0011] According to another aspect of the present disclosure, an apparatus for generating annotation data is provided. The apparatus includes:
[0012] A first image construction module, configured to construct a first 3D vector map of the first road section based on a first image set captured at a first moment of the first road section;
[0013] A second image construction module, configured to construct a second 3D vector map of the second road section based on a second image set captured at a second moment of the second road section; an area where the second road section is located overlaps with an area where the first road section is located;
[0014] A fusion module, configured to perform clustering and fusion based on the first 3D vector map obtained by the first image construction module and the second 3D vector map obtained by the second image construction module to generate a third 3D vector map of a third road section; an area where the third road section is located covers the area where the first road section is located and the area where the second road section is located;
[0015] A labeled data generation module, configured to generate labeled data of static objects in the third road section based on the third 3D vector map obtained by the fusion module.
[0016] According to another aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program for executing the above-described labeled data generation method.
[0017] According to another aspect of the present disclosure, there is provided an electronic device, which includes:
[0018] A processor;
[0019] A memory for storing executable instructions of the processor;
[0020] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the above-described labeled data generation method.
[0021] Based on the above solution provided by the present disclosure, a first 3D vector map of the first road section can be constructed respectively based on the first image set obtained by photographing the first road section at the first moment; and a second 3D vector map of the second road section can be constructed based on the second image set obtained by photographing the second road section at the second moment. Wherein, the area where the second road section is located overlaps with the area where the first road section is located. Then, based on the first 3D vector map and the second 3D vector map, a third 3D vector map of the third road section can be generated by clustering and fusing, wherein the area where the third road section is located covers the area where the first road section is located and the area where the second road section is located. Finally, based on the third 3D vector map, annotation data of static objects in the third road section can be generated. It can be seen that by adopting the above solution provided by the present disclosure, a 3D vector map of the third road section can be generated by fusing the reconstructed 3D vector maps of the first road section and the second road section, so that the information of static objects in the third road section can be clearly determined, and based on the 3D vector map of the third road section, each static object in the third road section can be automatically recognized, and then the annotation data of each static object can be obtained, greatly improving the acquisition efficiency of the annotation data. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] By describing the embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation to the present disclosure. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0023] Figure 1 It is a schematic diagram of an application scenario provided by an exemplary embodiment of the present disclosure.
[0024] Figure 2 It is a schematic flowchart of a method for generating annotation data provided by an exemplary embodiment of the present disclosure.
[0025] Figure 3 It is a schematic flowchart of a method for generating annotation data provided by another exemplary embodiment of the present disclosure.
[0026] Figure 4 It is a schematic flowchart of a method for generating annotation data provided by another exemplary embodiment of the present disclosure.
[0027] Figure 5 It is a schematic flowchart of a method for generating annotation data provided by another exemplary embodiment of the present disclosure.
[0028] Figure 6 It is a schematic flowchart of a method for generating annotation data provided by another exemplary embodiment of the present disclosure.
[0029] Figure 7It is a schematic diagram of an application scenario provided by another exemplary embodiment of the present disclosure.
[0030] Figure 8 It is a schematic diagram of an application scenario provided by another exemplary embodiment of the present disclosure.
[0031] Figure 9 It is a schematic flowchart of a method for generating labeled data provided by another exemplary embodiment of the present disclosure.
[0032] Figure 10 It is a block diagram of a device for generating labeled data provided by an exemplary embodiment of the present disclosure.
[0033] Figure 11 It is a block diagram of a device for generating labeled data provided by another exemplary embodiment of the present disclosure.
[0034] Figure 12 It is a block diagram of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners
[0035] Next, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.
[0036] It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0037] Those skilled in the art can understand that terms such as "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices, or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.
[0038] It should also be understood that in the embodiments of the present disclosure, "a plurality of" may refer to two or more, and "at least one" may refer to one, two, or more.
[0039] It should also be understood that for any component, data, or structure mentioned in the embodiments of the present disclosure, unless otherwise clearly defined or a contrary indication is given in the context, it is generally understood to be one or more.
[0040] In addition, the term "and / or" in the present disclosure is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.
[0041] It should also be understood that the description of each embodiment in the present disclosure emphasizes the differences between the various embodiments, and the similarities or similarities therebetween can be referred to each other. For the sake of brevity, they will not be elaborated one by one.
[0042] Meanwhile, it should be understood that, for the convenience of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationship.
[0043] The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present disclosure or its application or use.
[0044] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the specification.
[0045] It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, it does not need to be further discussed in subsequent figures.
[0046] The embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate together with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0047] Electronic devices such as terminal devices, computer systems, servers, etc. can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logics, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0048] Application Overview
[0049] The technical solution provided by the present disclosure can be applied to application scenarios of autonomous driving or assisted driving. During autonomous driving or assisted driving, an autonomous driving vehicle (such as a vehicle or an aircraft, etc.) usually needs to identify static objects (such as lane lines, traffic lights, road signs, etc.) in the surrounding geographical environment through an automatic recognition model (such as an image sensing and recognition model) and sensors (such as a camera), so as to plan a correct driving route.
[0050] Before identifying static objects using an automatic recognition model, it is necessary to train the automatic recognition model with known images with annotation data. Based on this, a large number of known images with annotation data need to be prepared in advance.
[0051] In the prior art, usually, the static objects in an image are annotated manually, so as to obtain known images with annotation data. By this means, it is necessary to manually annotate each frame of the image, with a large workload and extremely low efficiency.
[0052] To solve the above technical problems and improve the efficiency of obtaining known images with annotation data, the present disclosure provides a method and device for generating annotation data. Through the solution provided by the present disclosure, a three-dimensional vector map of a third section can be generated by fusing the three-dimensional vector maps of a reconstructed first section and a second section, so that the information of static objects in the third section can be clearly determined. Based on the three-dimensional vector map of the third section, each static object in the third section can be automatically identified, and then the annotation data of each static object can be obtained, greatly improving the efficiency of obtaining annotation data and being able to obtain known images with annotation data efficiently and quickly, with better applicability.
[0053] Exemplary System
[0054] To facilitate the understanding of the technical solution of the present disclosure, the application scenario of the technical solution provided by the present disclosure will be exemplarily described below with reference to the accompanying drawings.
[0055] Figure 1 is a schematic diagram of an application scenario provided by an exemplary embodiment of the present disclosure. As Figure 1 shown, the application scenario of the present disclosure may include: a terminal device 100 for generating annotation data (hereinafter referred to as the terminal device 100), and a vehicle 200 (hereinafter referred to as the vehicle 200) provided with a multi-camera system.
[0056] It should be noted that in some other exemplary embodiments, the terminal device 100 can also be replaced by a server for generating annotation data (hereinafter referred to as the server). Alternatively, the terminal device 100 can also be replaced by a controller, a data center, a cloud platform, etc. for generating annotation data. The present disclosure does not limit this. In the following exemplary embodiments, the terminal device 100 is taken as an example to exemplarily illustrate the technical solution provided by the present disclosure.
[0057] Among them, the multi-camera system can include multiple cameras, which are respectively arranged at different positions of the vehicle 200 and can capture the geographical environment around the vehicle 200 from different angles. Optionally, these multiple cameras can simultaneously capture the geographical environment around the vehicle 200 from different angles. Optionally, these multiple cameras can also respectively capture the geographical environment around the vehicle 200 at different times. Specifically, it can be set according to the requirements of the actual application scenario. For example, in the subsequent exemplary embodiments of the present disclosure, the multiple cameras of the multi-camera system can simultaneously capture the geographical environment around the vehicle 200 from different angles.
[0058] Exemplarily, the multi-camera system can include 6 cameras, which are respectively arranged at the front end, the rear end, the left front end, the left rear end, the right front end and the right rear end of the vehicle 200. Through these 6 cameras, the front view image, the rear view image, the left rear view image, the left front view image, the right rear view image and the right front view image of the vehicle 200 can be simultaneously collected, and the collection ranges of the 6 cameras can cover all areas around the vehicle 200.
[0059] The vehicle 200 can correspondingly store the image captured by any one of the cameras in the multi-camera system together with the shooting time of the image.
[0060] The terminal device 100 can communicate with the vehicle 200 through a wireless connection or a wired connection, and obtain the stored image and the shooting time corresponding to the image from the vehicle 200.
[0061] In addition, the exemplary functions of the terminal device 100, the multi-camera system and the vehicle 200 can also refer to the content of the subsequent embodiments, which will not be elaborated here.
[0062] Exemplary Method
[0063] Figure 2 It is a schematic flowchart of a method for generating annotation data provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to a terminal device for generating annotation data (hereinafter referred to as the terminal device), such as Figure 1 the terminal device 100 shown in Figure 2 shown, and includes the following steps:
[0064] Step S201: Based on the first image set obtained by photographing the first road section at the first moment, construct the first 3D vector map of the first road section.
[0065] As can be seen from the foregoing, a vehicle equipped with a multi-camera system, such as Figure 1 the vehicle 200 equipped with a multi-camera system shown, can photograph the geographical environment around the vehicle through its multi-camera system. Among them, the multi-camera system of the vehicle may include multiple cameras, and these multiple cameras may be arranged at different positions of the vehicle. Through the multiple cameras of the multi-camera system, the geographical environment around the vehicle can be photographed from different angles simultaneously.
[0066] Exemplarily, the viewing angles of these multiple cameras can be set to cover 360 degrees of the geographical environment around the vehicle. For each camera in the multi-camera system, its viewing angle can be set to 45 degrees or 60 degrees, etc. in the forward direction of the vehicle, or it can also be set to 45 degrees or 60 degrees, etc. in the direction opposite to the rear end of the vehicle. There may be partial overlap between the viewing angles of the multiple cameras of the multi-camera system so that information in the 360-degree geographical environment around the vehicle can be collected.
[0067] In a possible application scenario, the vehicle can run in the actual geographical environment at a certain operating speed. During the running of the vehicle, each camera of the multi-camera system on the vehicle can periodically photograph the geographical environment around the vehicle at the same frequency. That is, each camera of the multi-camera system on the vehicle can, at the same frequency, photograph the geographical environment around the vehicle from different angles simultaneously every preset time interval (the preset time interval can be set according to the requirements of the application scenario). In addition, each time a photograph is taken, the vehicle can store the images taken by each camera in the multi-camera system and the corresponding photographing times of each image.
[0068] The terminal device can be communicatively connected to the vehicle through a wired connection method or a wireless connection method. Then, the terminal device can obtain the images taken by each camera in the multi-camera system from the vehicle and the corresponding photographing times of each image.
[0069] In the present disclosure, when the vehicle runs in the geographical environment where the static object for which annotation data needs to be obtained is located, the photographing time of any one photograph can be recorded as the first moment. When the multiple cameras of the multi-camera system on the vehicle simultaneously photograph the geographical environment around the vehicle from different angles at the first moment, the geographical environment area around the vehicle covered by the viewing angles of these multiple cameras is recorded as the first road section, and the images obtained by each camera in the multi-camera system photographing the first road section are all recorded as the first images. Then, the first image set can include multiple frames of first images obtained by the multiple cameras of the multi-camera system on the vehicle simultaneously photographing the first road section from different angles.
[0070] It can be seen that the first image set contains multiple frames of first images, and these multiple frames of first images respectively reflect the information of static objects in the first section from different angles. After the terminal device obtains the first image set, it can reconstruct the three-dimensional vector map of the first section based on these multiple frames of first images. In the present disclosure, the three-dimensional vector map of the first section is denoted as the first three-dimensional vector map. Obviously, the information of static objects in the first section can be determined through the first three-dimensional vector map.
[0071] Exemplarily, the static objects in the first section may include objects of interest in the field of autonomous driving or assisted driving, and these objects of interest have a certain impact on the driving strategy planning in autonomous driving or assisted driving. For example, the static objects in the first section may include static objects such as lane lines, signs, traffic lights, etc.
[0072] Exemplarily, the first three-dimensional vector map of the first section may include information for indicating the positions of static objects in the first section, information for indicating the category attributes of static objects in the first section, etc.
[0073] It should be noted that the three-dimensional vector map can also be referred to as a three-dimensional semantic point cloud map. The first three-dimensional vector map of the first section can also be referred to as the first three-dimensional semantic point cloud map of the first section.
[0074] Step S202, construct a second three-dimensional vector map of the second section based on a second image set obtained by photographing the second section at a second moment.
[0075] Wherein, the second moment may be the shooting time of a certain shooting by the multi-camera system on the vehicle before the first moment, or the second moment may also be the shooting time of a certain shooting by the multi-camera system on the vehicle after the first moment. In the present disclosure, when multiple cameras of the multi-camera system on the vehicle simultaneously photograph the geographical environment around the vehicle from different angles at the second moment, the area of the geographical environment around the vehicle covered by the perspectives of these multiple cameras is denoted as the second section, and the images obtained by each camera in the multi-camera system photographing the second section are all denoted as second images. Then, the second image set may include multiple frames of second images obtained by multiple cameras of the multi-camera system on the vehicle simultaneously photographing the second section from different angles.
[0076] Similarly, the second image set contains multiple frames of second images, and these multiple frames of second images respectively reflect the information of static objects in the second section from different angles. After the terminal device obtains the second image set from the vehicle, it can reconstruct the three-dimensional vector map of the second section based on these multiple frames of second images. In the present disclosure, the three-dimensional vector map of the second section is denoted as the second three-dimensional vector map. Obviously, the information of static objects in the second section can be determined through the second three-dimensional vector map.
[0077] The area where the first road segment is located and the area where the second road segment is located may completely overlap, or the area where the first road segment is located and the area where the second road segment is located may partially overlap. The present disclosure does not limit this.
[0078] Exemplarily, the static objects in the second road segment may also include objects of interest in the field of autonomous driving or assisted driving, and such objects of interest have a certain impact on the driving strategy planning in autonomous driving or assisted driving. For example, the static objects in the second road segment may also include static objects such as lane lines, signs, traffic lights, etc.
[0079] Exemplarily, the second 3D vector map of the second road segment may include information for indicating the positions of the static objects in the second road segment, and information for indicating the category attributes of the static objects in the second road segment, etc.
[0080] It should be noted that the second 3D vector map of the second road segment may also be referred to as the second 3D semantic point cloud map of the second road segment.
[0081] Step S203: Based on the first 3D vector map and the second 3D vector map, cluster and fuse to generate a third 3D vector map of the third road segment.
[0082] After the terminal device constructs the first 3D vector map of the first road segment and the second 3D vector map of the second road segment, the first 3D vector map and the second 3D vector map can be cluster-fused to obtain a new 3D vector map after cluster fusion. In the present disclosure, the new 3D vector map obtained after cluster fusion is denoted as the third 3D vector map.
[0083] Corresponding to the cluster fusion of the first 3D vector map and the second 3D vector map, the geographical environment area corresponding to the first 3D vector map and the geographical environment area corresponding to the second 3D vector map are correspondingly merged into a new geographical environment area. In the present disclosure, the new geographical environment area obtained by merging is denoted as the third road segment. That is, the first road segment and the second road segment are merged into the third road segment. Obviously, the area where the third road segment is located can cover the area where the first road segment is located and the area where the second road segment is located.
[0084] Step S204: Based on the third 3D vector map, generate annotation data of the static objects in the third road segment.
[0085] In the method for generating labeled data provided by the present disclosure, a three-dimensional vector map of the third section can be generated by fusing the three-dimensional vector map of the reconstructed first section and the three-dimensional vector map of the second section. Thus, the information of static objects in the third section can be clearly determined. Based on the three-dimensional vector map of the third section, each static object in the third section can be automatically identified, and then the labeled data of each static object can be obtained, greatly improving the acquisition efficiency of the labeled data, and enabling the efficient and rapid acquisition of known images with labeled data, with better applicability.
[0086] In other exemplary embodiments of the present disclosure, in addition to the first three-dimensional vector map and the second three-dimensional vector map, more three-dimensional vector maps can be clustered and fused to obtain a new three-dimensional vector map after clustering and fusion. Then, based on the new three-dimensional vector map, the labeled data of static objects can be generated. Based on this, in another exemplary embodiment of the present disclosure, reference can also be made to Figure 3 . Such as Figure 3 shown, the method for generating labeled data may include the following steps:
[0087] Step S301, based on the first image set obtained by photographing the first section at the first moment, construct the first three-dimensional vector map of the first section.
[0088] The specific content of step S301 can refer to the content of the foregoing embodiment and will not be elaborated here.
[0089] Step S302, based on the second image set obtained by photographing the second section at the second moment, construct the second three-dimensional vector map of the second section.
[0090] The specific content of step S302 can refer to the content of the foregoing embodiment and will not be elaborated here.
[0091] Step S303, based on the third image set obtained by photographing the fourth section at the third moment, construct the fourth three-dimensional vector map of the fourth section.
[0092] Wherein, the third moment can be the shooting time of a certain shooting of the multi-camera system on the vehicle before the first moment and the second moment. Or, the third moment can also be the shooting time of a certain shooting of the multi-camera system on the vehicle after the first moment and the second moment. Or, the third moment can also be the shooting time of a certain shooting of the multi-camera system on the vehicle between the first moment and the second moment.
[0093] In the present disclosure, when multiple cameras of the multi-camera system on the vehicle simultaneously capture the geographical environment around the vehicle at the third moment, the perspectives of these multiple cameras cover the geographical environment area around the vehicle, which is denoted as the fourth road segment. The images obtained by each camera in the multi-camera system when separately capturing the fourth road segment are all denoted as the third images. Then, the third image set may include multiple frames of third images obtained by multiple cameras of the multi-camera system on the vehicle simultaneously capturing the fourth road segment from different perspectives.
[0094] Similarly, the third image set contains multiple frames of third images, which respectively reflect the information of static objects in the fourth road segment from different perspectives of the fourth road segment. After the terminal device obtains the third image set from the vehicle, it can reconstruct the three-dimensional vector map of the fourth road segment based on these multiple frames of third images. In the present disclosure, the three-dimensional vector map of the fourth road segment is denoted as the fourth three-dimensional vector map. Obviously, the information of static objects in the fourth road segment can be determined through the fourth three-dimensional vector map.
[0095] The area where the first road segment is located and the area where the second road segment is located can both completely overlap with the area where the fourth road segment is located, or the area where the first road segment is located and the area where the second road segment is located can both partially overlap with the area where the fourth road segment is located. The present disclosure does not limit this.
[0096] Exemplarily, the static objects in the fourth road segment may include objects of interest in the field of autonomous driving or assisted driving, and these objects of interest have a certain impact on the driving strategy planning in autonomous driving or assisted driving. For example, the static objects in the fourth road segment may include lane lines, road signs, traffic lights and other static objects.
[0097] Exemplarily, the fourth three-dimensional vector map of the fourth road segment may include information for indicating the positions of static objects in the fourth road segment, and information for indicating the category attributes of static objects in the fourth road segment, etc.
[0098] It should be noted that the fourth three-dimensional vector map of the fourth road segment may also be referred to as the fourth three-dimensional semantic point cloud map of the fourth road segment.
[0099] Step S304, based on the first three-dimensional vector map, the second three-dimensional vector map, and the fourth three-dimensional vector map, cluster and fuse to generate the third three-dimensional vector map of the third road segment.
[0100] After the terminal device constructs the first three-dimensional vector map of the first road segment, the second three-dimensional vector map of the second road segment, and the fourth three-dimensional vector map of the fourth road segment, it can perform clustering and fusion on the first three-dimensional vector map, the second three-dimensional vector map, and the fourth three-dimensional vector map to obtain a new three-dimensional vector map after clustering and fusion. In this embodiment, the new three-dimensional vector map obtained after clustering and fusion is denoted as the third three-dimensional vector map.
[0101] Corresponding to the clustering fusion of the first three-dimensional vector map, the second three-dimensional vector map, and the fourth three-dimensional vector map, the geographical environment regions corresponding to the first three-dimensional vector map, the second three-dimensional vector map, and the fourth three-dimensional vector map are correspondingly merged into a new geographical environment region. In this embodiment, the newly merged geographical environment region is denoted as the third road section. That is, the first road section, the second road section, and the fourth road section are merged into the third road section. Obviously, the area where the third road section is located can cover the areas where the first road section, the second road section, and the fourth road section are located.
[0102] Step S305: Generate annotation data of static objects in the third road section based on the third three-dimensional vector map.
[0103] It should be noted that more three-dimensional vector maps can also be clustered and fused to obtain a new three-dimensional vector map after clustering fusion. That is, in addition to the first three-dimensional vector map, the second three-dimensional vector map, and the fourth three-dimensional vector map, more three-dimensional vector maps can be clustered and fused to obtain the third three-dimensional vector map of the third road section. Then, based on the third three-dimensional vector map of the third road section, generate annotation data of static objects in the third road section. The specific implementation method is similar and can refer to the content of the foregoing embodiment, which will not be elaborated here.
[0104] Moreover, in order to obtain more accurate annotation data, the foregoing multiple three-dimensional vector maps can be three-dimensional vector maps of the same road section. That is, image sets of the same road section taken at different times are respectively used to construct a three-dimensional vector map of this road section, and multiple three-dimensional vector maps of this road section are obtained. Then, based on the clustering fusion of the multiple three-dimensional vector maps, a new three-dimensional vector map of this road section is generated. After that, based on the newly obtained three-dimensional vector map after clustering fusion, obtain the annotation data of static objects in this road section, and the obtained annotation data is more accurate.
[0105] In the method for generating annotation data provided by the present disclosure, based on the reconstructed three-dimensional vector maps of the first road section, the second road section, and the fourth road section, or more three-dimensional vector maps, a three-dimensional vector map of the third road section can be fused and generated, so that the information of static objects in the third road section can be clearly determined. Based on the three-dimensional vector map of the third road section, each static object in the third road section can be automatically recognized, and then the annotation data of each static object can be obtained, greatly improving the acquisition efficiency of the annotation data, and being able to efficiently and quickly obtain known images with annotation data, and having better applicability.
[0106] In the method for generating annotation data provided by another exemplary embodiment of the present disclosure, as Figure 4 shown, in the above Figure 2Based on the embodiments shown, step S201 may include the following steps:
[0107] Step S2011: Perform semantic segmentation processing on each frame of the first image included in the first image set, or perform semantic segmentation processing and object detection processing, to obtain first preprocessed images respectively corresponding to the first images.
[0108] Exemplarily, after the terminal device obtains the first image set captured at the first moment, it may perform semantic segmentation processing on each frame of the first image included in the first image set to obtain first preprocessed images respectively corresponding to the first images. In this application scenario, the first preprocessed image of the first image may include semantic information (or, also referred to as semantic features) of static objects in the first road segment. Exemplarily, the semantic information of static objects in the first road segment may be used to indicate the category attributes of static objects in the first road segment.
[0109] Exemplarily, after the terminal device obtains the first image set captured at the first moment, it may also perform semantic segmentation processing and object detection processing on each frame of the first image included in the first image set to obtain first preprocessed images respectively corresponding to the first images. In this application scenario, the first preprocessed image of the first image may include semantic information of static objects in the first road segment, as well as initial annotation information (or, also referred to as initial annotation features) of static objects in the first road segment.
[0110] The semantic information included in the first preprocessed image can be used to indicate the category attributes of static objects in the first preprocessed image. For example, it can be used to indicate that the static objects in the first preprocessed image are lane lines, traffic lights, or signs, etc. The initial annotation information included in the first preprocessed image can be used to annotate the corresponding static objects in the first preprocessed image.
[0111] Step S2012: Obtain first camera poses respectively corresponding to the first images.
[0112] After the terminal device obtains the first image set, it is necessary to respectively obtain the poses of the cameras used to capture each frame of the first image in the first image set. That is, the terminal device needs to obtain the poses of the cameras respectively corresponding to each frame of the first image. In the present disclosure, the poses of the cameras respectively used to capture the first images corresponding to each first image are all denoted as first camera poses.
[0113] Step S2013: Based on the first preprocessed images and the first camera poses respectively corresponding to the first preprocessed images, construct a first three-dimensional vector map.
[0114] Among them, the first camera pose corresponding to the first preprocessed image refers to the first camera pose of the first image that generates the first preprocessed image.
[0115] After the terminal device obtains each first pre - processed image and the corresponding first camera pose of each first pre - processed image, it can map the information of the static objects included in the corresponding first pre - processed image to the same three - dimensional vector map based on each first camera pose, so as to reconstruct the three - dimensional vector map of the first road section, that is, reconstruct the first three - dimensional vector map.
[0116] Based on the first three - dimensional vector map, the information of the static objects in the first road section can be determined, such as position information, semantic information, etc.
[0117] In the method for generating annotation data provided by the present disclosure, the three - dimensional vector map of the first road section can be reconstructed based on multiple frames of first images taken simultaneously from different angles of the first road section and the corresponding camera poses of each first image. Based on the three - dimensional vector map of the first road section, the information of the static objects in the first road section can be determined. It is convenient to automatically identify each static object in the first road section subsequently, and then obtain the annotation data of each static object, which can improve the acquisition efficiency of the annotation data.
[0118] In the method for generating annotation data provided by another exemplary embodiment of the present disclosure, as Figure 5 shown, based on the above - mentioned Figure 2 shown embodiment, step S202 may include the following steps:
[0119] Step S2021, perform semantic segmentation processing on each frame of the second image included in the second image set, or perform semantic segmentation processing and object detection processing, to obtain the second pre - processed image corresponding to each of the second images.
[0120] Exemplarily, after the terminal device obtains the second image set taken at the second moment, it can perform semantic segmentation processing on each frame of the second image included in the second image set to obtain the second pre - processed image corresponding to each of the second images. In this application scenario, the second pre - processed image of the second image may include the semantic information of the static objects in the second road section. Exemplarily, the semantic information of the static objects in the second road section can be used to indicate the category attributes of the static objects in the second road section.
[0121] Exemplarily, after the terminal device obtains the second image set taken at the second moment, it can also perform semantic segmentation processing and object detection processing on each frame of the second image included in the second image set to obtain the second pre - processed image corresponding to each of the second images. In this application scenario, the second pre - processed image of the second image may include the semantic information of the static objects in the second road section and the initial annotation information of the static objects in the second road section.
[0122] The semantic information included in the second preprocessed image can be used to indicate the category attributes of static objects in the second preprocessed image. For example, it can be used to indicate that the static objects in the second preprocessed image are lane lines, traffic lights, or signs, etc. The initial annotation information included in the second preprocessed image can be used to annotate the corresponding static objects in the second preprocessed image.
[0123] Step S2022: Obtain the second camera poses corresponding to each of the second images respectively.
[0124] After the terminal device obtains the second image set, it is necessary to obtain the poses of the cameras used to capture each frame of the second image in the second image set respectively. That is, the terminal device needs to obtain the poses of the cameras corresponding to each frame of the second image respectively. In this disclosure, the poses of the cameras respectively used to capture the corresponding second images are all denoted as the second camera poses.
[0125] Step S2023: Based on each of the second preprocessed images and the second camera poses corresponding to each of the second preprocessed images respectively, construct a second 3D vector map.
[0126] Among them, the second camera pose corresponding to the second preprocessed image refers to the second camera pose corresponding to the second image that generates the second preprocessed image.
[0127] After the terminal device obtains each second preprocessed image and the second camera poses corresponding to each of the second preprocessed images respectively, it can, based on each second camera pose, map the information of the static objects included in the corresponding second preprocessed image to the same 3D vector map, so as to reconstruct the 3D vector map of the second road section, that is, reconstruct the second 3D vector map.
[0128] Based on the second 3D vector map, the information of the static objects in the second road section can be determined, such as position information, semantic information, etc.
[0129] In the method for generating annotation data provided by this disclosure, a 3D vector map of the second road section can be reconstructed based on multiple frames of second images captured simultaneously from different angles of the second road section and the camera poses corresponding to each of the second images respectively. Based on the 3D vector map of the second road section, the information of the static objects in the second road section can be determined. It is convenient to subsequently automatically identify each static object in the second road section, and then obtain the annotation data of each static object, which can improve the acquisition efficiency of the annotation data.
[0130] In the method for generating annotation data provided by another exemplary embodiment of this disclosure, as Figure 6 shown, on the basis of the above Figure 2 shown embodiment, step S203 may include the following steps:
[0131] Step S2031: Align the first 3D vector map and the second 3D vector map.
[0132] After the terminal device reconstructs the first 3D vector map of the first road section and the second 3D vector map of the second road section, the first 3D vector map and the second 3D vector map can be aligned.
[0133] Exemplarily, when aligning the first 3D vector map and the second 3D vector map, the pose of the vehicle can be changed so that the same static objects in the first 3D vector map and the second 3D vector map can be aggregated together.
[0134] For example, as Figure 7 shown, Figure 7 in (a) shows the first 3D vector map and the second 3D vector map before alignment. Figure 7 in (b) shows the first 3D vector map and the second 3D vector map after alignment.
[0135] As Figure 7 shown in (a), before alignment, the lane lines in the first 3D vector map and the lane lines in the second 3D vector map are in a separated state from each other. The annotation boxes for indicating static objects in the first 3D vector map and the annotation boxes for indicating the same static object in the second 3D vector map are also in a separated state from each other.
[0136] As Figure 7 shown in (b), after alignment, the lane lines in the first 3D vector map and the lane lines in the second 3D vector map overlap with each other. The annotation boxes for indicating static objects in the first 3D vector map and the annotation boxes for indicating the same static object in the second 3D vector map are aggregated together.
[0137] It should be noted that in other exemplary embodiments of the present disclosure, a third 3D vector map of the third road section can also be obtained by clustering and fusing multiple 3D vector maps. In such an application scenario, for aligning these multiple 3D vector maps, reference can be made to Figure 8 . Figure 8 in (a) shows these multiple 3D vector maps and the reprojection images of these multiple 3D vector maps before alignment. Figure 8 in (b) shows these multiple 3D vector maps and the reprojection images of these multiple 3D vector maps after alignment.
[0138] As Figure 8 shown in (a), before alignment, the lane lines in these multiple 3D vector maps are separated from each other, and the annotation boxes of the same static objects are also separated from each other. As Figure 8As shown in (b) of , after alignment, the lane lines in these multiple 3D vector maps overlap with each other, and the bounding boxes of the same static objects also aggregate together.
[0139] Step S2032: Based on the aligned first 3D vector map and the second 3D vector map, perform clustering fusion to generate a third 3D vector map.
[0140] After the terminal device aligns the first 3D vector map and the second 3D vector map, it can perform clustering fusion on the aligned first 3D vector map and the second 3D vector map to generate a third 3D vector map.
[0141] In addition, after the terminal device aligns the first 3D vector map and the second 3D vector map, it is necessary to optimize the first camera poses of each camera based on the aligned first 3D vector map to obtain the first optimized camera poses corresponding to the respective first camera poses. In specific implementation, according to the pose change of the vehicle before and after the alignment of the first 3D vector map, the first camera poses of each camera on the vehicle can be optimized to obtain the first optimized camera poses of each first camera pose.
[0142] Similarly, after the terminal device aligns the first 3D vector map and the second 3D vector map, it is necessary to optimize the second camera poses of each camera based on the aligned second 3D vector map to obtain the second optimized camera poses corresponding to the respective second camera poses. In specific implementation, according to the pose change of the vehicle before and after the alignment of the second 3D vector map, the second camera poses of each camera on the vehicle can be optimized to obtain the second optimized camera poses of each second camera pose.
[0143] After that, the terminal device can generate the annotation data of the static objects in the third section based on the third 3D vector map, each first optimized camera pose, and each second optimized camera pose.
[0144] In the method for generating annotation data provided by the present disclosure, the first 3D vector map and the second 3D vector map can be aligned, or more 3D vector maps can be aligned. Then, the aligned first 3D vector map and the second 3D vector map, or more 3D vector maps, can be clustered and fused into a third 3D vector map. Moreover, based on the aligned first 3D vector map and the second 3D vector map, or more 3D vector maps, the camera poses can be optimized respectively to obtain the optimized camera poses. After that, based on the third 3D vector map and the optimized camera poses, the annotation data of the static objects in the third section can be generated. Since the third 3D vector map clusters and fuses at least two 3D vector maps, the information of the static objects is more complete, and the obtained annotation data is more accurate.
[0145] In the method for generating annotation data provided in another exemplary embodiment of the present disclosure, asFigure 9 As shown above Figure 6 Based on the third 3D vector map, each first optimized camera pose, and each second optimized camera pose, on the basis of the above-described embodiments, generating annotation data of static objects in the third road section may include the following steps:
[0146] Step S2041, by way of back-projection, generating a two-dimensional image of the third road section based on the third 3D vector map, each first optimized camera pose, and each second optimized camera pose.
[0147] After the terminal device obtains the third 3D vector map, each first optimized camera pose, and each second optimized camera pose, it can perform back-projection on the third 3D vector map respectively based on each first optimized camera pose and each second optimized camera pose to obtain two-dimensional images corresponding to each optimized camera pose. These two-dimensional images are two-dimensional images of different angles of the third road section, and can reflect the information of static objects in the third road section from different angles of the third road section.
[0148] Step S2042, generating a bird's-eye view of the third road section based on the third 3D vector map, each first optimized camera pose, and each second optimized camera pose.
[0149] After the terminal device obtains the third 3D vector map, each first optimized camera pose, and each second optimized camera pose, it can also construct a bird's-eye view of the third road section based on the third 3D vector map, each first optimized camera pose, and each second optimized camera pose. The bird's-eye view of the third road section can reflect the information of static objects in the third road section from a bird's-eye perspective.
[0150] Step S2043, generating two-dimensional annotation data of static objects in the third road section based on the two-dimensional image.
[0151] After the terminal device obtains two-dimensional images of different angles of the third road section by way of back-projection, it can annotate the static objects in the third road section in each frame of two-dimensional image to obtain two-dimensional annotation data of the corresponding static objects. Exemplarily, the two-dimensional annotation data of static objects may include information for indicating the position of the static objects and information for indicating the category attributes of the static objects.
[0152] Exemplarily, when the static objects are non-extensible static objects such as traffic lights and signs, the static objects in each frame of two-dimensional image can be annotated by a two-dimensional annotation box. In this scenario, the two-dimensional annotation data of static objects may include two-dimensional annotation box data of static objects, and the position of the corresponding static objects can be indicated by this two-dimensional annotation box data.
[0153] Exemplarily, when the static object is an extensible static object such as a lane line, in each frame of the two-dimensional image, the corresponding static object can be labeled through key points or curves. In such a scenario, the two-dimensional labeling data of the static object can include key point data or curve data, and through the key point data or curve data, the position of the corresponding static object can be indicated.
[0154] Step S2044, generating bird's-eye view labeling data of the static object in the third section based on the bird's-eye view.
[0155] After the terminal device constructs the bird's-eye view of the third section, based on the bird's-eye view, the static object in the third section can be labeled to obtain the bird's-eye view labeling data of the corresponding static object. Exemplarily, the bird's-eye view labeling data can be BEV labeling data from a bird's-eye view perspective.
[0156] In some other exemplary embodiments of the present disclosure, there may also be some dynamic objects in the geographical environment where the vehicle operates, such as moving vehicles or pedestrians, etc. In such an application scenario, when the terminal device constructs the first three-dimensional vector map of the first section based on the first image set captured at the first moment, it can also construct a three-dimensional annotation box of the dynamic object at the first moment (hereinafter simply referred to as the first three-dimensional annotation box) based on each first image included in the first image set and the first camera pose corresponding to each first image. That is, the terminal device can label the dynamic object in the first three-dimensional vector map through the first three-dimensional annotation box to obtain the first three-dimensional annotation box data of the dynamic object at the first moment.
[0157] Similarly, when the terminal device constructs the second three-dimensional vector map of the second section based on the second image set captured at the second moment, it can also construct a three-dimensional annotation box of the dynamic object at the second moment (hereinafter simply referred to as the second three-dimensional annotation box) based on each second image included in the second image set and the second camera pose corresponding to each second image. That is, the terminal device can label the dynamic object in the second three-dimensional vector map through the second three-dimensional annotation box to obtain the second three-dimensional annotation box data of the dynamic object at the second moment.
[0158] After that, after the terminal device clusters and fuses to obtain the third three-dimensional vector map, the terminal device can also generate the annotation data of the dynamic object at the first moment based on the third three-dimensional vector map, each first optimized camera pose, and the first three-dimensional annotation box data.
[0159] Similarly, after the terminal device clusters and fuses to obtain the third three-dimensional vector map, it can also generate the annotation data of the dynamic object at the second moment based on the third three-dimensional vector map, each second optimized camera pose, and the second three-dimensional annotation box data.
[0160] In addition, when a third 3D vector map is fused with multiple 3D vector maps, and these multiple 3D vector maps are respectively reconstructed from image sets at different times, the annotation data of the dynamic object at each time can be obtained in the above manner.
[0161] In the method for generating annotation data provided by the present disclosure, the 2D annotation data of the static object can be obtained by back-projection based on the 3D vector map obtained by clustering and fusing at least two 3D vector maps. A bird's-eye view can also be constructed based on the 3D vector map obtained by clustering and fusing. After that, the bird's-eye view annotation data of the static object can be obtained based on the bird's-eye view. In addition, based on the 3D annotation box data of the dynamic object at different times and the 3D vector map obtained by clustering and fusing, the annotation data of the dynamic object at different times can be obtained. It can be seen that more accurate annotation data with better applicability can be obtained through the method of the present disclosure.
[0162] Exemplary Apparatus
[0163] Figure 10 is a structural block diagram of a device for generating annotation data provided by an exemplary embodiment of the present disclosure. This device can be applied to a terminal device, such as Figure 1 the terminal device 100 shown. Alternatively, this device can also be the terminal device itself. By using this device, the method for generating annotation data provided by any of the above embodiments of the present disclosure can be executed, and corresponding beneficial effects can be obtained.
[0164] As Figure 10 shown, the device for generating annotation data provided by the present disclosure may include: a first image construction module 1001, a second image construction module 1002, a fusion module 1003, and an annotation data generation module 1004. Among them,
[0165] The first image construction module 1001 is configured to construct a first 3D vector map of the first road section based on a first image set captured of the first road section at a first time.
[0166] The second image construction module 1002 is configured to construct a second 3D vector map of the second road section based on a second image set captured of the second road section at a second time; the area where the second road section is located overlaps with the area where the first road section is located.
[0167] The fusion module 1003 is configured to cluster and fuse the first 3D vector map obtained by the first image construction module 1001 and the second 3D vector map obtained by the second image construction module 1002 to generate a third 3D vector map of a third road section; the area where the third road section is located covers the area where the first road section is located and the area where the second road section is located.
[0168] The labeled data generation module 1004 is configured to generate labeled data of static objects in the third road segment based on the third 3D vector map obtained by the fusion module 1003.
[0169] In another exemplary embodiment of the present disclosure, the apparatus for generating the labeled data may further include a third image construction module, configured to construct a fourth 3D vector map of the fourth road segment based on a third image set captured at a third moment; the regions where the first road segment is located and the second road segment is located both overlap with the region where the fourth road segment is located. The fusion module 1003 may further be configured to: based on the first 3D vector map obtained by the first image construction module 1001, the second 3D vector map obtained by the second image construction module 1002, and the fourth 3D vector map obtained by the third image construction module, perform clustering fusion to generate the third 3D vector map of the third road segment, wherein the region where the third road segment is located covers the region where the fourth road segment is located.
[0170] In another exemplary embodiment of the present disclosure, as Figure 11 shown, the first image construction module 1001 may include:
[0171] A first preprocessing unit 10011, configured to perform semantic segmentation processing, or perform semantic segmentation processing and object detection processing, on each frame of the first image included in the first image set, to obtain a first preprocessed image corresponding to each of the first images.
[0172] A first acquisition unit 10012, configured to acquire a first camera pose corresponding to each of the first images.
[0173] A first construction unit 10013, configured to construct the first 3D vector map based on the first preprocessed images obtained by the first preprocessing unit 10011 and the first camera poses corresponding to the first preprocessed images obtained by the first acquisition unit 10012.
[0174] In another exemplary embodiment of the present disclosure, as Figure 11 shown, the second image construction module 1002 may include:
[0175] A second preprocessing unit 10021, configured to perform semantic segmentation processing, or perform semantic segmentation processing and object detection processing, on each frame of the second image included in the second image set, to obtain a second preprocessed image corresponding to each of the second images.
[0176] A second acquisition unit 10022, configured to acquire a second camera pose corresponding to each of the second images.
[0177] The second construction unit 10023 is configured to construct the second 3D vector map based on each of the second preprocessed images obtained by the second preprocessing unit 10021 and the second camera poses respectively corresponding to each of the second preprocessed images obtained by the second acquisition unit 10022.
[0178] In another exemplary embodiment of the present disclosure, as Figure 11 shown, the fusion module 1003 may include:
[0179] An alignment unit 10031, configured to align the first 3D vector map obtained by the first image construction module 1001 and the second 3D vector map obtained by the second image construction module 1002.
[0180] A fusion unit 10032, configured to cluster and fuse the aligned first 3D vector map and second 3D vector map obtained by the alignment unit 10031 to generate the third 3D vector map.
[0181] In another exemplary embodiment of the present disclosure, the apparatus for generating the annotation data may further include:
[0182] A first optimization module, configured to optimize each of the first camera poses based on the aligned first 3D vector map to obtain first optimized camera poses respectively corresponding to each of the first camera poses.
[0183] A second optimization module, configured to optimize each of the second camera poses based on the aligned second 3D vector map to obtain second optimized camera poses respectively corresponding to each of the second camera poses.
[0184] The annotation data generation module 1004 may further be configured to generate annotation data of static objects in the third road segment based on the third 3D vector map obtained by the fusion unit 10032, each of the first optimized camera poses obtained by the first optimization module, and each of the second optimized camera poses obtained by the second optimization module.
[0185] In another exemplary embodiment of the present disclosure, as Figure 11 shown, the annotation data generation module 1004 may include:
[0186] A first generation unit 10041, configured to generate a two-dimensional image of the third road segment by way of back-projection based on the third 3D vector map obtained by the fusion unit 10032, each of the first optimized camera poses obtained by the first optimization module, and each of the second optimized camera poses obtained by the second optimization module.
[0187] A second generation unit 10042, configured to generate an aerial view of the third road segment based on the third 3D vector map obtained by the fusion unit 10032, the first optimized camera poses obtained by the first optimization module, and the second optimized camera poses obtained by the second optimization module.
[0188] A third generation unit 10043, configured to generate 2D annotation data of static objects in the third road segment based on the 2D images obtained by the first generation unit 10041.
[0189] A fourth generation unit 10044, configured to generate aerial annotation data of static objects in the third road segment based on the aerial view obtained by the second generation unit 10042.
[0190] In another exemplary embodiment of the present disclosure, the annotation data generation module 1004 may further include:
[0191] A fifth generation unit, configured to generate first 3D annotation box data of a dynamic object at the first moment based on each of the first images and the first camera poses respectively corresponding to the first images.
[0192] A sixth generation unit, configured to generate second 3D annotation box data of the dynamic object at the second moment based on each of the second images and the second camera poses respectively corresponding to the second images.
[0193] A seventh generation unit, configured to generate annotation data of the dynamic object at the first moment based on the third 3D vector map obtained by the fusion unit 10032, the first optimized camera poses obtained by the first optimization module, and the first 3D annotation box data obtained by the fifth generation unit.
[0194] An eighth generation unit, configured to generate annotation data of the dynamic object at the second moment based on the third 3D vector map obtained by the fusion unit 10032, the second optimized camera poses obtained by the second optimization module, and the second 3D annotation box data obtained by the sixth generation unit.
[0195] In another exemplary embodiment of the present disclosure, the first image set includes multiple frames of the first images obtained by a multi-camera system on a vehicle while simultaneously photographing the first road segment from different angles; the second image set includes multiple frames of the second images obtained by the multi-camera system on the vehicle while simultaneously photographing the second road segment from different angles.
[0196] Exemplary Electronic Device
[0197] Next, refer to Figure 12Describe an electronic device according to an embodiment of the present disclosure. The electronic device may be a terminal device for generating annotation data, such as Figure 1 Any of the terminal devices 100 shown and the server for generating annotation data. Alternatively, the electronic device may also be a stand-alone device independent of the foregoing devices, and the stand-alone device may communicate with each of the foregoing devices to receive the input signals collected therefrom.
[0198] Figure 12 The block diagram of an electronic device according to an embodiment of the present disclosure is illustrated.
[0199] As Figure 12 Shown, the electronic device 10 includes one or more processors 11 and a memory 12.
[0200] The processor 11 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.
[0201] The memory 12 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 11 may run the program instructions to implement the methods for generating annotation data of various embodiments of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, noise components, etc. may also be stored in the computer-readable storage media.
[0202] In one example, the electronic device 10 may further include: an input device 13 and an output device 14, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0203] For example, when the electronic device is a terminal device for generating annotation data, the input device 13 may be a microphone or a microphone array of the above terminal device for capturing the input signal of the sound source. When the electronic device is a stand-alone device, the input device 13 may be a communication network connector for receiving the input signals collected from the terminal device.
[0204] In addition, the input device 13 may further include, for example, a keyboard, a mouse, and so on.
[0205] The output device 14 can output various information to the outside, including the determined distance information, direction information, etc. The output device 14 can include, for example, a display, a speaker, a printer, a communication network, and remote output devices connected thereto, etc.
[0206] Of course, for simplicity, Figure 12 only some of the components related to the present disclosure in the electronic device 10 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 10 may further include any other appropriate components.
[0207] Exemplary Computer Program Product and Computer Readable Storage Medium
[0208] In addition to the above methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps in the method for generating annotation data according to various embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.
[0209] The computer program product can be written in any combination of one or more programming languages for the program code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0210] Furthermore, an embodiment of the present disclosure may also be a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the method for generating annotation data according to various embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.
[0211] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0212] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. Additionally, the above-disclosed specific details are only for illustrative and easy-to-understand purposes and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0213] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with each other.
[0214] It should also be noted that in the devices, equipment, and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.
[0215] The above description of the disclosed aspects enables any person skilled in the art to make or use the present disclosure. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0216] The foregoing description has been presented for purposes of illustration and description. In addition, this description is not intended to limit embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those of skill in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A method for generating labeled data, comprising: Based on a first image set obtained by photographing a first road section at a first moment, constructing a first three-dimensional vector map of the first road section; Based on a second image set obtained by photographing a second road section at a second moment, constructing a second three-dimensional vector map of the second road section; the area where the second road section is located overlaps with the area where the first road section is located; Based on the first three-dimensional vector map and the second three-dimensional vector map, clustering and fusing to generate a third three-dimensional vector map of a third road section; The area where the third road section is located covers the area where the first road section is located and the area where the second road section is located; Based on the third three-dimensional vector map, generating labeled data of static objects in the third road section; Wherein, the constructing the first three-dimensional vector map of the first road section based on the first image set obtained by photographing the first road section at the first moment includes: Performing semantic segmentation processing on each frame of the first image included in the first image set, or performing semantic segmentation processing and object detection processing, to obtain first preprocessed images respectively corresponding to the first images; the first preprocessed image of the first image includes semantic information and initial labeling information of static objects in the first road section; Obtaining first camera poses respectively corresponding to the first images; Based on the first preprocessed images and the first camera poses respectively corresponding to the first preprocessed images, constructing the first three-dimensional vector map.
2. The method according to claim 1, wherein, The method further includes: Based on a third image set obtained by photographing a fourth road section at a third moment, constructing a fourth three-dimensional vector map of the fourth road section; the area where the first road section is located and the area where the second road section is located both overlap with the area where the fourth road section is located; The clustering and fusing based on the first three-dimensional vector map and the second three-dimensional vector map to generate the third three-dimensional vector map of the third road section includes: Based on the first three-dimensional vector map, the second three-dimensional vector map and the fourth three-dimensional vector map, clustering and fusing to generate the third three-dimensional vector map of the third road section, wherein the area where the third road section is located covers the area where the fourth road section is located.
3. The method according to claim 1, wherein, The constructing the second three-dimensional vector map of the second road section based on the second image set obtained by photographing the second road section at the second moment includes: Performing semantic segmentation processing on each frame of the second image included in the second image set, or performing semantic segmentation processing and object detection processing, to obtain second preprocessed images respectively corresponding to the second images; Obtaining second camera poses respectively corresponding to the second images; Based on the second preprocessed images and the second camera poses respectively corresponding to the second preprocessed images, constructing the second three-dimensional vector map.
4. The method according to claim 3, wherein, The clustering and fusing based on the first three-dimensional vector map and the second three-dimensional vector map to generate the third three-dimensional vector map of the third road section includes: Aligning the first three-dimensional vector map and the second three-dimensional vector map; Based on the aligned first three-dimensional vector map and second three-dimensional vector map, clustering and fusing to generate the third three-dimensional vector map.
5. The method according to claim 4, wherein After aligning the first three-dimensional vector map and the second three-dimensional vector map, the method further includes: Based on the aligned first three-dimensional vector map, optimizing each of the first camera poses to obtain first optimized camera poses respectively corresponding to the first camera poses; Based on the aligned second three-dimensional vector map, optimizing each of the second camera poses to obtain second optimized camera poses respectively corresponding to the second camera poses; The generating the annotation data of the static objects in the third road segment based on the third three-dimensional vector map includes: Generating the annotation data of the static objects in the third road segment based on the third three-dimensional vector map, each of the first optimized camera poses, and each of the second optimized camera poses.
6. The method according to claim 5, wherein, The generating the annotation data of the static objects in the third road segment based on the third three-dimensional vector map, each of the first optimized camera poses, and each of the second optimized camera poses includes: Generating a two-dimensional image of the third road segment by back-projection based on the third three-dimensional vector map, each of the first optimized camera poses, and each of the second optimized camera poses; Generating an aerial view of the third road segment based on the third three-dimensional vector map, each of the first optimized camera poses, and each of the second optimized camera poses; Generating two-dimensional annotation data of the static objects in the third road segment based on the two-dimensional image; Generating aerial view annotation data of the static objects in the third road segment based on the aerial view.
7. The method according to claim 5 or 6, wherein The method further includes: Generating first three-dimensional annotation box data of the dynamic object at the first moment based on each of the first images and the first camera poses respectively corresponding to the first images; Generating second three-dimensional annotation box data of the dynamic object at the second moment based on each of the second images and the second camera poses respectively corresponding to the second images; Generating annotation data of the dynamic object at the first moment based on the third three-dimensional vector map, each of the first optimized camera poses, and the first three-dimensional annotation box data; Generating annotation data of the dynamic object at the second moment based on the third three-dimensional vector map, each of the second optimized camera poses, and the second three-dimensional annotation box data.
8. The method according to any one of claims 1 to 6, wherein, The first image set includes multiple frames of the first images obtained by a multi-camera system on a vehicle simultaneously shooting the first road segment from different angles; the second image set includes multiple frames of the second images obtained by the multi-camera system on the vehicle simultaneously shooting the second road segment from different angles.
9. A device for generating annotation data, comprising: A first image construction module, configured to construct a first three-dimensional vector map of the first road segment based on a first image set obtained by shooting the first road segment at a first moment; A second image construction module, configured to construct a second three-dimensional vector map of the second road segment based on a second image set obtained by shooting the second road segment at a second moment; the area where the second road segment is located overlaps with the area where the first road segment is located; A fusion module, configured to cluster and fuse a first three-dimensional vector map obtained by the first image construction module and a second three-dimensional vector map obtained by the second image construction module to generate a third three-dimensional vector map of a third road segment; an area where the third road segment is located covers an area where the first road segment is located and an area where the second road segment is located; An annotation data generation module, configured to generate annotation data of static objects in the third road segment based on the third three-dimensional vector map obtained by the fusion module; Wherein, the first image construction module includes: A first preprocessing unit, configured to perform semantic segmentation processing, or perform semantic segmentation processing and target detection processing, on each frame of the first image included in the first image set, to obtain a first preprocessed image corresponding to each of the first images; the first preprocessed image of the first image includes semantic information and initial annotation information of static objects in the first road segment; A first acquisition unit, configured to acquire a first camera pose corresponding to each of the first images; A first construction unit, configured to construct the first three-dimensional vector map based on each of the first preprocessed images and the first camera pose corresponding to each of the first preprocessed images.
10. A computer-readable storage medium, storing a computer program, where the computer program is used to execute the method for generating annotation data according to any one of claims 1-8 above.
11. An electronic device, including: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method for generating annotation data according to any one of claims 1-8 above.
Citation Information
Patent Citations
Point cloud data annotation method, segmentation model determination method, target detection method and related equipment
CN110264468A
Method, device and system for cooperatively constructing point cloud map
CN111681172A
Road picture marking method and device for lane line identification
CN113205447A