Method, device and storage medium for processing annotation information

By comparing the annotation information of point cloud data with the instance segmentation results of multiple reference models, multiple evaluation parameters are determined, and a decision model is used for comprehensive evaluation. This solves the problems of low accuracy and efficiency of point cloud data annotation information in traditional methods, and achieves efficient and accurate annotation information judgment.

CN122223682APending Publication Date: 2026-06-16BEIJING VOYAGER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING VOYAGER TECH CO LTD
Filing Date
2024-12-13
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

The technical challenge of point cloud data annotation in existing technologies lies in the fact that traditional point cloud data annotation methods are costly and inefficient, and traditional quality inspection methods cannot effectively solve the technical problem of how to improve the accuracy and efficiency of annotation information.

Method used

By comparing the annotation information of point cloud data with the instance segmentation results of multiple reference models, several evaluation parameters are determined, including projection region consistency, point set distance, point set consistency, point cloud reflectivity consistency, and shape consistency. A decision model is then used for comprehensive evaluation.

Benefits of technology

It enables the accurate judgment of point cloud data annotation information, improves the judgment efficiency, can accurately identify the accuracy of annotation information, and is adaptable to various judgment scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223682A_ABST
    Figure CN122223682A_ABST
Patent Text Reader

Abstract

According to embodiments of the present disclosure, a method, apparatus and storage medium for processing annotation information are provided. The method comprises obtaining annotation information of point cloud data, the annotation information indicating a target instance in the point cloud data; obtaining a plurality of instance segmentation results of a plurality of reference models, the plurality of instance segmentation results comprising a first segmentation result generated by an image segmentation model based on a target image associated with the point cloud data and a second segmentation result generated by a point cloud segmentation model based on the point cloud data; determining a first set of evaluation parameters based on a comparison between the annotation information and the plurality of instance segmentation results; and determining a target evaluation of the annotation information based on at least the first set of evaluation parameters. Thus, embodiments of the present disclosure can compare the annotation information of the point cloud data with a plurality of instance segmentation results obtained based on a plurality of reference models, accurately determine whether the annotation information of the point cloud data is accurately annotated, and improve the efficiency of determining whether the annotation information is accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, devices, computer-readable storage media, and computer program products for processing annotation information. Background Technology

[0002] With the rapid development of technology, point cloud data collected by sensors is being applied in various fields, such as autonomous driving. The demand for high-precision point cloud data annotation is also increasing. Accurate annotation of point cloud data is crucial for training Bird's Eye View (BEV) models and any other suitable models. Improving the annotation accuracy of point cloud data is a key focus. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for processing annotation information is provided. The method includes: acquiring annotation information of point cloud data, the annotation information indicating target instances in the point cloud data; acquiring multiple instance segmentation results from multiple reference models, the multiple instance segmentation results including a first segmentation result generated by an image segmentation model based on a target image associated with the point cloud data and a second segmentation result generated by a point cloud segmentation model based on the point cloud data; determining a first set of evaluation parameters based on a comparison of the annotation information and the multiple instance segmentation results; and determining a target evaluation of the annotation information based at least on the first set of evaluation parameters.

[0004] In a second aspect of this disclosure, an apparatus for processing annotation information is provided. The apparatus includes: a first acquisition module configured to acquire annotation information of point cloud data, the annotation information indicating target instances in the point cloud data; a second acquisition module configured to acquire multiple instance segmentation results from multiple reference models, the multiple instance segmentation results including a first segmentation result generated by an image segmentation model based on a target image associated with the point cloud data and a second segmentation result generated by a point cloud segmentation model based on the point cloud data; a first determination module configured to determine a first set of evaluation parameters based on a comparison of the annotation information and the multiple instance segmentation results; and a second determination module configured to determine a target evaluation of the annotation information based at least on the first set of evaluation parameters.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method of the first aspect.

[0008] It should be understood that the content described in this summary section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;

[0011] Figure 2 A schematic diagram illustrating the process of processing abnormal annotation information according to some embodiments of the present disclosure is shown;

[0012] Figure 3 This disclosure illustrates a flowchart of annotation information verification according to some embodiments of this disclosure;

[0013] Figure 4 A schematic structural block diagram of an apparatus for processing annotation information according to certain embodiments of the present disclosure is shown;

[0014] Figure 5 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0017] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0018] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0019] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0020] As used in this paper, the term "model" refers to a system that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. In this paper, "model" may also be referred to as a "machine learning model," a "machine learning network," or simply a "network," and these terms are used interchangeably.

[0021] As mentioned earlier, to improve the accuracy of point cloud data annotation, the annotation information of point cloud data can be quality checked to determine whether the standard information of the point cloud data is accurate. Compared with the annotation information of two-dimensional images, point cloud data has greater complexity, larger data volume, and a more cumbersome annotation process, which makes the quality check of point cloud data annotation information more technically challenging.

[0022] Traditional methods for quality control of point cloud data annotations primarily rely on manual inspection, which is costly and inefficient. Furthermore, due to the massive amount of annotation data required for training models, traditional methods can only assess the overall quality of data from the same batch through sampling, failing to provide fine-grained screening and hindering the timely detection and correction of annotation errors.

[0023] Embodiments of this disclosure provide a scheme for processing annotation information. According to various embodiments of this disclosure, annotation information of point cloud data is obtained, the annotation information indicating target instances in the point cloud data; multiple instance segmentation results from multiple reference models are obtained, the multiple instance segmentation results including a first segmentation result generated by an image segmentation model based on a target image associated with the point cloud data and a second segmentation result generated by a point cloud segmentation model based on the point cloud data; a first set of evaluation parameters is determined based on a comparison of the annotation information and the multiple instance segmentation results; and a target evaluation of the annotation information is determined based at least on the first set of evaluation parameters.

[0024] Therefore, the embodiments of this disclosure can compare the segmentation results of multiple instances obtained from multiple reference models with the annotation information of point cloud data, which can accurately determine whether the annotation information of point cloud data is accurate and improve the efficiency of determining whether the annotation information is accurate.

[0025] Example Environment

[0026] Figure 1 A schematic diagram of an example environment 100 in which several embodiments of the present disclosure can be implemented is shown.

[0027] Electronic device 110 is schematically shown in this example environment 100. Electronic device 110 can acquire annotation information corresponding to point cloud data, which can indicate target instances in the point cloud data. Target instances can be any appropriate object that can be identified and distinguished in three-dimensional space, such as people, animals, vehicles, traffic signs and signals (such as traffic lights, speed limit signs), roads and infrastructure (such as lane lines, road boundaries, guardrails), etc.

[0028] In some embodiments, this point cloud data can be 3D data collected by sensors pre-installed on the vehicle in an autonomous driving scenario. The vehicle can be any type of vehicle capable of carrying people and / or objects and moving via a power system such as an engine. Examples of vehicles include, but are not limited to, cars, trucks, buses, electric vehicles, motorcycles, RVs, trains, etc. As an example, the vehicle disclosed herein can be a vehicle with some level of assisted driving capability or autonomous driving capability; such vehicles are also referred to as intelligent driving vehicles. The sensors can be any suitable sensor capable of collecting point cloud data, such as LiDAR, radar, stereo cameras, or other 3D sensors. Of course, this point cloud data can also be data obtained by sensors in scenarios other than autonomous driving scenarios, which will not be elaborated upon here.

[0029] Furthermore, the electronic device 120 can determine whether the annotation information is accurate based on the annotation information of the acquired point cloud data.

[0030] Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).

[0031] Electronic device 110 can also be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Electronic device 110 may include, for example, computing systems / servers, such as mainframes, edge computing nodes, computing devices in cloud environments, and so on.

[0032] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0033] Example process

[0034] Figure 2 A flowchart of a process 200 for processing annotation information according to some embodiments of the present disclosure is shown. Process 200 can be implemented at electronic device 110. Reference is made below. Figure 1 Describe the process 200.

[0035] In box 210, electronic device 110 acquires annotation information of point cloud data, the annotation information indicating target instances in the point cloud data.

[0036] In some embodiments, point cloud data can be three-dimensional data obtained by capturing the vehicle's surrounding environment using any suitable sensor. This sensor can be LiDAR, radar, a stereo camera, or other three-dimensional sensors, etc. The sensor can measure and record light scattered or reflected from the surface of the target instance, thereby generating a set of surface points containing the target instance. For each point in the measured surrounding environment, the corresponding point cloud data can include spatial location information (typically X, Y, Z coordinates) and other attribute information, such as color, intensity, or temperature.

[0037] In some embodiments, the target instance can be any suitable object that can be identified and distinguished in three-dimensional space, such as people, animals, vehicles, traffic signs and signals (such as traffic lights, speed limit signs), roads and infrastructure (such as lane lines, road boundaries, guardrails), etc.

[0038] As an example, the annotation information in this point cloud data can indicate lane lines.

[0039] In box 220, electronic device 110 acquires multiple instance segmentation results of multiple reference models, including a first segmentation result generated by an image segmentation model based on a target image associated with point cloud data and a second segmentation result generated by a point cloud segmentation model based on point cloud data.

[0040] In some embodiments, the plurality of reference models can be any suitable machine learning model for segmenting instances, including not only image segmentation models and point cloud segmentation models, but also any other suitable segmentation models.

[0041] In some embodiments, the first segmentation result may indicate information corresponding to instances included in the target image identified and segmented based on the target image. In some embodiments, the target image may be a two-dimensional image corresponding to the vehicle's surrounding environment, which is synchronously acquired with point cloud data. This two-dimensional image may be acquired by an image sensor installed at a predetermined location on the vehicle, and the acquisition range of this image sensor is the same as that of the sensor that acquired the point cloud data, or the acquisition range of this image sensor includes the acquisition range of the sensor that acquired the point cloud data.

[0042] In some embodiments, the second segmentation result may indicate information about the individual instances identified and segmented based on point cloud data.

[0043] In some embodiments, the information of each instance indicated by the first segmentation result and the second segmentation result may include the location information of each identified and segmented instance, the type corresponding to each instance, and so on.

[0044] In box 230, electronic device 110 determines the first set of evaluation parameters based on the comparison between the annotation information and the segmentation results of multiple instances.

[0045] As an example, electronic device 110 can determine this first set of evaluation parameters based on a comparison between the annotation information and a first segmentation result, and based on a comparison between the annotation information and a second segmentation result. This first set of evaluation parameters can be parameters of predetermined dimensions that characterize whether the annotation information is accurate.

[0046] In some embodiments, the first set of evaluation parameters may include a first evaluation parameter and / or a second evaluation parameter, and may also include other evaluation parameters besides the first evaluation parameter and the second evaluation parameter.

[0047] The process of determining the first and second evaluation parameters will be explained below.

[0048] In some embodiments, the electronic device 110 can determine the projection region of the target instance in the target image, that is, the electronic device 110 can project point cloud data onto the corresponding two-dimensional image plane, and the region formed by the points of the point cloud data projected onto the two-dimensional image plane is the projection region in the target image. Further, the electronic device 110 can determine a first evaluation parameter based on a comparison between the projection region and a reference region indicated by a first segmentation result. The reference region indicated by the first segmentation result is the region in the target image corresponding to the target instance identified and segmented based on the target image. The first evaluation parameter can be any appropriate parameter that can characterize the consistency between the projection region and the reference region, wherein the more consistent the projection region and the reference region are, the greater the likelihood that the annotation information corresponding to the point cloud data is accurate.

[0049] As an example, the first evaluation parameter can be the area / number of points corresponding to the intersection. Specifically, the electronic device 110 can determine the intersection area between the projected area and the reference area. Further, the electronic device 110 can determine the area / number of points corresponding to this intersection area based on this intersection area.

[0050] As another example, the first evaluation parameter can be the intersection-union ratio (IU) between the projected region and the reference region. Specifically, the electronic device 110 can determine the intersection region between the projected region and the reference region, and determine the union region between the projected region and the reference region. Further, the electronic device 110 can determine the IU between the projected region and the reference region based on this intersection region and this union region.

[0051] In some embodiments, the electronic device 110 can determine the distance from the target point set corresponding to the annotation information to the reference point set indicated by the second segmentation result based on a comparison between the annotation information and the second segmentation result. The target point set corresponding to the annotation information is the set of points corresponding to the point cloud data projected onto a two-dimensional plane. The reference point set can be the set of points corresponding to other instances that are consistent with the classification corresponding to the target instance indicated by the second segmentation result and have the smallest distance to the target point set corresponding to the target instance. In some embodiments, the distance can be based on any suitable distance, such as Euclidean distance, Manhattan distance, etc. In some embodiments, this distance can be the distance between two points at predetermined positions corresponding to the target point set and the reference point set, such as the distance between the center point of the target point set and the center point of the reference point set. In other embodiments, this distance can be the average distance between the points included in the target point set and the corresponding points in the reference point set.

[0052] As an example, the distance from the target point set to the reference point set indicated by the second segmentation result can be determined based on the following formula:

[0053]

[0054] Where pi is the coordinate information of point i in the reference point set in the two-dimensional plane, l is the coordinate information of the point corresponding to point i in the target point set, ‖pi-l‖2 is the Euclidean distance between point i and the point corresponding to point i, N is the number of common points (point 1, point 2... point N) determined based on the reference point set and the target point set, and Sdist is the distance from the target point set to the reference point set indicated by the second segmentation result.

[0055] Furthermore, the electronic device 110 can determine a second evaluation parameter based on the distance. The second evaluation parameter can be any appropriate parameter that can characterize the consistency between the target point set and the reference point set, wherein the more consistent the target point set and the reference point set are (the smaller the distance), the greater the likelihood that the annotation information corresponding to the point cloud data is accurate.

[0056] As an example, electronic device 110 can determine this distance as a second evaluation parameter.

[0057] Since the information corresponding to the same instance observed at multiple consecutive time points exhibits spatial consistency, to improve the accuracy of determining whether the annotation information is accurate, as another example, the electronic device 110 can determine multiple point sets of the target instance at multiple time points based on the annotation information. These multiple point sets are collections of points corresponding to the target instance at different time points, and these multiple point sets can characterize the position and shape of the target instance in consecutive time frames. Further, the electronic device 110 can determine a second evaluation parameter based on the consistency degree and distance of the multiple point sets. The second evaluation parameter is positively correlated with the consistency degree and negatively correlated with the distance. The consistency degree of these multiple point sets can reflect the stability of the shape and position of the target instance at different time points. In some embodiments, the consistency degree of these multiple point sets can be determined based on the intersection-union ratio (IU) corresponding to these multiple point sets.

[0058] As an example, electronic device 110 can determine the intersection-union ratio corresponding to these multiple point sets based on the following formula:

[0059]

[0060] Where Siou is the intersection-union ratio of these multiple point sets, Poverlap is the area / number of points of the region corresponding to the intersection of these multiple point sets, and Punion is the area / number of points of the region corresponding to the union of these multiple point sets.

[0061] As an example, electronic device 110 can determine the second evaluation parameter based on the following formula:

[0062] score = Siou * (1 - Sdist)

[0063] Where Siou is the intersection-union ratio corresponding to these multiple point sets, Sdist is the distance from the target point set to the reference point set indicated by the second segmentation result, and score is the second evaluation parameter.

[0064] In box 240, electronic device 110 determines the target evaluation of the labeled information based at least on the first set of evaluation parameters.

[0065] In order to accurately and efficiently determine whether the annotation information of point cloud data is accurate, in some embodiments, in addition to determining the target evaluation based on the first set of evaluation parameters (the dimension corresponding to the sensor information), the electronic device 110 can also determine the target evaluation by comprehensively considering other evaluation parameters (such as the second set of evaluation parameters and / or the third set of evaluation parameters).

[0066] The following section explains the process of determining the second and third sets of evaluation parameters.

[0067] In some embodiments, the electronic device 110 may determine a second set of evaluation parameters based on a comparison of prior information and annotation information associated with the target instance. The second set of evaluation parameters may be any appropriate parameters that characterize whether the annotation information is consistent with the prior information, wherein the more consistent the annotation information is with the prior information, the greater the likelihood that the annotation information corresponding to the point cloud data is accurate.

[0068] In some embodiments, prior information may indicate, but is not limited to, at least one of the following: point cloud reflectance corresponding to the target instance, and shape features corresponding to the target instance.

[0069] In some embodiments, during sensor scanning, the point cloud reflectance corresponding to instances with different materials and surface properties may vary. For example, the point cloud reflectance corresponding to road markings such as lane lines differs significantly from that corresponding to other locations. In some embodiments, prior information may be the range of point cloud reflectance corresponding to the target instance. In some embodiments, the point cloud reflectance corresponding to the target instance may change in real time based on environmental information, such as weather, external resistance, etc.

[0070] In some embodiments, in response to the prior information indicating the point cloud reflectance corresponding to the target instance, the second set of evaluation parameters may include dynamic consistency of the point cloud reflectance. Dynamic consistency characterizes whether the point cloud reflectance corresponding to the target instance in different environments is consistent with the point cloud reflectance corresponding to the different environments indicated in the prior information.

[0071] In some embodiments, the electronic device 110 may also determine a second set of evaluation parameters based on the temporal consistency of the target entity. Temporal consistency characterizes whether the point cloud reflectance of the target instance is consistent across multiple time points.

[0072] In some embodiments, in response to this prior information indicating the shape features corresponding to the target instance, the second set of evaluation parameters may include shape consistency.

[0073] In some embodiments, the electronic device 110 can divide space into multiple grids at equal intervals. Further, the electronic device 110 can compute geometric features, such as curvature and normal vectors, for local points within each grid. Further, the electronic device 110 can determine target shape features based on Principal Component Analysis (PCA), whereby the target shape features are the shape features corresponding to the target instance and can exist in vector form. Further, the electronic device 110 can compare the target shape features with the shape features corresponding to the target instance indicated in prior information to determine shape consistency.

[0074] As an example, electronic device 110 can determine shape consistency based on the following formula:

[0075]

[0076] Here, λ1, λ2, and λ3 can represent the various shape features corresponding to the detected target instances, λ1′, λ2′, and λ3′ are the various shape features corresponding to the target instances indicated in the prior information, σ is a scaling factor that can be set according to requirements, and Sshape represents shape consistency.

[0077] In some embodiments, the annotation information may include four-dimensional annotation information. Four-dimensional annotation information may include, but is not limited to, any suitable dimension other than the dimension corresponding to spatial location (three-dimensional coordinates), such as information in the time dimension.

[0078] In some embodiments, the electronic device 110 can construct a global map based on multi-frame point cloud data. This multi-frame point cloud data can be point cloud data corresponding to target instances at multiple time points. The electronic device 110 can fully track the state of the target instance in consecutive time frames based on the global map.

[0079] As an example, electronic device 110 can utilize a predetermined first algorithm to determine inter-frame pose relationships based on multi-frame point cloud data. The predetermined first algorithm can be the Generalized Iterative Closest Point (GICP) algorithm, or any other suitable algorithm capable of determining inter-frame pose relationships. Furthermore, electronic device 110 can perform temporal registration of consecutive frames with labeled information based on a predetermined second algorithm to establish a global map. The predetermined second algorithm can be Simultaneous Localization and Mapping (SLAM), or any other suitable global map construction algorithm.

[0080] Furthermore, the electronic device 110 can construct a pose graph based on four-dimensional annotation information. Specifically, the electronic device 110 can establish a pose graph using the matching constraints between the four-dimensional annotation information and local frames of the base map data. The pose graph can include multiple nodes, and any two nodes may be connected by edges. Each node can represent point cloud data at different time points, and the edges represent the spatial relationship and temporal continuity between the point cloud data corresponding to two nodes.

[0081] Furthermore, the electronic device 110 can determine a third set of evaluation parameters based on the optimization results of the pose graph. These third set of evaluation parameters may include, but are not limited to, the optimization quality of the pose graph, the consistency of nodes and edges, and the stability of the entire pose graph.

[0082] As an example, the electronic device 110 can use the average residual value of the constrained edges after graph optimization, based on the annotation information of each frame, as a third set of evaluation parameters.

[0083] The process for determining the target score is explained below:

[0084] In some embodiments, the electronic device 110 can determine the target evaluation of the labeled information based on any one of the first set of evaluation parameters, the second set of evaluation parameters, and the third set of evaluation parameters.

[0085] In other embodiments, the electronic device 110 may determine the target evaluation of the labeled information based on a first set of evaluation parameters and a second set of evaluation parameters.

[0086] In other embodiments, the electronic device 110 may also determine the target evaluation of the labeled information based on the first set of evaluation parameters and the third set of evaluation parameters.

[0087] In other embodiments, the electronic device 110 may also determine the target evaluation of the labeled information based on the first set of evaluation parameters, the second set of evaluation parameters, and the third set of evaluation parameters.

[0088] In some embodiments, the electronic device 110 provides multiple sets of evaluation parameters to the decision model. These multiple sets of evaluation parameters may include a first set of evaluation parameters, a second set of evaluation parameters, a third set of evaluation parameters, and any number of other sets of evaluation parameters. The decision model can be any suitable machine learning model that uses a decision tree to determine the target evaluation of labeled information.

[0089] Furthermore, the electronic device 110 can acquire the target evaluation generated by the decision model. The target evaluation can indicate whether the annotation information corresponding to the point cloud data is accurate or the annotation accuracy rate of the annotation information corresponding to the point cloud data, etc.

[0090] This disclosure can output multi-dimensional evaluation results based on multiple evaluation parameters, and can accurately determine whether the annotation information is accurate. Furthermore, this disclosure generates target evaluations based on a decision model, which allows for more flexible judgment of the accuracy of annotation information and better adaptability to various judgment scenarios.

[0091] The training process of the decision-making model will be explained below.

[0092] In some embodiments, the electronic device 110 can use preset complete annotation information to generate a large amount of positive sample point cloud data and negative sample point cloud data to form a training sample set, wherein the positive sample point cloud data are sample point cloud data with correct annotation information, and the negative sample point cloud data are sample point cloud data with incorrect annotation information.

[0093] Furthermore, the electronic device 110 can acquire multiple sets of sample evaluation parameters corresponding to each training sample in the training sample set. The process of determining multiple sets of evaluation parameters corresponding to the same point cloud data is similar and will not be elaborated here.

[0094] Furthermore, the electronic device 110 can input these multiple sets of sample evaluation parameters into the decision model to be trained, so as to train the decision tree model based on ensemble learning methods such as Random Forest. Random Forest is an algorithm that integrates multiple decision trees, which can improve the generalization ability and accuracy of the model.

[0095] Furthermore, the electronic device 110 can obtain a fully trained decision model in response to the achievement of predetermined training conditions. The predetermined training conditions may include, but are not limited to, a threshold for training time, a threshold for the number of training iterations, a minimum target loss value, etc.

[0096] Figure 3 This disclosure illustrates a flowchart of annotation information verification according to some embodiments of this disclosure. Figure 3 Please provide an explanation.

[0097] In box 301, electronic device 110 inputs the annotation information of point cloud data.

[0098] In some embodiments, point cloud data can be three-dimensional data obtained by capturing the vehicle's surrounding environment using any suitable sensor. This sensor can be LiDAR, radar, a stereo camera, or other three-dimensional sensors, etc. The sensor can measure and record light scattered or reflected from the surface of the target instance, thereby generating a set of surface points containing the target instance. For each point in the measured surrounding environment, the corresponding point cloud data can include spatial location information (typically X, Y, Z coordinates) and other attribute information, such as color, intensity, or temperature.

[0099] In some embodiments, the annotation information of the point cloud data indicates target instances in the point cloud data. Target instances can be any suitable object that can be identified and distinguished in three-dimensional space, such as people, animals, vehicles, traffic signs and signals (such as traffic lights, speed limit signs), roads and infrastructure (such as lane lines, road boundaries, guardrails), etc.

[0100] In box 302, electronic device 110 performs cross-validation of multi-sensor information.

[0101] In some embodiments, the multi-sensor information may include image data acquired by an image sensor and point cloud data acquired by a lidar sensor. In some embodiments, after the electronic device 110 performs cross-validation of the multi-sensor information, it executes the operation in block 303.

[0102] In some embodiments, the electronic device 110 can determine the projection region of the target instance in the target image, that is, the electronic device 110 can project point cloud data onto the corresponding two-dimensional image plane, and the region formed by the points of the point cloud data projected onto the two-dimensional image plane is the projection region in the target image. Further, the electronic device 110 can compare the projection region with the reference region indicated by the first segmentation result.

[0103] In some embodiments, the electronic device 110 can compare the annotation information with the second segmentation result.

[0104] In box 303, electronic device 110 obtains multimodal verification parameters.

[0105] In some embodiments, the multimodal verification parameters may include a first evaluation parameter and a second evaluation parameter.

[0106] In some embodiments, the electronic device may determine a first evaluation parameter based on a comparison between the projected region and a reference region indicated by a first segmentation result. In some embodiments, the electronic device 110 may determine the distance from the target point set corresponding to the annotation information to the reference point set indicated by the second segmentation result based on a comparison between the annotation information and the second segmentation result, and determine a second evaluation parameter based on the distance.

[0107] In box 304, electronic device 110 performs timing verification.

[0108] In some embodiments, the electronic device 110 may perform timing verification before executing the operation of block 305.

[0109] In some embodiments, the electronic device 110 can utilize the GICP algorithm to determine inter-frame pose relationships based on multi-frame point cloud data. Further, the electronic device 110 can use the SLAM algorithm to perform temporal registration of consecutive frames of labeled information to establish a global map. Further, the electronic device 110 can construct a pose graph based on four-dimensional labeled information. Specifically, the electronic device 110 can utilize the matching constraints between four-dimensional labeled information and local frames of the base map data to establish the pose graph. The pose graph can include multiple nodes, and any two nodes may be connected by edges. Each node can represent point cloud data at different time points, and the edges represent the spatial relationship and temporal continuity between the point cloud data corresponding to two nodes. This pose graph is an intermediate product of the temporal verification process.

[0110] In box 305, electronic device 110 obtains frame conformance verification parameters.

[0111] In some embodiments, the electronic device 110 may determine whether the timing is consistent based on the optimization results of the pose graph.

[0112] In box 306, electronic device 110 performs feature verification.

[0113] In some embodiments, the electronic device 110 may verify the annotation information based on the point cloud reflectance and shape features associated with the target instance. In some embodiments, the electronic device may execute the operation of block 307 in response to the completion of feature verification.

[0114] As an example, the electronic device 110 can compare the detected point cloud reflectance of the target instance under different environments with the point cloud reflectance under different environments indicated in the prior information.

[0115] As another example, the electronic device 110 can divide the space into multiple grids at equal intervals. Further, the electronic device 110 can compute the curvature and normal vector of local points within each grid. Further, the electronic device 110 can determine the target shape feature based on PCA; this target shape feature is the shape feature corresponding to the target instance, which can exist in vector form. Further, the electronic device 110 can compare the target shape feature with the shape feature corresponding to the target instance indicated in prior information.

[0116] In box 307, electronic device 110 obtains point cloud reflectance and shape verification parameters.

[0117] In some embodiments, the electronic device 110 can determine point cloud reflection verification parameters based on a comparison between the detected point cloud reflectance of the target instance under different environments and the point cloud reflectance under different environments indicated in the prior information.

[0118] In some embodiments, the electronic device may determine shape verification parameters based on a comparison between the target shape features and the shape features corresponding to the target instance indicated in prior information.

[0119] In box 308, the aggregated verification results of electronic device 110 are displayed.

[0120] In some embodiments, the electronic device has already obtained verification results such as point cloud reflection verification parameters, shape verification parameters, frame consistency verification parameters, and multimodal verification parameters.

[0121] In box 309, electronic device 110 inputs the verification results into the decision model.

[0122] In some embodiments, the decision model can be any suitable machine learning model that uses a decision tree to determine the target evaluation of the labeled information.

[0123] In box 310, electronic device 110 obtains the target evaluation of the labeled information output by the decision model.

[0124] In some embodiments, target evaluation can indicate whether the annotation information corresponding to the point cloud data is accurate or the annotation accuracy rate of the annotation information corresponding to the point cloud data, etc.

[0125] Therefore, the embodiments of this disclosure can compare the segmentation results of multiple instances obtained from multiple reference models with the annotation information of point cloud data, which can accurately determine whether the annotation information of point cloud data is accurate and improve the efficiency of determining whether the annotation information is accurate.

[0126] Example devices and equipment

[0127] Figure 4 A schematic structural block diagram of an apparatus 400 for processing annotation information according to certain embodiments of the present disclosure is shown. The apparatus 400 may be implemented as or included in an electronic device 110. Various modules / components in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0128] As shown in the figure, the device 400 includes a first acquisition module 410 configured to acquire annotation information of point cloud data, wherein the annotation information indicates target instances in the point cloud data; a second acquisition module 420 configured to acquire multiple instance segmentation results of multiple reference models, wherein the multiple instance segmentation results include a first segmentation result generated by an image segmentation model based on a target image associated with the point cloud data and a second segmentation result generated by a point cloud segmentation model based on the point cloud data; a first determination module 430 configured to determine a first set of evaluation parameters based on a comparison between the annotation information and the multiple instance segmentation results; and a second determination module 440 configured to determine the target evaluation of the annotation information based at least on the first set of evaluation parameters.

[0129] In some embodiments, the first determining module 430 is further configured to determine the projection region of the target instance in the target image; and to determine a first evaluation parameter based on a comparison between the projection region and a reference region indicated by the first segmentation result.

[0130] In some embodiments, the first determining module 430 is further configured to determine the distance from the target point set corresponding to the annotation information to the reference point set indicated by the second segmentation result based on a comparison between the annotation information and the second segmentation result; and to determine a second evaluation parameter based on the distance.

[0131] In some embodiments, the first determining module 430 is further configured to determine multiple point sets of the target instance at multiple time points based on annotation information; and to determine a second evaluation parameter based on the consistency degree and distance of the multiple point sets, wherein the second evaluation parameter is positively correlated with the consistency degree and negatively correlated with the distance.

[0132] In some embodiments, the second determining module 440 is further configured to determine a second set of evaluation parameters based on a comparison of prior information and annotation information associated with the target instance; and to determine the target evaluation of the annotation information based on the first set of evaluation parameters and the second set of evaluation parameters.

[0133] In some embodiments, the prior information indicates at least one of the following: the point cloud reflectance corresponding to the target instance; and the shape features corresponding to the target instance.

[0134] In some embodiments, the annotation information includes four-dimensional annotation information, and the second determining module 440 is further configured to construct a global map based on multi-frame point cloud data; construct a pose graph based on the four-dimensional annotation information; determine a third set of evaluation parameters based on the optimization results of the pose graph; and determine the target evaluation of the annotation information based on the first set of evaluation parameters and the third set of evaluation parameters.

[0135] In some embodiments, the second determining module 440 is further configured to provide multiple sets of evaluation parameters to the decision model, the multiple sets of evaluation parameters including the first set of evaluation parameters; and to obtain the target evaluation generated by the decision model.

[0136] Figure 5 A block diagram is shown illustrating a computing device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that... Figure 5 The computing device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The computing device 500 shown can be used to implement Figure 1 Electronic devices 110.

[0137] like Figure 5 As shown, computing device 500 is in the form of a general-purpose computing device. Components of computing device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 500.

[0138] Computing device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to computing device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 500.

[0139] The computing device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0140] The communication unit 540 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 500 can be implemented as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 500 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.

[0141] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 500 can also communicate as needed with one or more external devices (not shown) via communication unit 540. These external devices, such as storage devices, display devices, etc., can communicate with one or more devices that enable user interaction with computing device 500, or with any device that enables computing device 500 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interfaces (not shown).

[0142] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0143] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0144] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0145] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0146] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0147] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for processing annotation information, comprising: Obtain annotation information from point cloud data, wherein the annotation information indicates target instances in the point cloud data; Obtain multiple instance segmentation results from multiple reference models, the multiple instance segmentation results including a first segmentation result generated by an image segmentation model based on a target image associated with the point cloud data and a second segmentation result generated by a point cloud segmentation model based on the point cloud data; Based on the comparison between the annotation information and the segmentation results of the multiple instances, a first set of evaluation parameters is determined; as well as The target evaluation of the labeled information is determined based at least on the first set of evaluation parameters.

2. The method according to claim 1, wherein determining the first set of evaluation parameters based on the comparison between the annotation information and the segmentation results of the plurality of instances includes: Determine the projection region of the target instance in the target image; as well as A first evaluation parameter is determined based on the comparison between the projected region and the reference region indicated by the first segmentation result.

3. The method according to claim 1, wherein determining the first set of evaluation parameters based on the comparison between the annotation information and the segmentation results of the plurality of instances includes: Based on the comparison between the annotation information and the second segmentation result, the distance from the target point set corresponding to the annotation information to the reference point set indicated by the second segmentation result is determined; as well as Based on the distance, a second evaluation parameter is determined.

4. The method according to claim 3, wherein determining the second evaluation parameter based on the distance includes: Based on the annotation information, multiple point sets of the target instance at multiple times are determined; as well as Based on the degree of consistency of the plurality of point sets and the distance, a second evaluation parameter is determined, wherein the second evaluation parameter is positively correlated with the degree of consistency and negatively correlated with the distance.

5. The method according to claim 1, wherein determining the target evaluation of the annotation information based at least on the first set of evaluation parameters includes: A second set of evaluation parameters is determined based on a comparison of prior information associated with the target instance and the annotation information. as well as Based on the first set of evaluation parameters and the second set of evaluation parameters, the target evaluation of the labeled information is determined.

6. The method of claim 5, wherein the prior information indicates at least one of the following: The point cloud reflectance corresponding to the target instance; The shape features corresponding to the target instance.

7. The method according to claim 1, wherein the annotation information includes four-dimensional annotation information, and determining the target evaluation of the annotation information based at least on the first set of evaluation parameters includes: Construct a global map based on multi-frame point cloud data; Based on the aforementioned four-dimensional annotation information, a pose graph is constructed; as well as Based on the optimization results of the pose graph, a third set of evaluation parameters is determined; as well as Based on the first set of evaluation parameters and the third set of evaluation parameters, the target evaluation of the labeled information is determined.

8. The method according to claim 1, wherein determining the target evaluation of the annotation information based at least on the first set of evaluation parameters includes: The decision-making model is provided with multiple sets of evaluation parameters, including the first set of evaluation parameters. as well as Obtain the target evaluation generated by the decision model.

9. An apparatus for processing annotation information, comprising: The first acquisition module is configured to acquire annotation information of point cloud data, wherein the annotation information indicates target instances in the point cloud data; The second acquisition module is configured to acquire multiple instance segmentation results of multiple reference models, the multiple instance segmentation results including a first segmentation result generated by an image segmentation model based on a target image associated with the point cloud data and a second segmentation result generated by a point cloud segmentation model based on the point cloud data; The first determining module is configured to determine a first set of evaluation parameters based on a comparison between the annotation information and the segmentation results of the multiple instances; as well as The second determining module is configured to determine the target evaluation of the labeled information based at least on the first set of evaluation parameters.

10. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 8.

12. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 8.