Label sample data generation method and electronic device
By introducing multiple marker objects into the target scene and establishing a unified target coordinate system, the problem is transformed into a 3D solution and multi-view projection problem. This solves the problems of low efficiency, large error, and difficulty in expansion of multi-view key point annotation in existing technologies, and achieves efficient and controllable annotation result expansion and stability improvement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD
- Filing Date
- 2026-02-12
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, multi-view key point annotation relies heavily on manual interaction, resulting in low efficiency, serious mislabeling or deviation, and difficulty in expansion, which affects annotation accuracy and algorithm reliability.
By introducing multiple marker objects into the target scene, acquiring multiple frames of images, establishing a unified target coordinate system, and transforming it into a 3D solution and multi-view projection problem, annotation sample data is generated, enabling efficient expansion of key point annotation results across multiple views and frames.
It improves annotation efficiency and consistency, reduces reliance on external 3D reconstruction processes, enhances the controllability and fault tolerance of annotation quality, and ensures the stability and anti-interference of the annotation process.
Smart Images

Figure CN122115727A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data generation technology, and more specifically, to a method and electronic device for generating labeled sample data. Background Technology
[0002] With the rapid development of computer vision and 3D perception technologies, multi-view keypoint annotation plays an increasingly important role in applications such as robot navigation, action recognition, 3D reconstruction, and embodied intelligence training. Especially in static scenes, accurate spatial localization of keypoints at specific locations on target objects or in the environment is a core step in building high-quality multi-view datasets. Such annotation results are not only used for training supervised learning models but also provide fundamental support for cross-view feature matching, pose estimation, and spatial reasoning.
[0003] In existing technologies, multi-view keypoint annotation mainly relies on human-interactive 2D annotation tools, where operators manually click on the pixel coordinates of specified keypoints in each frame of the image to complete the annotation process. For the same physical keypoint, independent manual annotation needs to be repeated in images from different viewpoints, and the annotation task is regarded as a series of isolated 2D image processing operations.
[0004] However, this purely manual frame-by-frame annotation method is highly dependent on human resources and is prone to mislabeling or deviations, severely affecting annotation accuracy and the reliability of subsequent algorithms. Furthermore, it has low scalability. Summary of the Invention
[0005] The purpose of this application is to provide a method and electronic device for generating labeled sample data, in order to address the shortcomings of the prior art, thereby solving the problems that the prior art relies heavily on human input and is prone to mislabeling or deviation, which seriously affects the labeling accuracy and the reliability of subsequent algorithms.
[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows: In a first aspect, one embodiment of this application provides a method for generating labeled sample data, the method comprising: Acquire multiple frames of images for a target scene, wherein the target scene includes: a target object and multiple pre-set marker objects; The marker position information of each of the marked objects in each frame image is determined, and the initial spatial pose of each of the marked objects in each frame image is determined. The marker position information includes the pixel coordinates of the marked object in the image coordinate system of the image. The target coordinate system corresponding to the multi-frame image is determined based on the initial spatial pose of each marked object in each frame image; Based on the multi-frame images, the target coordinate system corresponding to the multi-frame images, the marker position information of each marker object in each frame image, and the initial spatial pose of each marker object in each frame image, the global spatial structure information corresponding to the multi-frame images is determined. The global spatial structure information includes the global position information of each marker object and the camera pose of each image. The global position information is used to indicate the position information of the marker object in the target coordinate system, and the camera pose is used to indicate the camera pose of each frame image in the target coordinate system. Based on the global spatial structure information corresponding to the multi-frame images and the annotation information of the target object in the multi-frame images, annotation sample data is generated.
[0007] Secondly, another embodiment of this application provides a labeled sample data generation apparatus, the apparatus comprising: The acquisition module is used to acquire multiple frames of images of a target scene, wherein the target scene includes a target object and multiple pre-set marker objects; The determination module is used to determine the marker position information of each of the marker objects in each frame image, and to determine the initial spatial pose of each of the marker objects in each frame image, wherein the marker position information includes the pixel coordinates of the marker object in the image coordinate system of the image; The determining module is used to determine the target coordinate system corresponding to the multi-frame images based on the initial spatial pose of each marked object in each frame image; The determining module is used to determine global spatial structure information corresponding to the multi-frame images based on the multi-frame images, the target coordinate system corresponding to the multi-frame images, the marker position information of each marker object in each frame image, and the initial spatial pose of each marker object in each frame image. The global spatial structure information includes the global position information of each marker object and the camera pose of each image. The global position information is used to indicate the position information of the marker object in the target coordinate system, and the camera pose is used to indicate the camera pose of each frame image in the target coordinate system. The generation module is used to generate labeled sample data based on the global spatial structure information corresponding to the multi-frame images and the annotation information of the target object in the multi-frame images.
[0008] Thirdly, another embodiment of this application provides an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of any of the methods described in the first aspect above.
[0009] Fourthly, another embodiment of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of any of the methods described in the first aspect above.
[0010] The beneficial effects of this application are: by introducing multiple marked objects into the target scene and acquiring multiple frames of images of the target scene, determining the marked position information of each marked object in each frame of images, and determining the initial spatial pose of each marked object in each frame of images, and based on the initial spatial pose of each marked object in each frame of images, determining the target coordinate system corresponding to the multiple frames of images, a unified and stable target coordinate system corresponding to the multiple frames of images can be established. Therefore, based on the multiple frames of images, the target coordinate system corresponding to the multiple frames of images, the marked position information of each marked object in each frame of images, and the initial spatial pose of each marked object in each frame of images, the global spatial structure information corresponding to the multiple frames of images can be determined, which can transform the originally mutually exclusive... The independent multi-view 2D keypoint annotation problem is transformed into a keypoint 3D solution and multi-view projection problem under unified geometric constraints. Based on the global spatial structure information corresponding to multiple frames of images and the annotation information of the target object in multiple frames of images, annotation sample data is generated. It can solve the 3D spatial position of keypoints through multi-view geometric constraints and automatically project the 3D spatial position of keypoints to other views or other frames, thereby achieving efficient expansion of keypoint annotation results across multiple views and frames. It can achieve a synergistic improvement in efficiency, multi-view consistency and annotation quality controllability of multi-view keypoint annotation without relying on complex external 3D reconstruction processes or dedicated calibration software.
[0011] Meanwhile, even if individual labeled objects are occluded, images are blurred, or detection is abnormal, the stability of the labeled sample data generation process can still be maintained through multi-label redundancy and multi-frame fusion mechanisms, ensuring the continuous and reliable operation of the labeling process and improving the fault tolerance and anti-interference capability of the labeled sample data generation process. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A flowchart illustrating a method for generating labeled sample data provided in an embodiment of this application; Figure 2 A schematic diagram illustrating the marker position information of a marker object in the annotation sample data generation method provided in this application embodiment; Figure 3 This is a flowchart illustrating the process of determining the initial spatial pose of each marked object in each frame image in the annotation sample data generation method provided in the embodiments of this application. Figure 4 This is a flowchart illustrating the process of determining the target coordinate system corresponding to multiple frames of images in the annotation sample data generation method provided in the embodiments of this application. Figure 5 This is a flowchart illustrating the process of determining the global spatial structure information corresponding to multiple frames of images in the annotation sample data generation method provided in this application embodiment. Figure 6 This is a schematic flowchart illustrating the process of obtaining at least one intermediate image in the method for generating labeled sample data provided in the embodiments of this application. Figure 7 This is a flowchart illustrating the process of determining the global spatial structure information corresponding to multiple frames of images in the annotation sample data generation method provided in this application embodiment. Figure 8 Another flowchart illustrating the method for generating labeled sample data provided in this application embodiment; Figure 9 This is a schematic flowchart illustrating the process of generating labeled sample data in the labeled sample data generation method provided in the embodiments of this application. Figure 10 A schematic diagram of the annotation information of the target object in the annotation sample data generation method provided in the embodiments of this application; Figure 11 This is a schematic diagram of the electronic device structure provided in an embodiment of this application. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0015] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0016] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0017] In existing technologies, multi-view keypoint annotation mainly relies on human-interactive 2D annotation tools, where operators manually click on the pixel coordinates of specified keypoints in each frame of the image to complete the annotation process. For the same physical keypoint, independent manual annotation needs to be repeated in images from different viewpoints, and the annotation task is regarded as a series of isolated 2D image processing operations.
[0018] However, this purely manual frame-by-frame annotation method is highly dependent on human resources. The annotation workload increases linearly with the number of viewpoints and frames, resulting in low overall efficiency and making it difficult to meet the needs of rapidly constructing large-scale multi-view datasets. Furthermore, annotation results from different viewpoints are prone to spatial inconsistencies, especially under occlusion, deformation, or low-resolution conditions, which can easily lead to mislabeling or deviations, seriously affecting annotation accuracy and the reliability of subsequent algorithms. Additionally, it suffers from limited scalability.
[0019] Based on the aforementioned problems, this application proposes a method for generating labeled sample data. By introducing multiple labeled objects into the target scene and acquiring multiple frames of images of the target scene, a unified and stable target coordinate system corresponding to the multiple frames is established. This transforms the originally independent multi-view 2D keypoint annotation problem into a keypoint 3D solution and multi-view projection problem under unified geometric constraints. Thus, the 3D spatial position of the keypoint can be solved through multi-view geometric constraints, and the 3D spatial position of the keypoint can be automatically projected to other views or other frames. This achieves efficient expansion of keypoint annotation results across multiple views and frames. Without relying on complex external 3D reconstruction processes or dedicated calibration software, this method achieves a synergistic improvement in efficiency, multi-view consistency, and controllable annotation quality of multi-view keypoint annotation.
[0020] It is understood that by executing the labeled sample data generation method provided in the embodiments of this application, labeled sample data including labeled information can be automatically generated, which can then be applied to model training, data analysis or other computer vision-related applications to improve the richness of sample data.
[0021] The following describes in detail the method for generating labeled sample data provided in the embodiments of this application, with reference to several examples.
[0022] Figure 1 This is a flowchart illustrating a method for generating labeled sample data provided in an embodiment of this application. (Refer to...) Figure 1 As shown, the subject executing this method can be any electronic device with processing capabilities, and the method includes: S101. Acquire multiple frames of images for the target scene.
[0023] Optionally, multiple frames of images pre-captured for the target scene can be acquired.
[0024] The target scene includes a target object and multiple pre-set marker objects. Each frame image includes a target object and at least one marker object. The camera pose is different for each frame image.
[0025] Optionally, the target scene may include multiple target objects, which can be any object in the target scene other than the marked object. For example, in a transportation scene, the target object can be any object to be transported in the transportation scene.
[0026] Specifically, the marked object refers to a pre-set coded planar marker in the target scene. This coded planar marker is a planar coded pattern with unique identification information and known geometric dimensions. Examples include ArUco Marker, AprilTag, or other types.
[0027] Multiple frames can be captured from different spatial perspectives of the target scene. For example, a video can be recorded of the target scene, and each frame of the video can be used as a separate image. Furthermore, during the recording process, the target object and all marked objects in the target scene remain stationary; only the camera's perspective changes. The camera can be a monocular, binocular, or multi-view camera.
[0028] S102. Determine the marker position information of each marker object in each frame image, and determine the initial spatial pose of each marker object in each frame image.
[0029] Optionally, after obtaining each frame image, each marked object can be detected from each frame image, and the marking position information of each marked object in each frame image can be determined.
[0030] The marker location information includes the pixel coordinates of the marker object in the image coordinate system.
[0031] Optionally, each marked object may also carry a corresponding identifier to distinguish each marked object.
[0032] For example, each frame image can be preprocessed by grayscale conversion, filtering, and adaptive thresholding, and then contour detection and quadrilateral extraction can be performed to obtain multiple optional quadrilateral contours. Each optional quadrilateral contour can be corrected to a front view by perspective transformation and identified by pre-configured encoding rules to obtain each marked object, and the mark position information of each marked object in each frame image can be extracted.
[0033] For example, Figure 2 This is a schematic diagram illustrating the marker position information of a marker object in the annotation sample data generation method provided in this application embodiment, with reference to... Figure 2 As shown, taking a single image frame as an example, multiple marked objects can be detected in that frame, and the location information of each marked object within that frame can be determined. Figure 2 The two-dimensional pixel coordinates of the four corner points of each yellow rectangle.
[0034] Optionally, after obtaining the marker position information of each marker object in each frame image, the initial spatial pose of each marker object in each frame image can be detected to obtain the initial spatial pose of each marker object in each frame image.
[0035] The initial spatial pose of each marked object in each frame refers to the position and orientation of each marked object in the camera coordinate system of each image.
[0036] In one example, the initial spatial pose of each marked object in each frame image can be determined based on the marked position information of each marked object in each frame image.
[0037] In another example, each frame of images can be input into a pre-trained spatial pose detection model to identify the initial spatial pose of each labeled object in each frame. This spatial pose detection model can be implemented based on a deep learning model.
[0038] S103. Determine the target coordinate system corresponding to multiple frames of images based on the initial spatial pose of each marked object in each frame of images.
[0039] The target coordinate system corresponding to multiple frames of images is used to indicate the unified three-dimensional spatial reference corresponding to the multiple frames of images.
[0040] Optionally, after obtaining the initial spatial pose of each marked object in each frame image, the target marked object in each marked object can be determined by the initial spatial pose of each marked object in each frame image, thereby determining the target coordinate system corresponding to multiple frames images based on the coordinate system of the target marked object.
[0041] Optionally, a virtual reference coordinate system can be constructed by combining and analyzing the initial spatial poses of each marked object in each frame of the image, and the virtual reference coordinate system can be used as the target coordinate system for multiple frames of the image.
[0042] In one example, the spatial distribution stability of each labeled object can be determined based on the initial spatial pose of each labeled image in each frame image. Based on the spatial distribution stability of each labeled object, the labeled object with the highest spatial distribution stability can be selected from among the labeled objects and used as the target labeled object.
[0043] In another example, the geometric constraint reliability of each labeled object can be determined based on the initial spatial pose of each labeled image in each frame image. Based on the geometric constraint reliability of each labeled object, the labeled object with the highest geometric constraint reliability can be selected from the labeled objects and used as the target labeled object.
[0044] In another example, the number of observations of each marked object can be determined based on the initial spatial pose of each marked image in each frame image, and the marked object with the highest number of observations can be selected from all marked objects as the target marked object.
[0045] S104. Based on the multi-frame images, the target coordinate system corresponding to the multi-frame images, the marker position information of each marker object in each frame image, and the initial spatial pose of each marker object in each frame image, determine the global spatial structure information corresponding to the multi-frame images.
[0046] Optionally, after obtaining the target coordinate system corresponding to multiple frames of images, the marker position information of each marker object in each frame of images, and the initial spatial pose of each marker object in each frame of images, the marker position information of each marker object in each frame of images and the initial spatial pose of each marker object in each frame of images can be transformed according to the multiple frames of images and the target coordinate system corresponding to the multiple frames of images to determine the global spatial structure information corresponding to the multiple frames of images.
[0047] The global spatial structure information includes the global position information of each marked object and the camera pose of each image. The global position information is used to indicate the position of the marked object in the target coordinate system, and the camera pose is used to indicate the camera pose of each frame image in the target coordinate system.
[0048] Specifically, global spatial structure information is the set of spatial topological relationships obtained by aligning all images and all labeled objects to the target coordinate system. Global position information for each labeled object refers to its three-dimensional coordinates in the target coordinate system. Camera pose for each image refers to the camera pose of each image in the target coordinate system, including the camera's position and orientation.
[0049] In one example, the marker position information of each marked object in each frame image can be fused based on multiple frames of images and the target coordinate system corresponding to the multiple frames of images to obtain the global position information of each marked object. Then, the initial spatial pose of each marked object in each frame image can be fused and solved based on multiple frames of images and the target coordinate system corresponding to the multiple frames of images to obtain the camera pose of each image.
[0050] S105. Generate labeled sample data based on the global spatial structure information corresponding to multiple frames of images and the annotation information of the target object in multiple frames of images.
[0051] Optionally, after obtaining the global spatial structure information corresponding to multiple frames of images, labeled sample data can be generated based on the global spatial structure information corresponding to multiple frames of images and the annotation information of the target object in the multiple frames of images.
[0052] The annotation information of the target object in the multi-frame images includes the annotated image frame, the annotated target object, the two-dimensional coordinates of the annotated target object, and the camera parameter information of the annotated image frame.
[0053] For example, in response to a user's annotation operation in at least two of the multi-frame images, annotation information of the target object in the multi-frame images can be generated, and a three-dimensional solution can be performed based on the annotation information of the target object in the multi-frame images and the global spatial structure information corresponding to the multi-frame images to obtain key point information. The key point information is then extended to each frame image to generate annotation sample data.
[0054] In this embodiment, by introducing multiple marked objects into the target scene and acquiring multiple frames of images of the target scene, the marking position information of each marked object in each frame of the image is determined, and the initial spatial pose of each marked object in each frame of the image is determined. Based on the initial spatial pose of each marked object in each frame of the image, the target coordinate system corresponding to the multiple frames of the image is determined, thus establishing a unified and stable target coordinate system corresponding to the multiple frames of the image. Therefore, based on the multiple frames of the image, the target coordinate system corresponding to the multiple frames of the image, the marking position information of each marked object in each frame of the image, and the initial spatial pose of each marked object in each frame of the image, the global spatial structure information corresponding to the multiple frames of the image can be determined, which can transform the originally independent... The multi-view 2D keypoint annotation problem is transformed into a keypoint 3D solution and multi-view projection problem under unified geometric constraints. Based on the global spatial structure information corresponding to multiple frames of images and the annotation information of the target object in multiple frames of images, annotation sample data is generated. It can solve the 3D spatial position of keypoints through multi-view geometric constraints and automatically project the 3D spatial position of keypoints to other views or other frames, thereby realizing the efficient expansion of keypoint annotation results across multiple views and multiple frames. It can achieve a synergistic improvement in efficiency, multi-view consistency and annotation quality controllability of multi-view keypoint annotation without relying on complex external 3D reconstruction processes or dedicated calibration software.
[0055] Meanwhile, even if individual labeled objects are occluded, images are blurred, or detection is abnormal, the stability of the labeled sample data generation process can still be maintained through multi-label redundancy and multi-frame fusion mechanisms, ensuring the continuous and reliable operation of the labeling process and improving the fault tolerance and anti-interference capability of the labeled sample data generation process.
[0056] In one possible implementation, Figure 3 This is a flowchart illustrating the process of determining the initial spatial pose of each marked object in each frame image within the annotation sample data generation method provided in this application embodiment. (Refer to...) Figure 3 As shown, determining the initial spatial pose of each marked object in each frame image in S102 includes: S301. Obtain the attribute information of each marked object and the camera intrinsic parameters corresponding to multiple frames of images.
[0057] Optionally, attribute information of each marked object and camera intrinsic parameters corresponding to multiple frames of images can be obtained.
[0058] The attribute information for each marker object includes its length and width, among other things. The camera intrinsic parameters corresponding to multiple frames refer to the intrinsic parameters of the camera that captured the multiple frames.
[0059] S302. Based on the attribute information of each marked object corresponding to the multi-frame images, the camera intrinsic parameters corresponding to the multi-frame images, and the mark position information of each marked object, determine the initial spatial pose of each marked object in each frame image.
[0060] In one example, the initial spatial pose of each marked object in each frame image can be calculated based on the attribute information of each marked object corresponding to multiple frames of images, the camera intrinsic parameters corresponding to multiple frames of images, and the marked position information of each marked object, using the Perspective-n-Point (PnP) algorithm. This can achieve high-precision and interpretable initial spatial pose estimation, while avoiding problems such as poor model generalization ability and complex deployment, without the need for training data or deep learning models.
[0061] In another example, the initial spatial pose of each labeled object in each frame image can be obtained by using a pre-trained deep learning model based on the attribute information of each labeled object corresponding to multiple frames of images, the camera intrinsic parameters corresponding to multiple frames of images, and the label position information of each labeled object.
[0062] By using the attribute information of each labeled object corresponding to multiple frames of images, the camera intrinsic parameters corresponding to multiple frames of images, and the label position information of each labeled object, the initial spatial pose of each labeled object in each frame of images can be determined. This enables lightweight, low-cost, and easy-to-deploy initial spatial pose estimation without the need for additional sensors or high-precision robotic arms to assist in motion, thus reducing reliance on external equipment or human intervention.
[0063] In one possible implementation, Figure 4 This is a flowchart illustrating the process of determining the target coordinate system corresponding to multiple frames of images in the annotation sample data generation method provided in this application embodiment, with reference to... Figure 4 As shown, in step S103 above, determining the target coordinate system corresponding to multiple frames of images based on the initial spatial pose of each marked object in each frame image includes: S401. Based on the initial spatial pose of each marked object in each frame image, determine the co-occurrence information of each marked object in multiple frames images.
[0064] Optionally, the co-occurrence information of each marked object in multiple frames can be determined based on the initial spatial pose of each marked object in each frame image.
[0065] Co-occurrence information is used to indicate the information of the co-occurrence of each marked object in each frame image, or the intersection information of each marked object in each frame image.
[0066] In one example, co-occurrence information can indicate that both the first and second marked objects appear in a frame of an image.
[0067] In another example, co-occurrence information can indicate that a first marked object and a second marked object have corner points in a frame of an image.
[0068] For example, the initial spatial pose of each marked object in each frame image can be statistically analyzed to obtain the co-occurrence information of each marked object in multiple frames image.
[0069] S402. Based on the co-occurrence information of each marked object in multiple frames of images, the target marked object is obtained by filtering from each marked object.
[0070] Optionally, the target marker object can be selected from the marker objects based on the co-occurrence information of each marker object in multiple frames of images.
[0071] The target marker is the marker that appears most frequently in the co-occurrence information. That is, the target marker is the marker that co-occurs most frequently with other markers.
[0072] S403. Determine the target coordinate system based on the coordinate system of the target marked object.
[0073] Optionally, the center point of the target object is taken as the origin, the plane where the target object is located is taken as the xy plane, and the direction perpendicular to the target object and upward is taken as the z direction, so as to obtain the target coordinate system corresponding to multiple frames of images.
[0074] By determining the co-occurrence information of each marked object in multiple frames of images, and using this information to identify the target marked object, the target coordinate system can be determined based on the coordinate system of the target marked object. This improves the stability and robustness of the target coordinate system and enhances the consistency of cross-view geometric constraints. Furthermore, it enables automated selection of the coordinate origin, making it suitable for diverse deployment scenarios and improving scalability and ease of use.
[0075] In one possible implementation, Figure 5 This is a flowchart illustrating the process of determining the global spatial structure information corresponding to multiple frames of images in the annotation sample data generation method provided in this application embodiment, with reference to... Figure 5 As shown, in S104 above, the global spatial structure information corresponding to the multiple frames of images is determined based on the target coordinate system corresponding to the multiple frames of images, the marker position information of each marker object in each frame of images, and the initial spatial pose of each marker object in each frame of images, including: S501. Based on multiple frames of images, the target coordinate system corresponding to the multiple frames of images, the marker position information of each marker object in each frame of images, and the initial spatial pose of each marker object in each frame of images, perform multi-frame fusion position solution to determine the intermediate spatial pose of each marker object.
[0076] Optionally, multi-frame fusion position solution can be performed based on multiple frames of images, the target coordinate system corresponding to the multiple frames of images, the marker position information of each marker object in each frame of images, and the initial spatial pose of each marker object in each frame of images, to preliminarily estimate the three-dimensional position of each marker object in the target coordinate system and obtain the intermediate spatial pose of each marker object.
[0077] The intermediate spatial pose is used to indicate the spatial pose of each marked object in the target coordinate system. Specifically, the intermediate spatial pose is used to indicate the initial spatial pose of each marked object in the target coordinate system.
[0078] For example, in each frame of the image, candidate position information of each marked object in the target coordinate system can be deduced based on the marked position information and the initial spatial pose of each marked object. Then, the intermediate spatial pose of each marked object can be determined based on the candidate position information in the target coordinate system. For instance, the intermediate spatial pose of each marked object can be determined by filtering the candidate position information in the target coordinate system.
[0079] S502. Based on multiple frames of images, the target coordinate system corresponding to the multiple frames of images, and the marker position information of each marker object in each frame of images, solve the camera pose of multiple marker objects, determine the intermediate camera pose of each frame of images, and filter the multiple frames of images based on the intermediate camera pose of each frame of images to obtain at least one intermediate image.
[0080] Optionally, the camera pose of each frame image can be solved by multi-marker object camera pose based on multiple frames of images, the target coordinate system corresponding to the multiple frames of images, and the mark position information of each marked object in each frame image, so as to obtain the intermediate camera pose of each frame image.
[0081] The intermediate camera pose indicates the camera's pose in the target coordinate system within the current frame image. Specifically, the intermediate camera pose indicates the camera's initial pose in the target coordinate system within the current frame image.
[0082] Optionally, after obtaining the intermediate camera pose of each frame, the quality of each frame can be evaluated and filtered using the intermediate camera pose of each frame to obtain at least one intermediate image.
[0083] For example, the intermediate camera pose of each frame image can be verified, and the images corresponding to intermediate camera poses at abnormal angles can be removed according to preset verification rules to obtain at least one intermediate image.
[0084] S503. Perform joint optimization based on each intermediate image, the intermediate camera pose of each intermediate image, and the intermediate spatial pose of each marked object to determine the global spatial structure information corresponding to multiple frames of images.
[0085] Optionally, after obtaining each intermediate image, global consistency joint optimization can be performed based on each intermediate image, the intermediate camera pose of each intermediate image, and the intermediate spatial pose of each labeled object to determine the global spatial structure information corresponding to multiple frames of images, so that the overall consistency reaches the optimal level, thereby improving the overall accuracy and consistency of the obtained global spatial structure information.
[0086] For example, a joint optimization function can be constructed based on the intermediate camera pose of each intermediate image and the intermediate spatial pose of each labeled object, and then iteratively solved using an optimization algorithm to obtain the global spatial structure information corresponding to multiple frames of images.
[0087] By determining the intermediate spatial poses of each labeled object and the intermediate camera poses of each frame, and then filtering multiple frames based on these intermediate camera poses to obtain at least one intermediate image, joint optimization is performed to determine the global spatial structure information corresponding to multiple frames. This avoids the convergence difficulties caused by performing complex optimizations directly on the original noisy data, effectively suppresses the propagation of single-frame errors, ensures computational efficiency, and significantly improves the accuracy of 3D reconstruction, ensuring that the final output global spatial structure information has high geometric consistency. It also significantly reduces labor costs, supports the rapid generation of large-scale datasets, and is suitable for industrial-grade AI training needs.
[0088] In one possible implementation, step S501 above involves multi-frame fusion position solving based on multiple frames of images, the target coordinate system corresponding to the multiple frames of images, the marker position information of each marker object in each frame of images, and the initial spatial pose of each marker object in each frame of images, to determine the intermediate spatial pose of each marker object, including: Traverse multiple frames of images. For the current frame image, determine the spatial information of each marked object in the current frame image based on the marked position information, target coordinate system, and initial spatial pose of each marked object in the current frame image. After traversal, fuse the spatial information of each marked object in each frame image to obtain the intermediate spatial pose of each marked object.
[0089] Optionally, multiple frames of images can be traversed. For the current frame image that has been traversed, all marked objects in the current frame image are traversed. For the current marked object that has been traversed, coordinate system transformation is performed based on the marked position information of the current marked object in the current frame, the initial spatial pose of the current marked object in the current frame image, and the target coordinate system. The spatial information of the current marked object in the current frame image is calculated. After all marked objects have been traversed, the spatial information of all marked objects in the current frame image is obtained.
[0090] Optionally, after all images have been traversed, spatial information of all marked objects in all images can be obtained.
[0091] Spatial information is used to indicate the spatial transformation relationship of the marked object relative to the target coordinate system, such as the angle between the plane where the marked object is located and the target coordinate system, and the position coordinates of the four vertices of the marked object in the target coordinate system.
[0092] Optionally, the spatial information of each labeled object in each frame image is fused over multiple frames to obtain the intermediate spatial pose of each labeled object. The fusion process includes weighted fusion, Random Sample Consensus Algorithm (RANSAC), or other fusion methods.
[0093] By determining the spatial information of each labeled object in each frame of the image frame and fusing it globally to obtain the intermediate spatial pose of each labeled object, the estimation accuracy of the intermediate spatial pose of each labeled object can be improved by utilizing multi-frame observations, significantly reducing random errors in single-frame estimation. It can also improve the robustness of the intermediate spatial pose determination process.
[0094] In one possible implementation, step S502 above involves solving the camera pose of multiple marked objects based on multiple frames of images, the target coordinate system corresponding to the multiple frames of images, and the marked position information of each marked object in each frame of images, to determine the intermediate camera pose of each frame of images, including: Based on the marker position information of all marked objects in the current frame image, the target coordinate system, and the preset camera pose solving algorithm, the intermediate camera pose of the current frame image is obtained.
[0095] Optionally, a globally consistent solution can be performed based on the marker position information of all marked objects in the current frame image, the target coordinate system, and a preset camera pose solving algorithm to obtain the intermediate camera pose of the current frame image. This increases the number of feature points involved in the calculation, providing stronger geometric constraints and significantly improving the pose solving accuracy. Furthermore, it enhances robustness and avoids over-reliance on a single marked object.
[0096] Among them, the camera pose solving algorithm can be the PnP algorithm.
[0097] The intermediate camera pose specifically refers to the extrinsic parameters of the camera in the target coordinate system within that frame of the image.
[0098] For example, taking the current frame image as an example, the intermediate camera pose of the current frame image can be calculated based on the PnP algorithm using the marker position information of all marked objects in the current frame image as joint constraints.
[0099] In one possible implementation, Figure 6 This is a schematic flowchart illustrating the process of obtaining at least one intermediate image in the method for generating labeled sample data provided in this application embodiment, with reference to... Figure 6 As shown, in step S502 above, multiple frames of images are filtered based on the intermediate camera pose of each frame to obtain at least one intermediate frame, including: S601. Based on the intermediate camera pose of the current frame image, reproject each marked object in the current frame image and determine the reprojection error.
[0100] Optionally, taking the current frame image as an example, each marked object in the current frame image can be reprojected according to the intermediate camera pose of the current frame image to obtain the reprojection position of each marked object in the current frame image. Based on the reprojection position of each marked object in the current frame image and the marked position information of each marked object in the current frame image, the reprojection error of each marked object in the current frame image can be calculated.
[0101] S602. Determine the number of corner points corresponding to each marked object in the current frame image.
[0102] Optionally, the number of in-corner points corresponding to each marked object in the current frame image can be counted to obtain the total number of in-corner points corresponding to each marked object in the current frame image. Here, the number of in-corner points refers to the total number of in-corner points for each marked object in the current frame image.
[0103] S603. Determine the validity result of the current frame image based on the number of corner points and reprojection error of each marked object in the current frame image.
[0104] Optionally, the validity result of the current frame image can be calculated based on the number of corner points and reprojection error of each marked object in the current frame image.
[0105] In one example, the number of inliers at the corners of each marked object in the current frame image and the reprojection error can be judged according to preset judgment conditions. If the number of inliers at the corners of each marked object in the current frame image is greater than the preset threshold for the number of inliers and the reprojection error is less than the preset error threshold, then the validity result corresponding to the current frame image can be determined as valid.
[0106] In another example, the number of corner points and reprojection error corresponding to each marked object in the current frame image can be input into a pre-built validity determination model to calculate the validity value corresponding to the current frame image, and the validity value can be used as the validity result.
[0107] The validity result is used to indicate the reliability of the current frame image.
[0108] S604. Based on the validity result corresponding to the current frame image, determine whether to use the current frame image as an intermediate frame image. If so, use the current frame image as an intermediate frame image.
[0109] Optionally, the validity result corresponding to the current frame image can be judged. If the validity result corresponding to the current frame image is valid, or the validity result corresponding to the current frame image is greater than the preset validity threshold, then the current frame image is determined to be an intermediate frame image.
[0110] Optionally, if the validity result corresponding to the current frame image is invalid, or the validity result corresponding to the current frame image is less than or equal to a preset validity threshold, then it is determined that the current frame image will not be used as an intermediate frame image.
[0111] By using the intermediate camera pose of the current frame image, each labeled object in the current frame image is reprojected, and the reprojection error is determined. Based on the number of corner points and inliers corresponding to each labeled object in the current frame image and the reprojection error, the validity result of the current frame image is determined. Thus, by using the validity result of the current frame image, each frame image is filtered, and images with blurriness, occlusion, or severe false detection can be excluded. This achieves automatic image quality assessment without manual intervention, thereby improving the standardization and repeatability of the entire annotation process.
[0112] In one possible implementation, Figure 7 This is a flowchart illustrating the process of determining the global spatial structure information corresponding to multiple frames of images in the annotation sample data generation method provided in this application embodiment, with reference to... Figure 7 As shown, in S503 above, joint optimization is performed based on each intermediate image, the intermediate camera pose of each intermediate image, and the intermediate spatial pose of each marked object to determine the global spatial structure information corresponding to multiple frames of images, including: S701. Using the intermediate camera pose of each intermediate image and the intermediate spatial pose of each marked object as optimization variables, determine the minimum reprojection error of each marked object in each intermediate image through a preset joint optimization algorithm.
[0113] Optionally, the intermediate camera pose of each intermediate image and the intermediate spatial pose of each marked object are used as optimization variables. Based on each intermediate image, the intermediate camera pose of each intermediate image and the intermediate spatial pose of each marked object are optimized through a preset joint optimization algorithm until the overall reprojection error of each marked object in multiple frames of images is minimized.
[0114] The minimum reprojection error refers to the minimum overall reprojection error of each marked object across multiple frames of images. Joint optimization algorithms include: gradient backpropagation, nonlinear least squares, or other numerical optimization methods.
[0115] S702. Based on the target camera pose of each intermediate image when the minimum reprojection error is obtained and the target spatial pose of each marked object, determine the global spatial structure information corresponding to the multi-frame images.
[0116] Optionally, the target camera pose of each intermediate image at which the minimum reprojection error is obtained is taken as the camera pose of each image, and the target spatial pose of each marked object is taken as the global position information of each marked object to obtain global spatial structure information.
[0117] By treating the camera poses of all intermediate images and the spatial poses of all labeled objects as unified optimization variables and simultaneously optimizing them under a joint objective function, global consistency optimization under multi-view geometric constraints can be achieved, eliminating the overall drift caused by the accumulation of local errors. Simultaneously, it can significantly improve the reconstruction accuracy of 3D spatial structures, resulting in higher geometric fidelity of the obtained global spatial structure information.
[0118] In one possible implementation, after obtaining the global spatial information corresponding to multiple frames of images, further filtering can be performed on the intermediate images of the multiple frames. Figure 8 This is another flowchart illustrating the method for generating labeled sample data provided in the embodiments of this application, referred to... Figure 8 As shown, the above method also includes: S801. Based on the target camera pose of each intermediate image and the target spatial pose of each marked object, project each marked object into each intermediate image to obtain the reference position information of each marked object.
[0119] Optionally, taking an intermediate frame as an example, each marked object can be reprojected into the intermediate frame based on the target camera pose and the target spatial pose of each marked object in the intermediate frame, so as to obtain the reference position information of each marked object in the intermediate frame.
[0120] S802. Based on the reference position information and the mark position information of each marked object, filter each intermediate image to obtain at least one target image.
[0121] Optionally, continuing to take a frame of intermediate image as an example, the position offset of each marked object can be calculated based on the reference position information of each marked object in the frame of intermediate image and the marked position information of each marked object in the frame image. The position offset of each marked object is then compared with a preset position offset threshold. If the position offset of each marked object is greater than or equal to the preset position offset threshold, it is determined that the frame of intermediate image has significant inconsistencies under the unified geometric constraints, and the frame of intermediate image is then removed.
[0122] By obtaining the reference position information of each labeled object and filtering the intermediate images based on the reference position information and the label position information of each labeled object, at least one target image is obtained. This can identify and remove image frames that seem reasonable in the initial processing stage but expose geometric inconsistencies after global optimization (such as image frames with false detections, occlusion, blurring or significant noise). This ensures that the final image used to generate labeled sample data is of high quality and conforms to unified three-dimensional geometric constraints, improves the spatial consistency and robustness of multi-view annotation results, and avoids the introduction of systematic bias due to individual low-quality frames.
[0123] In one possible implementation, Figure 9 This is a flowchart illustrating the process of generating labeled sample data in the labeled sample data generation method provided in this application embodiment, with reference to... Figure 9 As shown, in step S105 above, based on the global spatial structure information corresponding to multiple frames of images, labeled sample data is generated, including: S901. Based on the annotation information of the target object, the global spatial structure information, and the target coordinate system, the three-dimensional spatial coordinates of each annotation point are obtained.
[0124] Optionally, based on the annotation information of the target object, the global spatial structure information, and the target coordinate system, a multi-view projection relationship is constructed, and the three-dimensional spatial coordinates of each annotation point are solved.
[0125] In one example, taking two images as the labeled image frames, the three-dimensional spatial coordinates of each labeled point can be calculated using a dual-view triangulation method, global spatial structure information, and the target coordinate system.
[0126] In another example, taking three frames of labeled images as an example, the three-dimensional spatial coordinates of each labeled point can be calculated by fusing multi-view triangulation, global spatial structure information and target coordinate system.
[0127] For example, Figure 10 This is a schematic diagram illustrating the annotation information of the target object in the annotation sample data generation method provided in this application embodiment, with reference to... Figure 10As shown, the target object can be Figure 10 The green border in the image indicates the object, and the annotation information can include the object's four annotation points.
[0128] S902. Based on the three-dimensional spatial coordinates of each annotation point, project each annotation point onto other target images in the multi-frame image to obtain annotation sample data.
[0129] Optionally, after obtaining the three-dimensional spatial coordinates of each annotation point, the annotation points can be projected onto other target images in multiple frames based on global spatial structure information to generate corresponding two-dimensional key point candidate positions, thereby obtaining annotation sample data.
[0130] By utilizing the annotation information of each target object, global spatial structure information, and target coordinate system, the 3D spatial coordinates of each annotation point are obtained. Based on these coordinates, each annotation point is projected onto other target images in multiple frames to obtain annotation sample data. This enables automatic propagation of cross-view annotations, significantly improving annotation efficiency and reducing labor costs and time expenditure. Simultaneously, it ensures spatial consistency across multiple viewpoints, eliminating human error.
[0131] Based on the same inventive concept, this application also provides a labeled sample data generation device corresponding to the labeled sample data generation method. Since the principle of the device in this application is similar to the labeled sample data generation method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0132] This application provides an embodiment of a labeled sample data generation device, comprising: an acquisition module, a determination module, and a generation module; The acquisition module is used to acquire multiple frames of images of a target scene, which includes the target object and multiple pre-set marker objects. The determination module is used to determine the marker position information of each marker object in each frame image and to determine the initial spatial pose of each marker object in each frame image. The marker position information includes the pixel coordinates of the marker object in the image coordinate system of the image. The determination module is used to determine the target coordinate system corresponding to multiple frames of images based on the initial spatial pose of each marked object in each frame of images; The determination module is used to determine the global spatial structure information corresponding to the multi-frame images based on the multi-frame images, the target coordinate system corresponding to the multi-frame images, the marker position information of each marker object in each frame image, and the initial spatial pose of each marker object in each frame image. The global spatial structure information includes the global position information of each marker object and the camera pose of each image. The global position information is used to indicate the position information of the marker object in the target coordinate system, and the camera pose is used to indicate the camera pose of each frame image in the target coordinate system. The generation module is used to generate labeled sample data based on the global spatial structure information corresponding to multiple frames of images and the annotation information of the target object in the multiple frames of images.
[0133] Optionally, a module is defined, specifically for: Obtain the attribute information of each marked object and the camera intrinsic parameters corresponding to multiple frames of images; Based on the attribute information of each marked object corresponding to multiple frames of images, the camera intrinsic parameters corresponding to multiple frames of images, and the marked position information of each marked object, the initial spatial pose of each marked object in each frame of images is determined.
[0134] Optionally, a module is defined, specifically for: Based on the initial spatial pose of each labeled object in each frame image, determine the co-occurrence information of each labeled object in multiple frames images; Based on the co-occurrence information of each labeled object in multiple frames of images, the target labeled object is obtained by filtering from each labeled object; Determine the target coordinate system based on the coordinate system of the target marked object.
[0135] Optionally, a module is defined, specifically for: Based on multiple frames of images, the target coordinate system corresponding to the multiple frames of images, the marker position information of each marker object in each frame of images, and the initial spatial pose of each marker object in each frame of images, the multi-frame fusion position solution is performed to determine the intermediate spatial pose of each marker object. The intermediate spatial pose is used to indicate the spatial pose of each marker object in the target coordinate system. The camera pose of multiple marked objects is solved based on multiple frames of images, the target coordinate system corresponding to the multiple frames of images, and the mark position information of each marked object in each frame of images. The intermediate camera pose of each frame of images is determined, and the multiple frames of images are filtered based on the intermediate camera pose of each frame of images to obtain at least one intermediate image. The intermediate camera pose is used to indicate the pose of the camera in the target coordinate system in the current frame of images. By jointly optimizing each intermediate image, the intermediate camera pose of each intermediate image, and the intermediate spatial pose of each labeled object, the global spatial structure information corresponding to multiple frames of images is determined.
[0136] Optionally, a module is defined, specifically for: Traverse multiple frames of images. For the current frame image, determine the spatial information of each marked object in the current frame image based on the marked position information, target coordinate system, and initial spatial pose of each marked object in the current frame image. After the traversal is completed, the spatial information of each marked object in each frame image is fused to obtain the intermediate spatial pose of each marked object.
[0137] Optionally, a module is defined, specifically for: Based on the marker position information of all marked objects in the current frame image, the target coordinate system, and the preset camera pose solving algorithm, the intermediate camera pose of the current frame image is obtained.
[0138] Optionally, a module is defined, specifically for: Based on the intermediate camera pose of the current frame image, reproject each marked object in the current frame image and determine the reprojection error; Determine the number of corner points corresponding to each marked object in the current frame image; The validity result of the current frame image is determined based on the number of corner points and reprojection error of each marked object in the current frame image. Based on the validity result corresponding to the current frame image, determine whether to use the current frame image as an intermediate frame image. If so, use the current frame image as an intermediate frame image.
[0139] Optionally, a module is defined, specifically for: Using the intermediate camera pose of each intermediate image and the intermediate spatial pose of each labeled object as optimization variables, the minimum reprojection error of each labeled object in each intermediate image is determined by a preset joint optimization algorithm based on each intermediate image. Based on the target camera pose of each intermediate image when the minimum reprojection error is obtained, and the target spatial pose of each marked object, the global spatial structure information corresponding to multiple frames of images is determined.
[0140] Optionally, the determining module is also used for: Based on the target camera pose of each intermediate image and the target spatial pose of each marked object, each marked object is projected into each intermediate image to obtain the reference position information of each marked object; Based on the reference position information and the marker position information of each marker object, each intermediate image is filtered to obtain at least one target image.
[0141] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.
[0142] This application also provides an electronic device, such as... Figure 11 As shown, Figure 11 The schematic diagram of the electronic device structure provided in the embodiments of this application includes: a processor 1101 and a memory 1102, and optionally, a bus 1103. The memory 1102 stores machine-readable instructions executable by the processor 1101. When the electronic device is running, the processor 1101 and the memory 1102 communicate through the bus 1103, and the processor 1101 executes the machine-readable instructions to perform the steps of the above-described annotation sample data generation method.
[0143] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described method for generating labeled sample data.
[0144] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.
[0145] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0146] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for generating labeled sample data, characterized in that, include: Acquire multiple frames of images for a target scene, wherein the target scene includes: a target object and multiple pre-set marker objects; The marker position information of each of the marked objects in each frame image is determined, and the initial spatial pose of each of the marked objects in each frame image is determined. The marker position information includes the pixel coordinates of the marked object in the image coordinate system of the image. The target coordinate system corresponding to the multi-frame image is determined based on the initial spatial pose of each marked object in each frame image; Based on the multi-frame images, the target coordinate system corresponding to the multi-frame images, the marker position information of each marker object in each frame image, and the initial spatial pose of each marker object in each frame image, the global spatial structure information corresponding to the multi-frame images is determined. The global spatial structure information includes the global position information of each marker object and the camera pose of each image. The global position information is used to indicate the position information of the marker object in the target coordinate system, and the camera pose is used to indicate the camera pose of each frame image in the target coordinate system. Based on the global spatial structure information corresponding to the multi-frame images and the annotation information of the target object in the multi-frame images, annotation sample data is generated.
2. The method for generating labeled sample data according to claim 1, characterized in that, Determining the initial spatial pose of each of the marked objects in each frame image includes: Obtain the attribute information of each of the marked objects and the camera intrinsic parameters corresponding to the multi-frame images; Based on the attribute information of each marked object corresponding to the multi-frame images, the camera intrinsic parameters corresponding to the multi-frame images, and the mark position information of each marked object, the initial spatial pose of each marked object in each frame image is determined.
3. The method for generating labeled sample data according to claim 1, characterized in that, The step of determining the target coordinate system corresponding to the multi-frame images based on the initial spatial pose of each marked object in each frame image includes: Based on the initial spatial pose of each marked object in each frame image, determine the co-occurrence information of each marked object in the multi-frame image; Based on the co-occurrence information of each marked object in the multi-frame images, the target marked object is obtained by filtering from each marked object; The target coordinate system is determined based on the coordinate system of the target marked object.
4. The method for generating labeled sample data according to claim 1, characterized in that, The step of determining the global spatial structure information corresponding to the multi-frame images based on the multi-frame images, the target coordinate system corresponding to the multi-frame images, the marker position information of each marker object in each frame image, and the initial spatial pose of each marker object in each frame image includes: Based on the multi-frame images, the target coordinate system corresponding to the multi-frame images, the marking position information of each marked object in each frame image, and the initial spatial pose of each marked object in each frame image, the multi-frame fusion position solution is performed to determine the intermediate spatial pose of each marked object. The intermediate spatial pose is used to indicate the spatial pose of each marked object in the target coordinate system. The camera pose of multiple marked objects is solved based on the multiple frames of images, the target coordinate system corresponding to the multiple frames of images, and the mark position information of each marked object in each frame of images. The intermediate camera pose of each frame of images is determined, and the multiple frames of images are filtered based on the intermediate camera pose of each frame of images to obtain at least one intermediate image. The intermediate camera pose is used to indicate the pose of the camera in the target coordinate system in the current frame of images. Based on the intermediate images, the intermediate camera poses of the intermediate images, and the intermediate spatial poses of the marked objects, joint optimization is performed to determine the global spatial structure information corresponding to the multi-frame images.
5. The method for generating labeled sample data according to claim 4, characterized in that, The step of performing multi-frame fusion position calculation based on the multi-frame images, the target coordinate system corresponding to the multi-frame images, the marker position information of each marker object in each frame image, and the initial spatial pose of each marker object in each frame image, to determine the intermediate spatial pose of each marker object, includes: Traverse the multiple frames of images. For the current frame image, determine the spatial information of each marked object in the current frame image based on the marked position information of each marked object in the current frame image, the target coordinate system, and the initial spatial pose of each marked object in the current frame image. After the traversal is completed, the spatial information of each marked object in each frame image is fused to obtain the intermediate spatial pose of each marked object.
6. The method for generating labeled sample data according to claim 4, characterized in that, The step of solving the camera pose of multiple marked objects based on the multiple frames of images, the target coordinate system corresponding to the multiple frames of images, and the marked position information of each marked object in each frame of images, and determining the intermediate camera pose of each frame of images, includes: The intermediate camera pose of the current frame image is obtained by using the marker position information of all marked objects in the current frame image, the target coordinate system, and the preset camera pose solving algorithm.
7. The method for generating labeled sample data according to claim 4, characterized in that, The step of filtering the multiple frames of images based on the intermediate camera pose of each frame to obtain at least one intermediate frame includes: Based on the intermediate camera pose of the current frame image, each marked object in the current frame image is reprojected, and the reprojection error is determined. Determine the number of corner points corresponding to each marked object in the current frame image; The validity result of the current frame image is determined based on the number of corner points and reprojection error of each marked object in the current frame image. Based on the validity result corresponding to the current frame image, determine whether to use the current frame image as an intermediate frame image. If so, use the current frame image as an intermediate frame image.
8. The method for generating labeled sample data according to claim 4, characterized in that, The step of jointly optimizing the intermediate images, the intermediate camera poses of the intermediate images, and the intermediate spatial poses of the marked objects to determine the global spatial structure information corresponding to the multi-frame images includes: Using the intermediate camera pose of each intermediate image and the intermediate spatial pose of each labeled object as optimization variables, the minimum reprojection error of each labeled object in each intermediate image is determined by a preset joint optimization algorithm based on each intermediate image. Based on the target camera pose of each intermediate image at the minimum reprojection error and the target spatial pose of each marked object, the global spatial structure information corresponding to the multi-frame images is determined.
9. The method for generating labeled sample data according to claim 8, characterized in that, The method further includes: Based on the target camera pose of each intermediate image and the target spatial pose of each marker object, each marker object is projected into each intermediate image to obtain the reference position information of each marker object; Based on the reference position information and the mark position information of each of the marked objects, the intermediate images are filtered to obtain at least one target image.
10. An electronic device, characterized in that, include: The processor and memory, the memory storing machine-readable instructions executable by the processor, which, when the electronic device is running, are executed by the processor to perform the steps of the labeled sample data generation method as described in any one of claims 1 to 9.