Dataset collection method, training method, image recognition method, device, medium and product

By generating and synthesizing scene image sets and combining them with simulation acquisition methods, the problems of high dataset labeling costs and insufficient accuracy of synthesized data in supervised training are solved. This provides diverse and highly generalizable target datasets, thereby improving the model training effect.

CN122454571APending Publication Date: 2026-07-24BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2025-01-22
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing supervised training methods require a large amount of manually labeled datasets, which is costly and makes it difficult to guarantee the quality of the labels. Furthermore, the accuracy and realism of the synthetic data are insufficient, making it impossible to take into account the generalization ability of the model.

Method used

By acquiring a first set of images with semantic labels, a scene image set is generated and a target dataset is synthesized. A rich and diverse target dataset is generated using simulation acquisition methods, including outdoor and indoor scene images. Skeletal system objects are added, and simulation shooting is performed to determine the target semantic labels.

Benefits of technology

It enables the low-cost and efficient generation of diverse and highly generalizable target datasets, improves the richness and applicability of model training data, reduces the cost of manual annotation, and enhances the realism and interactivity of the dataset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122454571A_ABST
    Figure CN122454571A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a dataset collection method, a training method, an image recognition method, an apparatus, a medium and a product. The dataset collection method comprises: obtaining a first image set, each image in the first image set corresponding to a three-dimensional (3D) object and having a semantic label; for each image in the first image set, generating a scene image set based on the semantic label, synthesizing each image with each scene image in the scene image set to obtain a scene image set of each image, and obtaining all scene image sets of the images in the first image set as a second image set; and performing simulation collection based on the second image set to obtain a target dataset. Through the present disclosure, the richness and fidelity of the target dataset can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence, and in particular to dataset acquisition methods, training methods, image recognition methods, devices, media and products. Background Technology

[0002] In the fields of artificial intelligence and computer vision, supervised training plays a crucial role, often significantly improving the performance of AI models in specific task scenarios, including but not limited to depth estimation, object detection and segmentation, motion detection, and human pose estimation. However, supervised training requires large datasets, and the cost of manually labeling and building these datasets is high. Furthermore, it is often designed only for a single task, failing to consider model generalization and making it difficult to guarantee label quality. Summary of the Invention

[0003] To overcome the problems existing in related technologies, this disclosure provides a dataset acquisition method, training method, image recognition method, device, medium and product.

[0004] According to a first aspect of the present disclosure, a dataset acquisition method is provided, comprising: acquiring a first image set, wherein each image in the first image set corresponds to a three-dimensional 3D object and has a corresponding semantic label; for each image in the first image set, generating a scene image set based on the semantic label, and compositing each image with each scene image in the scene image set to obtain a scene image set for each image, and using the complete scene image set of each image in the first image set as a second image set; performing simulation acquisition based on the second image set to obtain a target dataset, wherein the target dataset includes a third image set and labeled data of target objects in the third image set, wherein the third image set is obtained based on simulation rendering of the second image set, and the target objects are determined based on target semantic labels determined in the semantic labels.

[0005] In one embodiment, the scene image set includes an outdoor scene image set; the step of generating the scene image set based on the semantic tag and compositing each image with each scene image in the scene image set to obtain a scene image set for each image includes: obtaining outdoor scene parameters associated with the semantic tag; randomly generating outdoor scene component resources based on the outdoor scene parameters; obtaining an outdoor scene image set based on the outdoor scene parameters and the outdoor scene component resources; and compositing each image with each scene image in the outdoor scene image set to obtain an outdoor scene image set for each image.

[0006] In one embodiment, the scene image set includes an indoor scene image set; the step of generating a scene image set based on the semantic tags and compositing each image with each scene image in the scene image set to obtain a scene image set for each image includes: obtaining indoor scene parameters, the indoor scene parameters including parameters for performing Venn diagram partitioning; performing Venn diagram partitioning on the regions of the indoor scene based on the indoor scene parameters to obtain multiple regions; matching indoor scene component resources in the multiple regions, the indoor scene component resources being associated with the semantic tags; obtaining an indoor scene image set based on the indoor scene parameters and the indoor scene component resources; and compositing each image with each scene image in the indoor scene image set to obtain an indoor scene image set for each image.

[0007] In one implementation, the method further includes adding a skeletal system object to the images of the second image set.

[0008] In one embodiment, adding a skeletal system object to the images in the second image set includes: creating an object template corresponding to the skeletal system object, wherein the object template is a preset character model; randomly setting appearance features for the object template and obtaining the skeletal system object; obtaining the target position of the 3D object in the scene image set; determining the target point of each bone in the skeletal system object based on the target position, wherein the target point is used to determine the position of the bone end of the skeletal system object; calculating the joint angle required to reach the position of the bone end based on the position of the bone end, and adjusting the bones of the skeletal system object according to the joint angle to obtain an adjusted skeletal system object; and adding the adjusted skeletal system object to the target object in the second image set.

[0009] In one implementation, the step of simulating acquisition based on the second image set to obtain a target dataset includes: acquiring preset camera parameters and determining the shooting pose information corresponding to each image in the second image set based on each image in the second image set; performing simulated shooting based on the camera parameters and the shooting pose information to obtain a third dataset after simulated shooting; determining target semantic labels in the semantic labels of the third dataset; and generating labeled data of target objects matching the target semantic labels in the third dataset based on the target semantic labels to obtain the target dataset.

[0010] In one embodiment, the simulated shooting based on the camera parameters and the shooting pose information includes: determining multiple shooting positions for each second image in the second image set; determining the pixels of the target object corresponding to each shooting position based on multiple shooting orientations at each shooting position, obtaining multiple pixel images; determining the image with a sufficient number of pixels among the multiple pixel images as the target pixel image; determining the shooting orientation corresponding to the target pixel image as the target shooting pose of the current shooting position, and obtaining the target shooting pose corresponding to each shooting position among the multiple shooting positions; performing interpolation calculation on the multiple target shooting poses to obtain the shooting pose information of the second image; and performing simulated shooting for each second image based on the camera parameters and the shooting pose information.

[0011] According to a second aspect of the present disclosure, an image recognition model training method is provided, comprising: acquiring a target dataset, wherein the training dataset is obtained using the data acquisition method described in the first aspect or any one of the first aspects; and training an image recognition model based on the target dataset to obtain an image recognition model.

[0012] According to a third aspect of the present disclosure, an image recognition method is provided, comprising: acquiring an image to be recognized, the image to be recognized including an object to be recognized; inputting the image to be recognized into an image recognition model to obtain target annotation data of the object to be recognized; wherein the image recognition model is trained based on a target dataset, and the target dataset is obtained using the data acquisition method described in the first aspect or any one of the first aspects.

[0013] According to a fourth aspect of the present disclosure, a data acquisition device is provided, comprising: an acquisition unit, configured to acquire a first image set, wherein each image in the first image set corresponds to a three-dimensional 3D object and a semantic tag;

[0014] The processing unit is configured to generate a scene image set for each image in the first image set based on the semantic tags, and to synthesize each image with each scene image in the scene image set to obtain a scene image set for each image, and to use the complete scene image set of each image in the first image set as a second image set; and to perform simulation acquisition based on the second image set to obtain a target dataset, wherein the target dataset includes a third image set and labeled data of target objects in the third image set, wherein the third image set is obtained based on simulation rendering of the second image set, and the target objects are determined based on target semantic tags determined in the semantic tags.

[0015] In one embodiment, the scene image set includes an outdoor scene image set; the processing unit generates the scene image set based on the semantic tag in the following manner, and synthesizes each image with each scene image in the scene image set to obtain a scene image set for each image: obtaining outdoor scene parameters associated with the semantic tag; randomly generating outdoor scene component resources based on the outdoor scene parameters; obtaining an outdoor scene image set based on the outdoor scene parameters and the outdoor scene component resources; and synthesizing each image with each scene image in the outdoor scene image set to obtain an outdoor scene image set for each image.

[0016] In one embodiment, the scene image set includes an indoor scene image set; the processing unit generates the scene image set based on the semantic tags in the following manner, and synthesizes each image with each scene image in the scene image set to obtain a scene image set for each image: obtaining indoor scene parameters, the indoor scene parameters including parameters for performing Venn diagram partitioning; performing Venn diagram partitioning on the regions of the indoor scene based on the indoor scene parameters to obtain multiple regions; matching indoor scene component resources in the multiple regions, the indoor scene component resources being associated with the semantic tags; obtaining an indoor scene image set based on the indoor scene parameters and the indoor scene component resources; and synthesizing each image with each scene image in the indoor scene image set to obtain an indoor scene image set for each image.

[0017] In one implementation, the processing unit is further configured to: add a skeletal system object to the images of the second image set.

[0018] In one embodiment, the processing unit adds a skeletal system object to the images of the second image set in the following manner: creating an object template corresponding to the skeletal system object, wherein the object template is a preset character model; randomly setting appearance features for the object template and obtaining the skeletal system object; obtaining the target position of the 3D object in the scene image set; determining the target point of each bone in the skeletal system object based on the target position, wherein the target point is used to determine the position of the bone end of the skeletal system object; calculating the joint angle required to reach the position of the bone end based on the position of the bone end, and adjusting the bones of the skeletal system object according to the joint angle to obtain the adjusted skeletal system object; and adding the adjusted skeletal system object to the target object in the second image set.

[0019] In one embodiment, the processing unit performs simulated acquisition based on the second image set to obtain a target dataset in the following manner: acquiring preset camera parameters, and determining the shooting pose information corresponding to each image in the second image set based on each image in the second image set; performing simulated shooting based on the camera parameters and the shooting pose information to obtain a third dataset after simulated shooting; determining target semantic labels in the semantic labels of the third dataset; and generating labeled data of target objects matching the target semantic labels in the third dataset based on the target semantic labels to obtain the target dataset.

[0020] In one embodiment, the processing unit performs simulated shooting based on the camera parameters and the shooting pose information as follows: determining multiple shooting positions for each second image in the second image set; determining the pixels of the target object corresponding to each shooting position based on multiple shooting orientations, thereby obtaining multiple pixel images; identifying the image with a sufficient number of pixels as the target pixel image from among the multiple pixel images; determining the shooting orientation corresponding to the target pixel image as the target shooting pose of the current shooting position, and obtaining the target shooting pose corresponding to each shooting position among the multiple shooting positions; performing interpolation calculation on the multiple target shooting poses to obtain the shooting pose information of the second image; and performing simulated shooting for each second image based on the camera parameters and the shooting pose information.

[0021] According to a fifth aspect of the present disclosure, an image recognition model training apparatus is provided, comprising: an acquisition unit for acquiring a target dataset, wherein the training dataset is obtained using the data acquisition method described in the first aspect or any one of the first aspects; and a processing unit for training an image recognition model based on the target dataset to obtain an image recognition model.

[0022] According to a sixth aspect of the present disclosure, an image recognition apparatus is provided, comprising: an acquisition unit for acquiring an image to be recognized, the image to be recognized including an object to be recognized; and a processing unit for inputting the image to be recognized into an image recognition model to obtain target annotation data of the object to be recognized; wherein the image recognition model is trained based on a target dataset, and the target dataset is obtained using the data acquisition method described in the first aspect or any one of the first aspects.

[0023] According to a seventh aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to: execute a data acquisition method as described in the first aspect or any embodiment of the first aspect, or execute an image recognition model training method as described in the second aspect, or execute an image recognition method as described in the third aspect.

[0024] According to an eighth aspect of the present disclosure, a storage medium is provided, wherein the storage medium stores instructions that, when executed by a processor, enable the processor to perform the data acquisition method described in the first aspect or any embodiment of the first aspect, or to perform the image recognition model training method described in the second aspect, or to perform the image recognition method described in the third aspect.

[0025] According to a ninth aspect of the present disclosure, a computer program product is provided, the computer program product including a computer program, which, when executed by a processor, implements the data acquisition method described in the first aspect or any embodiment of the first aspect, or implements the image recognition model training method described in the second aspect embodiment, or implements the image recognition method described in the third aspect embodiment.

[0026] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: A first image set is obtained, where each image in the first image set corresponds to a semantic tag. For each image in the first image set, a scene image set related to the semantic tag is generated based on the semantic tag. By synthesizing each image with each scene image in the scene image set, the diversity and generalization ability of the scene image set are increased, and a scene image set for each image can be obtained. This leads to the complete scene image set for all images in the first image set, which is then used as a second image set. Simulation acquisition is performed based on the second image set to simulate complex display environments. The target dataset is obtained at a lower cost, providing a flexible and scalable target dataset for model training and improving the richness of the target dataset.

[0027] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0028] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0029] Figure 1This is a flowchart illustrating a dataset collection method according to an exemplary embodiment.

[0030] Figure 2 This is a flowchart illustrating a method for determining an outdoor scene image set according to an exemplary embodiment.

[0031] Figure 3 This is a flowchart illustrating a method for determining an indoor scene image set according to an exemplary embodiment.

[0032] Figure 4 This is a flowchart illustrating a method for adding a skeletal system object according to an exemplary embodiment.

[0033] Figure 5 This is a flowchart illustrating a method for determining a target dataset according to an exemplary embodiment.

[0034] Figure 6 This is a flowchart illustrating a simulation shooting method according to an exemplary embodiment.

[0035] Figure 7 This is a flowchart illustrating a simulation shooting method according to an exemplary embodiment.

[0036] Figure 8 This is a flowchart illustrating an image recognition model training method according to an exemplary embodiment.

[0037] Figure 9 This is a flowchart illustrating an image recognition method according to an exemplary embodiment.

[0038] Figure 10 This is a schematic diagram illustrating a target dataset and an image recognition model application according to an exemplary embodiment.

[0039] Figure 11 This is a block diagram illustrating a data acquisition device according to an exemplary embodiment.

[0040] Figure 12 This is a block diagram illustrating an image recognition model training apparatus according to an exemplary embodiment.

[0041] Figure 13 This is a block diagram illustrating an image recognition device according to an exemplary embodiment.

[0042] Figure 14 This is a frame of an apparatus shown according to an exemplary embodiment. Figure 1 .

[0043] Figure 15 This is a frame of an apparatus shown according to an exemplary embodiment. Figure 2 . Detailed Implementation

[0044] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure.

[0045] Manually labeled high-precision data is time-consuming, and it is impossible to label various weather conditions, lighting, climates, or dangerous scenarios. Therefore, the dataset collection method provided in this disclosure can collect diverse and highly realistic datasets, which is beneficial for various models to obtain datasets as training data.

[0046] To avoid the high costs and lack of generalization associated with manual annotation, related technologies utilize synthetic data as datasets. Synthetic data is relatively simple to generate, allowing for easy, pixel-level annotation of images programmatically, avoiding ethical and privacy concerns, and enabling the collection of diverse data. Therefore, the use of synthetic data to assist training has been widely adopted in recent years for training and testing various artificial intelligence models. However, when using offline rendering methods to acquire synthetic data, each pair of data requires tens of minutes to generate and render the scene, resulting in low efficiency and difficulty in quickly generating data and iteratively training models. Furthermore, the accuracy, realism, and semantic richness of some synthetic data are limited, with significant domain differences compared to real data, leading to poor performance of training results when tested on real-world datasets.

[0047] In related technologies, the InfiGen procedural generator can be used to generate ultra-large-scale, highly rich outdoor environment scenes, and then rendered offline to obtain relatively realistic synthetic data and annotations. The FaceSynthetics dataset can also be used to learn parameterized facial shape and expression models from three-dimensional (3D) scanned facial data. Shape and expression parameter vectors are randomized, and facial detail textures, hair, clothing, and lighting conditions are randomly selected from an asset library and rendered offline to obtain a facial dataset. The synthetic face dataset can be used to train neural network models for facial region segmentation, facial navigation line prediction, and eye tracking. However, InfiGen can only generate synthetic data for outdoor environments, while FaceSynthetics is limited by the generation of human-related facial data.

[0048] In view of this, this disclosure proposes a dataset collection method that can avoid manual annotation and obtain a dataset with richness and high authenticity.

[0049] Figure 1This is a flowchart illustrating a dataset collection method according to an exemplary embodiment. Figure 1 As shown, the method includes steps S11 to S14.

[0050] In step S11, a first image set is obtained, where each image in the first image set corresponds to a three-dimensional 3D object and has a corresponding semantic tag.

[0051] In this embodiment of the disclosure, the first image set can be an image set of the object obtained by scanning the object from multiple perspectives using an array of multiple cameras. Each image in the first image set corresponds to a 3D object, and each image corresponds to its own semantic tag. The semantic tag can be information for classifying and labeling the 3D objects in the image. The semantic tag can be filtered and eliminated according to the semantic type during simulation and rendering to obtain the target object, and annotation data can also be added to the image.

[0052] In this embodiment, a light field device capable of capturing light information from all directions can be used to scan the optical properties and geometric details of an object's surface. This allows for the extraction and reconstruction of high-resolution base albedo maps, roughness maps, normal maps, and displacement maps, among other textures. The base albedo map represents the object's color information, the roughness map represents the object's surface smoothness, the normal map represents the object's surface orientation information, and the displacement map is used to simulate complex surface details.

[0053] In this embodiment, the extracted and reconstructed image can be imported into a digital content production tool such as an engine. Within this tool, materials, lighting, shadows, and post-processing effects can be further adjusted to ensure that the 3D object visually approximates real-world objects, thus achieving a realistic visual effect. The engine can be a Unity or Blender rendering engine; this disclosure does not specifically limit its use.

[0054] In step S12, for each image in the first image set, a scene image set is generated based on semantic tags.

[0055] In this embodiment of the disclosure, scene images related to the semantic tags can be generated based on the semantic tags corresponding to each object. These scene images can be understood as environmental scene images related to the semantic tags, such as street images. Multiple scene images can be generated for each object, i.e., a scene image set can be generated. It should be noted that a scene image set can also be randomly generated based on each image in the first image set.

[0056] In step S13, each image is synthesized with each scene image in the scene image set to obtain a scene image set for each image, and the complete scene image set of each image in the first image set is used as the second image set.

[0057] In this embodiment of the disclosure, each image corresponds to a scene image set. Each image is synthesized with each scene image in the scene image set to obtain a scene image set for each image. It can be understood that the scene image set for each image includes multiple scene images for each image, and the scene images for each image include the 3D object corresponding to that image.

[0058] In this embodiment of the disclosure, each image corresponds to multiple scene images. Each image is synthesized with each scene image in the multiple scene images to obtain a scene image set corresponding to each image. Since the first image set includes multiple images, the complete scene image set of each image in the first image set can be obtained based on the scene image set corresponding to each image. The complete scene image set corresponding to the first image set is then used as the second image set.

[0059] In step S14, simulation acquisition is performed based on the second image set to obtain the target dataset.

[0060] In this embodiment of the disclosure, the target dataset includes a third image set and labeled data of target objects within the third image set. The third image set is obtained by simulating and rendering a second image set, and the target objects are determined based on target semantic labels identified in the semantic tags.

[0061] Specifically, obtaining the target dataset can be achieved by simulating and rendering a second image set to obtain a third dataset. Based on the semantic labels of 3D objects in the second image set, target semantic labels are determined. The 3D objects corresponding to these target semantic labels are then used as target objects, and these target objects are annotated to obtain their labeled data. Based on the third dataset and the labeled data of the target objects within it, the target dataset is obtained.

[0062] In this embodiment, a scene image set is generated based on the semantic tags corresponding to the images, resulting in semantically complete scene images. Each image is synthesized with each scene image in the scene image set to obtain a scene image set for each image, thus obtaining the complete scene image set for all images in the first image set, achieving scene image diversity. Furthermore, the complete scene image set is used as a second image set. Based on simulation acquisition of the second image set, a target dataset is obtained, which increases the diversity of the target dataset and improves its generalization ability. This facilitates the rapid acquisition of the target dataset as training data for large visual base models and various small models, thereby improving model performance and applicability.

[0063] In this embodiment of the disclosure, the scene may include an outdoor scene and an indoor scene, so the scene image set may include an outdoor scene image set and an indoor scene image set.

[0064] Figure 2 This is a flowchart illustrating a method for determining an outdoor scene image set according to an exemplary embodiment. Figure 2 As shown, the method includes steps S21 to S24.

[0065] In step S21, outdoor scene parameters associated with semantic tags are obtained.

[0066] In this embodiment of the disclosure, outdoor scene parameters can be road-related parameters, such as road width parameters. Semantic tag-related outdoor scene parameters can be outdoor scene parameters associated with 3D objects. For example, if the 3D object is a vehicle, the associated outdoor scene parameters can be road-related parameters.

[0067] In step S22, outdoor scene component resources are randomly generated based on outdoor scene parameters.

[0068] In this embodiment of the disclosure, an outdoor scene is obtained based on outdoor scene parameters. The outdoor scene component resources can be various elements and objects that constitute the outdoor scene, such as green plants, benches, traffic signs, etc. The generated outdoor scene component resources can be randomly scattered in the outdoor scene. For example, a sidewalk with a first road width has been obtained based on the outdoor scene parameters, and outdoor scene component resources such as green plants and benches can be randomly scattered on the sidewalk.

[0069] In step S23, an outdoor scene image set is obtained based on outdoor scene parameters and outdoor scene component resources.

[0070] In this embodiment of the disclosure, based on outdoor scene parameters and outdoor scene component resources, an outdoor scene image set of outdoor scene component resources randomly distributed in the outdoor scene can be obtained.

[0071] In step S24, each image is combined with each scene image in the outdoor scene image set to obtain the outdoor scene image set for each image.

[0072] In this embodiment of the disclosure, an outdoor scene image set is obtained based on the semantic tag corresponding to each image. Each image is then synthesized with each scene image in the outdoor scene image set to obtain an outdoor scene image set for each image.

[0073] For example, a procedural scene generator can be used to automatically synthesize diverse sets of outdoor scene images. This procedural scene generator can be a tool for automatically creating virtual scenes. Each image is input into the procedural scene generator, and the main road mesh is automatically constructed based on several Bézier curves and outdoor scene parameters to obtain the outdoor road. Outdoor scene component resources such as road signs, traffic lines, or manhole covers can be automatically placed at road intersections. The surfaces of road lines, manhole covers, or other outdoor scene component resources are rendered by covering them with graphic decals on the road material. Sidewalks can also be formed on both sides of the outdoor road, and outdoor scene component resources such as greenery and streetlights can be scattered on the sidewalks according to outdoor scene parameters, or preset building models can be randomly placed at intervals on both sides of the sidewalk along the street.

[0074] In this embodiment of the disclosure, by combining outdoor scene component resources and outdoor scene parameters in a random or preset manner, a set of outdoor scene images with a high degree of realism can be generated, which can be used in fields such as game development, film production, urban planning and virtual reality.

[0075] Figure 3 This is a flowchart illustrating a method for determining an indoor scene image set according to an exemplary embodiment. Figure 3 As shown, the method includes steps S31 to S35.

[0076] In step S31, indoor scene parameters are obtained.

[0077] In this embodiment of the disclosure, the indoor scene parameters include parameters for performing Voronoi diagram partitioning. A Voronoi diagram is a method for dividing space into multiple regions; if the indoor scene parameters are obtained, the indoor scene can be partitioned using the Voronoi diagram.

[0078] In step S32, based on the indoor scene parameters, the indoor scene area is divided into multiple regions using a Venn diagram.

[0079] In this embodiment of the disclosure, the indoor scene parameters can be the number of bounding edges and the maximum and minimum interior angle values. Based on the number of bounding edges and the maximum and minimum interior angle values, the floor shape and size can be randomly generated to determine the geometry of the indoor scene.

[0080] In this embodiment of the disclosure, a Voronoi diagram is divided based on indoor scene parameters and Poisson random distribution scatter points to obtain multiple regions of the outdoor scene.

[0081] In step S33, indoor scene component resources are matched in multiple regions, and the indoor scene component resources are associated with semantic tags.

[0082] In this embodiment, the indoor scene component resources can be 3D models of furniture, decorations, etc., placed in the indoor scene. Multiple regions are divided based on a Voronoi diagram, and different functional areas can be defined within these regions. Indoor scene component resources are then matched to these different functional areas to achieve a reasonable distribution of furniture across them. These indoor scene component resources can be of a preset type or associated with semantic tags to facilitate the generation of indoor scenes related to the first image set, thereby improving the semantic realism of the indoor scene.

[0083] In step S34, an indoor scene image set is obtained based on indoor scene parameters and indoor scene component resources.

[0084] In this embodiment of the disclosure, 3D modeling software and rendering technology can be used to obtain an indoor scene image set by combining indoor scene parameters and indoor scene component resources. It is understood that these indoor scene image sets can include multiple perspectives and different combinations of indoor scene parameters and indoor scene component resources to showcase the diversity and realism of the indoor scenes.

[0085] In step S35, each image is synthesized with each scene image in the indoor scene image set to obtain an indoor scene image set for each image.

[0086] In this embodiment of the disclosure, an indoor scene image set is obtained based on indoor scene parameters and indoor scene component resources. Each image is then combined with each scene image in the indoor scene image set to obtain an indoor scene image set for each image.

[0087] Understandably, outdoor and indoor scenes can be preset, generated randomly, or generated based on the semantic tags of the first image.

[0088] In this embodiment of the disclosure, by effectively utilizing indoor scene parameters and indoor scene component resources, a rich set of indoor scene images can be generated, which can be applied to fields such as computer vision, interior design, and virtual reality.

[0089] In this embodiment of the disclosure, when generating a scene image set, in order to improve the realism of the scene, a certain number of high-precision scenes created by artists can be introduced and combined with the indoor scene image set or the outdoor scene image set, thereby reducing the semantic distribution difference between the synthesized scene and the real scene.

[0090] In this embodiment of the disclosure, the second image set may include a scene image set, that is, the second image set is the complete scene image set of all images in the first image set. Furthermore, a skeletal system object may be added to the images in the second image set to obtain a second image set including the skeletal system object, the scene image set, and the first image set.

[0091] In this embodiment of the disclosure, a skeletal system object may be added to the images of the second image set in the following manner. Figure 4 This is a flowchart illustrating a method for adding a skeletal system object according to an exemplary embodiment. Figure 4 As shown, the method includes steps S41 to S46.

[0092] In step S41, an object template corresponding to the skeletal system object is created.

[0093] In this embodiment of the disclosure, the object template is a preset character model that includes a complete skeletal system and can be used for animation and pose adjustments. For example, the object template can be created in CharacterCreator software.

[0094] In step S42, appearance features are randomly set for the object template, and a skeletal system object is obtained.

[0095] In this embodiment of the disclosure, appearance features may include clothing and accessories, skin color, hairstyle, height, facial features and clothing, etc. Randomly setting the above appearance features for the object template results in diverse and personalized skeletal system objects.

[0096] In step S43, the target position of the 3D object in the scene image set is obtained.

[0097] In this embodiment of the disclosure, the skeletal system object can interact with 3D objects in the scene image set to ensure the pose diversity of the skeletal system object and the rationality of the interaction semantics between the skeletal system object and the scene in the scene image set.

[0098] In this embodiment of the disclosure, the target position of a 3D object is obtained. For example, if the 3D object is a sofa, the target position of the sofa can be obtained by obtaining the position and size of the sofa in the scene image, or by obtaining the shape, surface angle or physical properties of the sofa.

[0099] In step S44, the target point of each bone in the skeletal system object is determined based on the target location.

[0100] In this embodiment, the target point is used to determine the position of the bone ends of the skeletal system object, such as the position of the sofa where the buttocks of the skeletal system object contact when sitting down. Since a skeletal system is created for the object template when it is created, the skeletal system object has a skeletal structure. Different types of inverse dynamics constraint components are automatically bound to the bones of the skeletal system object, such as the hands and feet, so that the data acquisition personnel can control the movement of the skeletal system object by moving the bones, ensuring that each bone in the skeletal system object reaches its end-point position.

[0101] In this embodiment of the disclosure, based on the target position of the 3D object, the target point of each bone in the skeletal system object can be determined, and the position of the bone end of the skeletal system object can be determined based on the target point, thereby controlling each bone of the skeletal system object to reach the position of the bone end.

[0102] In step S45, based on the position of the bone end, the joint angle required to reach the position of the bone end is calculated, and the bones of the skeletal system object are adjusted according to the joint angle to obtain the adjusted skeletal system object.

[0103] In this embodiment of the disclosure, since each bone of the skeletal system object has reached the end position of the bone, the joint angle required for the skeletal system object to reach the end position of the bone can be calculated in reverse from the end position of the bone, and the bones of the skeletal system object can be adjusted according to the joint angle, thereby obtaining the adjusted skeletal system object.

[0104] In step S46, the adjusted skeletal system object is added to the target object of the second image set.

[0105] In this embodiment of the disclosure, the adjusted skeletal system object is the skeletal system object after interacting with the scene. The adjusted skeletal system object is added to the target object of the second image set so that information such as facial feature point annotations of the skeletal system object can be output.

[0106] In this embodiment, a 3D object is used as a sofa, and a skeletal system object is used as a character. The process of adding a skeletal system object is described below. Several character templates are created, and appearance features are randomly replaced. Skin and clothing are automatically bound to skeletons to obtain a virtual character. The virtual character is imported into the Unity engine, and different types of inverse dynamics constraint components are automatically bound to the skeletal points of the virtual character, such as hands, feet, and hips. The target position of the 3D sofa in the scene image is determined, and a suitable animation type is randomly selected from the animation library, such as a talking standing posture, a sitting posture, or a hand-holding posture. Using the inverse dynamics constraint components and the target position of the sofa, a target point is determined. Based on the target point, the local posture and position of the virtual character are automatically adjusted and blended with the animation from the animation library to obtain a virtual character sitting on a sofa.

[0107] In this embodiment of the disclosure, the interaction between the scene image set and the skeletal system object is based on the scene image set to obtain the skeletal system object that is semantically related to the scene image set, thereby improving the realism and interactivity between the scene and the object, and generating image sets with complex actions and interactions.

[0108] In this embodiment of the disclosure, since the standing posture, sitting posture, etc. of the skeletal system object can be determined based on the scene image, the skeletal system object is determined in real time. Based on the adjusted skeletal system object, facial feature point annotations, segmentation annotations of different parts, and skeletal joint axis-angle vectors are generated. Furthermore, the pose parameters corresponding to the current posture of the skeletal system object can be converted into a vertex-based Skinned Multi-Person Linear Model (SMPL) for easy subsequent direct use.

[0109] In this embodiment of the disclosure, the second image set can be simulated and rendered to obtain the rendered target dataset.

[0110] Figure 5 This is a flowchart illustrating a method for determining a target dataset according to an exemplary embodiment. Figure 5 As shown, the method includes steps S51 to S54.

[0111] In step S51, preset camera parameters are obtained, and shooting pose information corresponding to each image in the second image set is determined based on each image in the second image set.

[0112] In this embodiment of the disclosure, the camera parameters may be resolution, focal length, exposure parameters, and imaging parameters, etc. That is, the camera parameters are the parameters involved in the shooting process, and the camera parameters can be preset.

[0113] In this embodiment of the disclosure, the irradiance energy value of the light source in each image can be determined based on each image in the second image set, and the irradiance energy value of the light source can also be physically controlled.

[0114] In this embodiment of the disclosure, the shooting pose information can be the position and orientation of the camera. It is understood that the camera shooting pose information involved here is not the actual shooting using the camera in a real-world environment, but rather the determination of the position and orientation to simulate a camera shooting environment, obtaining images with different orientations and positions under the simulated shooting environment.

[0115] In step S52, based on camera parameters and shooting pose information, a simulated shooting is performed to obtain a third dataset after the simulated shooting.

[0116] In this embodiment of the disclosure, a third dataset is obtained by performing simulated shooting based on preset camera parameters and shooting pose information. The simulated shooting does not involve shooting every image in the second image set, but rather shooting the environment composed of 3D objects, scenes, and skeletal system objects in each image in the second image set, resulting in multiple images as the third dataset.

[0117] In step S53, the target semantic label is determined from the semantic labels of the third dataset.

[0118] In this embodiment of the disclosure, since the third dataset contains a first image set of 3D objects with their own semantic labels, the semantic label of the data to be annotated can be determined from the semantic labels of the third dataset as the target semantic label.

[0119] In step S54, based on the target semantic label, labeled data of the target objects that match the target semantic label in the third dataset are generated to obtain the target dataset.

[0120] In this embodiment of the disclosure, the third dataset includes rendered images. Based on the target semantic labels, annotation data of target objects matching the target semantic labels in the third dataset can be generated. For example, if the 3D object is an electric vehicle, the semantic labels of the electric vehicle include headlights and wheels, etc., and the wheels are determined to be the target semantic label. The annotation data of the wheels in the third image set is obtained, such as the color and material of the wheels. This embodiment of the disclosure is only a brief illustrative example and is not intended to be specific.

[0121] In this embodiment of the disclosure, real-time ray tracing can be performed using DirectX RayTracing (DXR) to reflect the dynamic performance under different lighting conditions such as intensity and day / night, generating a highly realistic third image set.

[0122] In this embodiment of the disclosure, when simulating and rendering the second dataset, a data acquisition environment based on a real-time rendering pipeline can be constructed. Within this acquisition environment, pixel shaders or computation shaders can be written to call graphics processor drawing commands to annotate the target objects corresponding to the target semantic labels. Asynchronous multithreading can also be used to read back and store the image and labeled image of each frame in the graphics processor, thereby optimizing runtime performance.

[0123] In this embodiment of the disclosure, simulated shooting can be performed based on camera parameters and shooting pose information in the following manner. Figure 6 This is a flowchart illustrating a simulation shooting method according to an exemplary embodiment. For example... Figure 6 As shown, the method includes steps S61 to S66.

[0124] In step S61, multiple shooting positions for each second image in the second image set are determined.

[0125] In this embodiment of the disclosure, the camera's movement trajectory can be determined according to the A* pathfinding algorithm. In each second image of the second image set, multiple starting points and target points can be determined to obtain multiple movement trajectories. A target object is preset, such as an object of interest, and the movement trajectory including the object of interest is determined as the target movement trajectory. Multiple nodes on the target movement trajectory are determined as multiple shooting positions.

[0126] In this embodiment of the disclosure, the starting point and the target point can be pixels, grid points, or other forms of discrete points.

[0127] In step S62, based on multiple shooting orientations, the pixels of the target object corresponding to each shooting position are determined at each shooting position to obtain multiple pixel images.

[0128] In this embodiment of the disclosure, the shooting orientation can be a 360-degree virtual viewpoint in six directions: up, down, left, right, front, and back. Based on multiple shooting orientations at each shooting position, multiple virtual viewpoints are determined, and the pixels of the object of interest at each shooting position within the virtual viewpoints are obtained, resulting in a pixel image. The 360-degree images captured from the six directions can be merged into a complete spherical image, thereby creating a seamless panoramic view. Therefore, each shooting orientation can be an orientation that allows viewing one-sixth of the spherical image.

[0129] In step S63, the image whose number of pixels meets the pixel count requirement is determined from the multiple pixel images as the target pixel image.

[0130] In this embodiment of the disclosure, the number of pixels of the object of interest within the rectangular view range can be specified as a metric to determine the target pixel image. For example, among multiple pixel images, the image with the most pixels can be determined as the target pixel image, since the area corresponding to the image with the most pixels is more significant, and therefore can be used as the target pixel image.

[0131] In step S64, the target shooting pose corresponding to the shooting orientation of the target pixel image is determined as the target shooting pose of the current shooting position, and the target shooting pose corresponding to each shooting position among multiple shooting positions is obtained.

[0132] In this embodiment of the disclosure, the target shooting pose of the current shooting position is determined based on the shooting orientation corresponding to the target pixel image. It is understood that there is a one-to-one correspondence between shooting position and shooting orientation; each shooting position has one shooting orientation. Therefore, both the shooting position and shooting orientation can be determined simultaneously based on the target pixel image, thereby obtaining the target shooting pose. It is also understood that the target shooting pose can be obtained for each of multiple shooting positions.

[0133] In step S65, interpolation calculations are performed on the shooting poses of multiple targets to obtain the shooting pose information of the second image.

[0134] In this embodiment of the disclosure, interpolation calculation is performed on multiple target shooting poses to obtain the lens transition between multiple shooting positions of the camera, and then the shooting pose obtained after interpolation calculation is used as the shooting pose information of the second image.

[0135] In this embodiment of the disclosure, since there may be multiple shooting positions in the target movement path, in order to avoid excessive computing resources, an interpolation calculation method can be adopted to reduce the shooting orientation at each shooting position.

[0136] In step S66, each second image is simulated and captured based on camera parameters and shooting pose information.

[0137] In this embodiment of the disclosure, based on preset camera parameters such as resolution and focal length, and shooting pose information with different shooting orientations at different shooting positions, the scene in each second image is simulated and shot, resulting in multiple frames of images at different shooting positions and with different shooting orientations. The shooting pose information is stored in the data packet of the corresponding second image so that it can be used in video or models with continuous information perception.

[0138] In this embodiment of the disclosure, camera shake can also be applied along the target movement path to simulate handheld shooting, thereby improving the realism of the third image set.

[0139] In this embodiment of the disclosure, the pose information of each frame can be randomly randomized in a skip-style manner to determine the shooting pose information. It can also automatically detect and filter out image frames with erroneous semantics, low pixel count, or a small number of pixels. Erroneous semantics could include images such as an electronic device appearing on a wall.

[0140] In this embodiment of the disclosure, data users can also manually and interactively set shooting pose information and collect image data in a scene according to the specific data requirements of a specific model.

[0141] In this embodiment of the disclosure, the use of omnidirectional perspective shooting can create a diverse set of images. Combining camera parameters and shooting pose parameters improves the realism of image generation, providing an accurate foundation for subsequent training of the dataset. Furthermore, it creates smooth shot transitions between each frame to achieve visual coherence and appeal, ensuring that the target dataset has a high degree of realism and diversity.

[0142] The embodiments disclosed herein are described below in conjunction with Figure 7 The data collection method is explained. Figure 7 This is a flowchart illustrating a simulation shooting method according to an exemplary embodiment. For example... Figure 7 As shown, in high-precision scanning resources, in order to fit the semantic scenes of daily life, camera arrays are usually used to reconstruct the shape mesh of objects and optimize the topology. At the same time, a 360-degree light field device is used to scan and measure the optical properties and geometric details of the object surface, extract and reconstruct 8K resolution base albedo maps, roughness maps, normal maps, displacement maps and other textures, and import them into Unity engine or digital content production tools such as Blender to synthesize hyper-realistic 3D objects.

[0143] In this embodiment of the disclosure, a procedural scene generator is used to automatically synthesize diverse and semantically complete virtual 3D object scenes by inputting 3D objects into the generator. Examples include outdoor or indoor scenes. For the procedural generation of outdoor street scene image sets, the process involves automatically constructing a main road grid surface based on several Bézier curves and road width parameters. Road signs and traffic lines are automatically placed at road intersections, and surfaces such as traffic lines and manhole covers are rendered by covering them with graphic decals on the road material. Sidewalks are formed on both sides of the road, and green plants and various road component resources are scattered on the sidewalks according to scene parameters. Preset building models are randomly placed at intervals on both sides of the sidewalks along the street. For the procedural generation of indoor scene image sets, the process first randomly generates the floor shape and size based on the number of bounding edges and the maximum and minimum interior angle values. Then, a Voronoi diagram is created based on a Poisson random distribution, and furniture is then reasonably distributed within each area.

[0144] In this embodiment of the disclosure, in addition to the procedurally generated scenes, the scene set will also include a certain number of high-precision environmental scenes manually created by artists, which will be transmitted into the simulation acquisition system to reduce the difference in semantic distribution between the synthesized scene environment and the real environment.

[0145] In this embodiment, for the skeletal binding task and the procedural environmental interaction animation process based on inverse dynamics, several preset object templates can be created in CharacterCreator software, and clothing accessories, skin color, and hairstyle can be randomly replaced. Height and facial features can be randomly adjusted within a small range, and skeletal binding of skin and clothing can be performed automatically. During the import of the skeletal system object into Unity, different types of inverse dynamics constraint components are automatically bound to its hand, foot, and hip bone points. This allows the skeletal system object to be placed in the environment, and to randomly select appropriate animation types from the animation library based on other surrounding skeletal system objects and interactive environmental objects such as beds and chairs. Examples include standing postures, sitting postures, and hand-holding postures. Simultaneously, the local posture and position of the skeletal system object are automatically adjusted through inverse dynamics constraints and mixed with the animations in the animation library. This achieves the diversity of automated posture sampling for the skeletal system object while ensuring the semantic rationality of the interaction between the skeletal system object and its surrounding environment.

[0146] In this embodiment, outdoor or indoor scenes and skeletal system objects are input into a simulation acquisition system to form a data acquisition environment based on a real-time rendering pipeline. According to pre-configured scene parameters, the lighting environment, i.e., the irradiance energy value of the light source, and camera parameters, such as the camera's exposure parameters, are physically controlled. Since the surface attributes of scene resources and characters are also physically based, and real-time ray tracing lighting calculation technology based on DirectX RayTracing is used, the dynamic appearance under different lighting conditions such as intensity and day / night can be reflected, generating highly realistic rendered images.

[0147] In the camera parameter control of this disclosure, in addition to controlling the output image resolution and focal length, the control of the camera position and orientation has two modes: automatic and interactive. The automatic control mode is divided into two schemes: the first is a progressive approach, which first determines the camera's position trajectory based on the A* pathfinding algorithm, then equally selects several nodes on the path. At each node, the pixels of the object of interest are drawn using a 360-degree virtual viewpoint in six directions (up, down, left, right, front, and back). The number of pixels of the object of interest within a specified rectangular view is used as a measure of regional saliency, resulting in a saliency detection map. The most suitable viewpoint in the panoramic image is selected as the camera pose for each node, thereby interpolating the lens transition between these nodes. Furthermore, camera shake can be applied on the path to simulate handheld shooting. This method can output continuous frame data and can be used for video or deep learning models with temporal continuous information awareness. The second is a discrete approach, which randomly randomizes the camera pose in a jump-like manner for each frame and automatically detects and filters out image frames with erroneous semantics or low saliency. The interactive control mode allows data users to manually and interactively collect image data in the scene according to the specific data requirements of a specific model.

[0148] In this embodiment, the required data annotation is mainly achieved by writing pixel shaders or computation shaders, and then calling graphics processor drawing commands to perform corresponding annotation data on images or image sequences, thereby obtaining images or image sequences and corresponding annotation data. This output data can be used for model training. In the acquisition environment, asynchronous multithreading can be used to read back and store the main image and annotation image of each frame from the graphics processor to optimize runtime performance.

[0149] In this embodiment, diverse and highly realistic visual datasets can be automatically collected in virtual scenes, with resolutions ranging from 4K to 8K, eliminating the time and expense of manual data collection and annotation. It can cover numerous different types of labeled data output. Existing resources can be reused for different annotation needs. It facilitates the rapid acquisition of training data for large visual models and various small models, enabling iterative improvement of model capabilities.

[0150] Based on the same concept, this disclosure also provides a method for training an image recognition model. Figure 8 This is a flowchart illustrating an image recognition model training method according to an exemplary embodiment. For example... Figure 8 As shown, the method includes steps S71 to S72.

[0151] In step S71, the target dataset is obtained.

[0152] In this embodiment of the disclosure, the training dataset can be adopted. Figures 1 to 6 The dataset collection method involved in the specific embodiments or related embodiments is obtained.

[0153] In step S72, the image recognition model is trained based on the target dataset to obtain the image recognition model.

[0154] In this embodiment of the disclosure, an image recognition model is trained based on a target dataset to obtain the image recognition model.

[0155] In this embodiment of the disclosure, the target dataset can be used to estimate image depth on an electronic device chip to perform a focus blurring algorithm, or to train a model with segmentation and bounding box data and use it for target detection and tracking in an electronic device chip, or for model training in robotics or autonomous driving tasks, or human image-related data can be used for human pose estimation to perform machine control, facial expression recognition, and automatic image retouching models.

[0156] Based on the same concept, this disclosure also provides an image recognition method. Figure 9 This is a flowchart illustrating an image recognition method according to an exemplary embodiment. Figure 9 As shown, the method includes steps S81 to S82.

[0157] In step S81, the image to be identified is acquired.

[0158] In this embodiment of the disclosure, the image to be identified includes an object to be identified.

[0159] In step S82, the image to be identified is input into the image recognition model to obtain the target annotation data of the object to be identified.

[0160] In this embodiment of the disclosure, the image recognition model is trained based on the target dataset.

[0161] For example, in combination Figure 10 The application scenarios for the target dataset and image recognition model are described. For example, it could involve the automated generation of datasets for visual perception deep learning neural networks and the rapid iterative optimization of algorithm models, replacing some of the time-consuming and expensive manual data annotation process. Figure 10 As shown, Figure 10 This is a schematic diagram illustrating a target dataset and an image recognition model application according to an exemplary embodiment. By inputting 3D objects and setting scene parameters, a virtual scene is generated programmatically, continuously constructing diverse and semantically complete scenes, and placing skeletal system objects within them to select appropriate interactive poses. Annotated images or continuous frame data are acquired through an interactive or automated acquisition program to obtain a target dataset. The image recognition model is then trained using this target dataset, and the dataset and model can be iterated to obtain a semantically complete image recognition model. The target dataset includes camera depth maps, semantic segmentation and instance segmentation maps, 2D / 3D bounding boxes of objects, skeletal axis and angle data of human figures, and facial keypoint coordinates. For example, if a business model without labeled data needs annotation, scene parameters from similar previous business operations can be reused to automatically add annotations to the dataset, eliminating the need to re-acquire and annotate the original images one by one, significantly saving costs and time compared to manual annotation. Furthermore, due to the real-time interactivity of the acquisition environment, a robot control system model can be integrated for visual feedback-based physical simulation training within the environment.

[0162] Based on the same concept, this disclosure also provides a dataset acquisition device 100.

[0163] It is understood that the dataset acquisition device 100 provided in this disclosure includes hardware structures and / or software modules corresponding to each function in order to achieve the above-mentioned functions. In conjunction with the units and algorithm steps of the various examples disclosed in this disclosure, this disclosure can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the technical solutions of this disclosure.

[0164] Figure 11 This is a block diagram illustrating a data acquisition device according to an exemplary embodiment. (Refer to...) Figure 11 The device 100 includes an acquisition unit 101 and a processing unit 102.

[0165] The acquisition unit 101 is used to acquire a first image set, wherein each image in the first image set corresponds to a three-dimensional 3D object and has a corresponding semantic tag.

[0166] The processing unit 102 is configured to generate a scene image set for each image in the first image set based on the semantic tags, and to synthesize each image with each scene image in the scene image set to obtain a scene image set for each image, and to take the complete scene image set of each image in the first image set as a second image set; and to perform simulation acquisition based on the second image set to obtain a target dataset, wherein the target dataset includes a third image set and labeled data of target objects in the third image set, wherein the third image set is obtained based on simulation rendering of the second image set, and the target objects are determined based on target semantic tags determined in the semantic tags.

[0167] In some embodiments, the scene image set includes an outdoor scene image set; the processing unit 102 generates the scene image set based on semantic tags in the following manner, and synthesizes each image with each scene image in the scene image set to obtain the scene image set for each image: obtaining outdoor scene parameters associated with semantic tags; randomly generating outdoor scene component resources based on outdoor scene parameters; obtaining an outdoor scene image set based on outdoor scene parameters and outdoor scene component resources; and synthesizing each image with each scene image in the outdoor scene image set to obtain the outdoor scene image set for each image.

[0168] In some embodiments, the scene image set includes an indoor scene image set; the processing unit 102 generates the scene image set based on semantic tags in the following manner, and synthesizes each image with each scene image in the scene image set to obtain a scene image set for each image: obtaining indoor scene parameters, including parameters for performing Venn diagram partitioning; performing Venn diagram partitioning on the regions of the indoor scene based on the indoor scene parameters to obtain multiple regions; matching indoor scene component resources in the multiple regions, the indoor scene component resources being associated with semantic tags; obtaining an indoor scene image set based on the indoor scene parameters and the indoor scene component resources; and synthesizing each image with each scene image in the indoor scene image set to obtain an indoor scene image set for each image.

[0169] In some embodiments, the processing unit 102 is further configured to: add a skeletal system object to an image of the second image set.

[0170] In some embodiments, the processing unit 102 adds a skeletal system object to the images of the second image set in the following manner: creating an object template corresponding to the skeletal system object, wherein the object template is a preset character model; randomly setting appearance features for the object template and obtaining the skeletal system object; obtaining the target position of the 3D object in the scene image set; determining the target point of each bone in the skeletal system object based on the target position, wherein the target point is used to determine the position of the bone end of the skeletal system object; calculating the joint angle required to reach the position of the bone end based on the position of the bone end, and adjusting the bones of the skeletal system object according to the joint angle to obtain the adjusted skeletal system object; and adding the adjusted skeletal system object to the target object in the second image set.

[0171] In some embodiments, the processing unit 102 performs simulated acquisition based on the second image set to obtain the target dataset in the following manner: acquiring preset camera parameters, and determining the shooting pose information corresponding to each image in the second image set based on each image in the second image set; performing simulated shooting based on the camera parameters and shooting pose information to obtain the third dataset after simulated shooting; determining the target semantic label in the semantic label of the third dataset; and generating labeled data of the target objects matching the target semantic label in the third dataset based on the target semantic label to obtain the target dataset.

[0172] In some embodiments, the processing unit 102 performs simulated shooting based on camera parameters and shooting pose information in the following manner: determining multiple shooting positions for each second image in the second image set; determining the pixels of the target object corresponding to each shooting position based on multiple shooting orientations, and obtaining multiple pixel images; determining the image with a sufficient number of pixels as the target pixel image among the multiple pixel images; determining the shooting orientation corresponding to the target pixel image as the target shooting pose of the current shooting position, and obtaining the target shooting pose corresponding to each shooting position among the multiple shooting positions; performing interpolation calculation on the multiple target shooting poses to obtain the shooting pose information of the second image; and performing simulated shooting for each second image based on camera parameters and shooting pose information.

[0173] Figure 12 This is a block diagram illustrating an image recognition model training apparatus 200 according to an exemplary embodiment. (Refer to...) Figure 12 The device 200 includes an acquisition unit 201 and a processing unit 202.

[0174] Acquisition unit 201 is used to acquire the target dataset. The training dataset adopts the above-mentioned... Figures 1-6 The data was acquired using the data acquisition method described in the embodiments or related embodiments.

[0175] The processing unit 202 is used to train the image recognition model based on the target dataset to obtain the image recognition model.

[0176] Figure 13 This is a block diagram illustrating an image recognition device 300 according to an exemplary embodiment. (Refer to...) Figure 13 The device 300 includes an acquisition unit 301 and a processing unit 302.

[0177] The acquisition unit 301 is used to acquire an image to be recognized, which includes an object to be recognized.

[0178] Processing unit 302 is used to input the image to be recognized into an image recognition model to obtain target annotation data of the object to be recognized. The image recognition model is trained based on a target dataset, which uses the aforementioned... Figures 1-6 The data was acquired using the data acquisition method described in the embodiments or related embodiments.

[0179] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0180] Figure 14 This is a frame of an apparatus shown according to an exemplary embodiment. Figure 1Device 400 can be provided as a terminal. Device 400 can be used to perform the above-described data acquisition method, image recognition model training method, or image recognition method. For example, device 400 can be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0181] Reference Figure 14 The device 400 may include one or more of the following components: processing component 402, memory 404, power component 406, multimedia component 408, audio component 410, input / output (I / O) interface 412, sensor component 414, and communication component 416.

[0182] Processing component 402 typically controls the overall operation of device 400, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 402 may include one or more processors 420 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 402 may include one or more modules to facilitate interaction between processing component 402 and other components. For example, processing component 402 may include a multimedia module to facilitate interaction between multimedia component 408 and processing component 402.

[0183] Memory 404 is configured to store various types of data to support the operation of device 400. Examples of such data include instructions for any application or method operating on device 400, contact data, phonebook data, messages, pictures, videos, etc. Memory 404 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0184] The power supply component 406 provides power to the various components of the device 400. The power supply component 406 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 400.

[0185] Multimedia component 408 includes a screen that provides an output interface between the device 400 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 408 includes a front-facing camera and / or a rear-facing camera. When the device 400 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0186] Audio component 410 is configured to output and / or input audio signals. For example, audio component 410 includes a microphone (MIC) configured to receive external audio signals when device 400 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 404 or transmitted via communication component 416. In some embodiments, audio component 410 also includes a speaker for outputting audio signals.

[0187] I / O interface 412 provides an interface between processing component 402 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0188] Sensor assembly 414 includes one or more sensors for providing status assessments of various aspects of device 400. For example, sensor assembly 414 may detect the on / off state of device 400, the relative positioning of components such as the display and keypad of device 400, changes in the position of device 400 or a component of device 400, the presence or absence of user contact with device 400, the orientation or acceleration / deceleration of device 400, and temperature changes of device 400. Sensor assembly 414 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 414 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 414 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0189] Communication component 416 is configured to facilitate wired or wireless communication between device 400 and other devices. Device 400 can access wireless networks based on communication standards, such as WiFi, 4G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 416 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 416 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0190] In an exemplary embodiment, the apparatus 400 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0191] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 404 including instructions, which can be executed by a processor 420 of the device 400 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0192] Figure 15 This is a frame of an apparatus shown according to an exemplary embodiment. Figure 2 For example, device 500 can be provided as a server. Device 500 can be used to perform the above-described data acquisition method, image recognition model training method, or image recognition method. (See reference...) Figure 15 The apparatus 500 includes a processing component 522, which further includes one or more processors, and memory resources represented by memory 532 for storing instructions, such as application programs, that can be executed by the processing component 522. The application programs stored in memory 532 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 522 is configured to execute instructions to perform the methods described above.

[0193] Device 500 may also include a power supply component 526 configured to perform power management of device 500, a wired or wireless network interface 550 configured to connect device 500 to a network, and an input / output (I / O) interface 558. Device 500 may operate on an operating system stored in memory 532, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.

[0194] Based on the same concept, this disclosure also provides a computer program product, wherein the computer program product includes a computer program. This computer program can be executed by a processor, and when executed by the processor, it can perform any of the data acquisition methods described in the above embodiments, or perform the image recognition model training method described in the above embodiments, or perform the image recognition method described in the above embodiments.

[0195] It is understood that in this disclosure, "multiple" refers to two or more, and other quantifiers are similar. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. The singular forms "a," "the," and "the" are also intended to include the plural forms unless the context clearly indicates otherwise.

[0196] It is further understood that the terms "first," "second," etc., are used to describe various types of information, but this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another, and do not indicate a specific order or degree of importance. In fact, the expressions "first," "second," etc., are completely interchangeable. For example, without departing from the scope of this disclosure, first information can also be referred to as second information, and similarly, second information can also be referred to as first information.

[0197] It is further understood that the terms “center,” “longitudinal,” “lateral,” “front,” “rear,” “up,” “down,” “left,” “right,” “vertical,” “horizontal,” “top,” “bottom,” “inner,” and “outer,” etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this embodiment and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation.

[0198] It can be further understood that, unless otherwise specified, "connection" includes both direct connections where no other components exist between the two parties and indirect connections where other components exist between them.

[0199] It is further understood that although operations are described in a specific order in the accompanying drawings in the embodiments of this disclosure, this should not be construed as requiring these operations to be performed in the specific order or serial order shown, or requiring all of the shown operations to be performed to obtain the desired result. In certain environments, multitasking and parallel processing may be advantageous.

[0200] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein.

[0201] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A dataset collection method, characterized in that, include: Obtain a first image set, where each image in the first image set corresponds to a three-dimensional 3D object and has a corresponding semantic tag; For each image in the first image set, a scene image set is generated based on the semantic tag, and each image is synthesized with each scene image in the scene image set to obtain a scene image set for each image. The complete scene image set of each image in the first image set is used as the second image set. Based on the second image set, a target dataset is obtained by simulation acquisition. The target dataset includes a third image set and labeled data of target objects in the third image set. The third image set is obtained by simulation rendering of the second image set, and the target objects are determined based on the target semantic labels determined in the semantic labels.

2. The method according to claim 1, characterized in that, The scene image set includes an outdoor scene image set; The process of generating a scene image set based on the semantic tags, and then combining each image with each scene image in the scene image set to obtain a scene image set for each image, includes: Obtain the outdoor scene parameters associated with the semantic tag; Based on the outdoor scene parameters, outdoor scene component resources are randomly generated; Based on the outdoor scene parameters and the outdoor scene component resources, an outdoor scene image set is obtained; Each image is synthesized with each scene image in the outdoor scene image set to obtain an outdoor scene image set for each image.

3. The method according to claim 1 or 2, characterized in that, The scene image set includes an indoor scene image set; The process of generating a scene image set based on the semantic tags, and then combining each image with each scene image in the scene image set to obtain a scene image set for each image, includes: Obtain indoor scene parameters, including parameters for performing Venn diagram partitioning; Based on the indoor scene parameters, the indoor scene area is divided into multiple regions using a Venn diagram. Match indoor scene component resources in the multiple regions, wherein the indoor scene component resources are associated with the semantic tags; Based on the indoor scene parameters and the indoor scene component resources, an indoor scene image set is obtained; Each image is synthesized with each scene image in the indoor scene image set to obtain an indoor scene image set for each image.

4. The method according to claim 1, characterized in that, The method further includes: Add a skeletal system object to the images in the second image set.

5. The method according to claim 4, characterized in that, Adding a skeletal system object to the images in the second image set includes: Create an object template for the corresponding skeletal system object, wherein the object template is a preset character model; Randomly set appearance features for the object template to obtain a skeletal system object; Obtain the target position of the 3D object in the scene image set; Based on the target location, a target point is determined for each bone in the skeletal system object, and the target point is used to determine the position of the bone end of the skeletal system object. Based on the position of the bone end, the joint angle required to reach the position of the bone end is calculated, and the bones of the skeletal system object are adjusted according to the joint angle to obtain the adjusted skeletal system object. Add the adjusted skeletal system object to the target object of the second image set.

6. The method according to claim 5, characterized in that, The simulation acquisition based on the second image set to obtain the target dataset includes: Obtain preset camera parameters, and based on each image in the second image set, determine the shooting pose information corresponding to each image in the second image set; Based on the camera parameters and the shooting pose information, a simulated shooting is performed to obtain a third dataset after the simulated shooting. Determine the target semantic label from the semantic labels in the third dataset; Based on the target semantic label, labeled data of target objects matching the target semantic label in the third dataset are generated to obtain the target dataset.

7. The method according to claim 6, characterized in that, The simulated shooting based on the camera parameters and the shooting pose information includes: Determine multiple shooting locations for each second image in the second image set; Based on multiple shooting orientations at each shooting position, the pixels of the target object corresponding to each shooting position are determined to obtain multiple pixel images; Among the plurality of pixel images, the image whose number of pixels meets the pixel count requirement is determined as the target pixel image; The target shooting pose corresponding to the shooting orientation of the target pixel image is determined as the target shooting pose of the current shooting position, and the target shooting pose corresponding to each shooting position among the multiple shooting positions is obtained; Interpolation calculations are performed on the shooting poses of the multiple targets to obtain the shooting pose information of the second image; Based on the camera parameters and the shooting pose information, each second image is simulated for shooting.

8. A method for training an image recognition model, characterized in that, include: The target dataset is obtained, wherein the training dataset is collected using the method described in any one of claims 1 to 7; Based on the target dataset, an image recognition model is trained to obtain the image recognition model.

9. An image recognition method, characterized in that, include: Obtain an image to be identified, wherein the image to be identified includes an object to be identified; The image to be identified is input into the image recognition model to obtain the target annotation data of the object to be identified; The image recognition model is trained based on a target dataset, which is trained using the method described in any one of claims 1 to 7.

10. A dataset acquisition device, characterized in that, include: The acquisition unit is used to acquire a first image set, wherein each image in the first image set corresponds to a three-dimensional 3D object and has a corresponding semantic tag; The processing unit is configured to generate a scene image set based on the semantic tag for each image in the first image set, and to synthesize each image with each scene image in the scene image set to obtain a scene image set for each image, and to use the complete scene image set of each image in the first image set as a second image set; It is used to perform simulation acquisition based on the second image set to obtain a target dataset, which includes a third image set and labeled data of target objects in the third image set. The third image set is obtained based on simulation rendering of the second image set, and the target objects are determined based on the target semantic labels determined in the semantic labels.

11. The apparatus according to claim 10, characterized in that, The scene image set includes an outdoor scene image set; The processing unit generates a scene image set based on the semantic tags in the following manner, and synthesizes each image with each scene image in the scene image set to obtain a scene image set for each image: Obtain the outdoor scene parameters associated with the semantic tag; Based on the outdoor scene parameters, outdoor scene component resources are randomly generated; Based on the outdoor scene parameters and the outdoor scene component resources, an outdoor scene image set is obtained; Each image is synthesized with each scene image in the outdoor scene image set to obtain an outdoor scene image set for each image.

12. The apparatus according to claim 10 or 11, characterized in that, The scene image set includes an indoor scene image set; The processing unit generates a scene image set based on the semantic tags in the following manner, and synthesizes each image with each scene image in the scene image set to obtain a scene image set for each image: Obtain indoor scene parameters, including parameters for performing Venn diagram partitioning; Based on the indoor scene parameters, the indoor scene area is divided into multiple regions using a Venn diagram. Match indoor scene component resources in the multiple regions, wherein the indoor scene component resources are associated with the semantic tags; Based on the indoor scene parameters and the indoor scene component resources, an indoor scene image set is obtained; Each image is synthesized with each scene image in the indoor scene image set to obtain an indoor scene image set for each image.

13. The apparatus according to claim 10, characterized in that, The processing unit is also used for: Add a skeletal system object to the images in the second image set.

14. The apparatus according to claim 13, characterized in that, The processing unit adds skeletal system objects to the images in the second image set in the following manner: Create an object template for the corresponding skeletal system object, wherein the object template is a preset character model; Randomly set appearance features for the object template to obtain a skeletal system object; Obtain the target position of the 3D object in the scene image set; Based on the target location, a target point is determined for each bone in the skeletal system object, and the target point is used to determine the position of the bone end of the skeletal system object. Based on the position of the bone end, the joint angle required to reach the position of the bone end is calculated, and the bones of the skeletal system object are adjusted according to the joint angle to obtain the adjusted skeletal system object. Add the adjusted skeletal system object to the target object of the second image set.

15. The apparatus according to claim 14, characterized in that, The processing unit performs simulation acquisition based on the second image set to obtain the target dataset in the following manner: Obtain preset camera parameters, and based on each image in the second image set, determine the shooting pose information corresponding to each image in the second image set; Based on the camera parameters and the shooting pose information, a simulated shooting is performed to obtain a third dataset after the simulated shooting. Determine the target semantic label from the semantic labels in the third dataset; Based on the target semantic label, labeled data of target objects matching the target semantic label in the third dataset are generated to obtain the target dataset.

16. The apparatus according to claim 15, characterized in that, The processing unit performs simulated shooting based on the camera parameters and the shooting pose information in the following manner: Determine multiple shooting positions for each second image in the second image set; at each shooting position, based on multiple shooting orientations, determine the pixels of the target object corresponding to each shooting position to obtain multiple pixel images; Among the plurality of pixel images, the image whose number of pixels meets the pixel count requirement is determined as the target pixel image; The target shooting pose corresponding to the shooting orientation of the target pixel image is determined as the target shooting pose of the current shooting position, and the target shooting pose corresponding to each shooting position among the multiple shooting positions is obtained; Interpolation calculations are performed on the shooting poses of the multiple targets to obtain the shooting pose information of the second image; Based on the camera parameters and the shooting pose information, each second image is simulated for shooting.

17. An image recognition model training device, characterized in that, include: An acquisition unit is used to acquire a target dataset, wherein the training dataset is acquired using the data acquisition method described in any one of claims 1 to 7; The processing unit is used to train the image recognition model based on the target dataset to obtain the image recognition model.

18. An image recognition device, characterized in that, include: An acquisition unit is used to acquire an image to be identified, wherein the image to be identified includes an object to be identified; The processing unit is used to input the image to be identified into the image recognition model to obtain the target annotation data of the object to be identified; The image recognition model is trained based on a target dataset, which is trained using the data acquisition method described in any one of claims 1 to 7.

19. An electronic device, characterized in that, include: processor: Memory used to store processor-executable instructions; The processor is configured to: perform the data acquisition method according to any one of claims 1 to 7, or perform the image recognition model training method according to claim 8, or perform the image recognition method according to claim 9.

20. A storage medium, characterized in that, The storage medium stores instructions that, when executed by a processor, enable the processor to perform the data acquisition method of any one of claims 1 to 7, or the image recognition model training method of claim 8, or the image recognition method of claim 9.

21. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the data acquisition method as described in any one of claims 1 to 7, or performs the image recognition model training method as described in claim 8, or performs the image recognition method as described in claim 9.