Method, device and storage medium for constructing dataset for model training
By constructing a dataset of fine-grained 3D scene information and multi-step spatial reasoning information, the problem of multi-step spatial reasoning in complex tasks of existing visual language models is solved, and the spatial understanding and reasoning ability of the model is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE
- Filing Date
- 2025-06-04
- Publication Date
- 2026-04-28
AI Technical Summary
Existing visual language model datasets lack multi-step spatial relationships and complex spatial concepts, making it difficult to effectively support multi-step spatial reasoning and limiting their application in complex tasks.
By acquiring a set of natural content images of samples, we can identify fine-grained 3D scene information, generate multi-step spatial reasoning information, and construct a dataset that includes fine-grained 3D scene information and multi-step spatial reasoning information to guide the visual language model to reason step by step.
It improves the ability of visual language models in multi-step spatial reasoning tasks, supports multi-step reasoning for complex tasks, provides detailed 3D scene information, and enhances the model's spatial understanding and reasoning capabilities.
Smart Images

Figure CN121053479B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus and storage medium for constructing a dataset for model training. Background Technology
[0002] In the field of spatial reference reasoning within Visual Language Models (VLMs), the dataset is a core factor influencing the spatial reference reasoning capabilities of the trained VLM. Existing datasets, containing only limited spatial relationships and simple spatial concepts, make it difficult for VLMs trained on these datasets to effectively support multi-step spatial reasoning, thus limiting their application in complex tasks requiring multi-step reasoning. How to enable VLMs to achieve accurate reasoning in spatial reference tasks is a crucial issue that urgently needs to be addressed in the industry. Summary of the Invention
[0003] In view of the problems existing in the prior art, the present invention provides a method, apparatus and storage medium for constructing a dataset for model training.
[0004] This invention provides a method for constructing a dataset for model training, comprising:
[0005] A set of sample natural content images is acquired, and the sample natural content images in the set are identified to obtain fine-grained three-dimensional scene information of the sample natural content images.
[0006] A set template corresponding to the fine-grained 3D scene information is determined, and multi-step spatial reasoning information of the sample natural content image is generated based on the set template and the fine-grained 3D scene information.
[0007] The dataset is obtained based on the sample natural content images, the fine-grained 3D scene information of the sample natural content images, and the multi-step spatial reasoning information of the sample natural content images.
[0008] According to a method for constructing a dataset for model training provided by the present invention, the step of obtaining a sample natural content image set includes:
[0009] Obtain a first sample image dataset and perform vectorization processing on the first sample images in the first sample image dataset to obtain the embedding of the first sample images;
[0010] The similarity is determined based on the embedding of the first sample image and the embedding of each label in the set of labels;
[0011] The label with the highest similarity is obtained, and if the label with the highest similarity belongs to the positive label subset of the set label set, the first sample image after coarse screening is obtained;
[0012] A set of natural content images of the samples is obtained based on the first sample image after coarse screening.
[0013] According to a method for constructing a dataset for model training provided by the present invention, the step of obtaining a set of sample natural content images based on the first sample images after coarse screening further includes:
[0014] The first sample image after coarse screening is input into the large language model, and the large language model is guided by the first prompt word to output the discrimination result of the first sample image after coarse screening;
[0015] Based on the discrimination result, at least a portion of the coarsely screened first sample images are selected from the retained first sample images to obtain finely screened first sample images, so as to obtain a sample natural content image set based on the finely screened first sample images.
[0016] According to a method for constructing a dataset for model training provided by the present invention, the step of obtaining a sample natural content image set includes:
[0017] Obtain the second sample image dataset, and perform frame sampling processing on the sample video image data in the second sample image dataset to obtain the second sample image;
[0018] The sample natural content image set is obtained based on the second sample image.
[0019] According to a method for constructing a dataset for model training provided by the present invention, the method identifies sample natural content images in the sample natural content image set to obtain fine-grained three-dimensional scene information of the sample natural content images, including:
[0020] The sample natural content images in the sample natural content image set are identified to obtain category description information of objects in the sample natural content images and image-level point clouds of the sample natural content images;
[0021] The sample natural content images in the sample natural content image set are input into the visual detection model, and the category description information corresponding to the sample natural content images guides the visual detection model to output the bounding boxes of objects in the sample natural content images;
[0022] Based on the bounding boxes of objects in the sample natural content image, an instance mask of the objects in the sample natural content image is obtained, so as to obtain the object-level point cloud of the objects in the sample natural content image according to the instance mask and the image-level point cloud of the sample natural content image.
[0023] According to a method for constructing a dataset for model training provided by the present invention, the method identifies sample natural content images in the sample natural content image set to obtain category description information of objects in the sample natural content images, including:
[0024] If the sample natural content images in the sample natural content image set are identified and it is determined that the sample natural content images include at least two objects of the same category, the sample natural content images in the sample natural content image set are input into the category detection model, and the category detection model is guided by a second prompt word to output at least one initial object description information of the sample natural content images.
[0025] A set template corresponding to the sample natural content image is determined, and a hierarchical object description is generated by expanding the initial object description information based on the set template.
[0026] According to a method for constructing a dataset for model training provided by the present invention, the multi-step spatial reasoning information for generating the sample natural content image based on the set template and the fine-grained 3D scene information includes:
[0027] Based on the set template and the fine-grained 3D scene information, simple spatial reasoning information is generated for the sample natural content image; the simple spatial reasoning information includes at least one of preliminary question-answer pairs, multiple-choice questions, and factual statements; image description information of the sample natural content image is determined;
[0028] The simple spatial reasoning information is input into the reasoning big language model. Based on the image description information of the sample natural content image and the category description information of the objects in the sample natural content image, the reasoning big language model is guided to output the complex spatial reasoning information of the sample natural content image.
[0029] Based on the simple spatial reasoning information and complex spatial reasoning information of the sample natural content image, multi-step spatial reasoning information of the sample natural content image is obtained.
[0030] According to a method for constructing a dataset for model training provided by the present invention, the second sample image includes the original bounding boxes of objects in the second sample image;
[0031] Identify the sample natural content images in the sample natural content image set to obtain fine-grained three-dimensional scene information of the sample natural content images, including:
[0032] The sample natural content images in the sample natural content image set are identified to obtain category description information of objects in the sample natural content images;
[0033] The sample natural content images in the sample natural content image set are input into the visual detection model. The visual detection model is guided to output the predicted bounding boxes of the objects in the sample natural content images by the category description information of the objects in the sample natural content images.
[0034] The predicted bounding box and the original bounding box are matched to obtain a matching result, and the predicted bounding box that matches at least one of the original bounding boxes is retained based on the matching result.
[0035] The matching degree between the predicted bounding box and each of the original bounding boxes that match it is determined. Based on the matching degree, the original bounding box with the highest matching degree with the predicted bounding box is retained, so as to obtain fine-grained 3D scene information of the sample natural content image according to the original bounding box with the highest matching degree.
[0036] According to a method for constructing a dataset for model training provided by the present invention, the method identifies sample natural content images in the sample natural content image set to obtain fine-grained three-dimensional scene information of the sample natural content images, including:
[0037] The sample natural content images in the sample natural content image set are identified to obtain object-level point clouds and 3D bounding boxes of objects in the sample natural content images;
[0038] Based on the gravity alignment matrix of the second sample image dataset, the object-level point cloud and 3D bounding box of the object are transformed to obtain the gravity-aligned object-level point cloud and 3D bounding box.
[0039] Based on gravity-aligned object-level point clouds and 3D bounding boxes, free space for object placement is identified.
[0040] According to a method for constructing a dataset for model training provided by the present invention, the step of identifying free space for object placement based on gravity-aligned object-level point clouds and 3D bounding boxes includes:
[0041] When the volume of an object in the sample natural content image is within a set volume range, the free space below the object in the sample natural content image for object placement is identified based on the gravity-aligned object-level point cloud and 3D bounding box.
[0042] According to a method for constructing a dataset for model training provided by the present invention, the step of obtaining a sample natural content image set includes:
[0043] Generate a set number of indoor scenes; wherein, no objects exist on the target surface in the generated indoor scenes;
[0044] Obtain category description information of sample objects in the sample object dataset, and filter the sample object dataset based on the category description information to obtain coarsely screened sample objects;
[0045] Obtain object description information of sample objects in the sample object dataset, and filter the coarse-screened sample objects based on the object description information to obtain fine-screened sample objects;
[0046] Obtain the selection command input by the user, and filter the fine-screen sample object based on the selection command input by the user to obtain the target sample object;
[0047] The target sample pair is placed in the indoor scene and rendered to obtain a third sample image, and a sample natural content image set is obtained based on the third sample image.
[0048] According to the present invention, a method for constructing a dataset for model training is provided, wherein the setting template includes a simulation setting template that formalizes strong dependency spatial relationships;
[0049] Multi-step spatial reasoning information for generating the sample natural content image based on the set template and the fine-grained 3D scene information includes:
[0050] Based on the category description information and object description information of the sample object, generate the reference expression of the sample object;
[0051] Multi-step spatial reasoning information for generating the natural content image of the sample is generated based on the simulation setting template and the referential expression of the sample object.
[0052] The present invention also provides an apparatus for constructing a dataset for model training, comprising:
[0053] The scene information acquisition module is used to acquire a set of sample natural content images and to identify the sample natural content images in the set of sample natural content images in order to obtain fine-grained three-dimensional scene information of the sample natural content images.
[0054] The reasoning information generation module is used to determine the set template corresponding to the fine-grained three-dimensional scene information, and generate multi-step spatial reasoning information of the sample natural content image based on the set template and the fine-grained three-dimensional scene information;
[0055] The dataset acquisition module is used to obtain a dataset based on sample natural content images, fine-grained 3D scene information of the sample natural content images, and multi-step spatial reasoning information of the sample natural content images.
[0056] The present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement a method for constructing a dataset for model training as described in any of the preceding claims.
[0057] The present invention provides a computer storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements a method for constructing a dataset for model training as described in any of the preceding claims.
[0058] The present invention provides a method, apparatus, and storage medium for constructing a dataset for model training. The dataset is constructed based on fine-grained 3D scene information and multi-step spatial reasoning information of sample natural content images. It can be used to guide the visual language model to perform step-by-step reasoning during training and provides detailed 3D scene information required in the step-by-step reasoning process. This enables the trained visual language model to support multi-step spatial reasoning and to be applied to complex tasks that require multi-step reasoning. Attached Figure Description
[0059] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0060] Figure 1 This is a flowchart illustrating the method for constructing a dataset for model training provided by the present invention.
[0061] Figure 2 This is one of the schematic diagrams illustrating an example of the method for constructing a dataset for model training provided by the present invention.
[0062] Figure 3 This is the second example of a schematic diagram illustrating the method for constructing a dataset for model training provided by this invention.
[0063] Figure 4 This is the third example of a schematic diagram illustrating the method for constructing a dataset for model training provided by this invention.
[0064] Figure 5 This is the fourth example of a schematic diagram illustrating the method for constructing a dataset for model training provided by this invention.
[0065] Figure 6 This is the fifth example of a schematic diagram illustrating the method for constructing a dataset for model training provided by this invention.
[0066] Figure 7 This is the sixth example of a schematic diagram illustrating the method for constructing a dataset for model training provided by this invention.
[0067] Figure 8 This is the seventh example of a method for constructing a dataset for model training provided by the present invention.
[0068] Figure 9 This is the eighth example of a schematic diagram illustrating the method for constructing a dataset for model training provided by this invention.
[0069] Figure 10 This is the ninth example of a method for constructing a dataset for model training provided by the present invention.
[0070] Figure 11 This is the tenth example of a method for constructing a dataset for model training provided by the present invention.
[0071] Figure 12 This is illustrative diagram eleven of an example of the method for constructing a dataset for model training provided by the present invention.
[0072] Figure 13 This is a schematic diagram of the structure of the dataset construction device for model training provided by the present invention.
[0073] Figure 14 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0074] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0075] The following is combined with Figures 1-14 The present invention describes a method, apparatus, and storage medium for constructing a dataset for model training.
[0076] Figure 1 This is a flowchart illustrating the method for constructing a dataset for model training provided by the present invention, as shown below. Figure 1 As shown, the method includes the following steps:
[0077] Step 101: Obtain a set of sample natural content images and identify the sample natural content images in the set to obtain fine-grained three-dimensional scene information of the sample natural content images.
[0078] There are many ways to identify the sample natural content images in the sample natural content image set, such as model recognition, machine learning recognition, etc., and this invention does not limit this method.
[0079] Natural content images refer to images that include three-dimensional spatial information. These can be pseudo-3D scene images derived from 2D images, 3D video frames containing various 3D information, and image frames synthesized through simulation. Simulation-synthesized data can also be called analog data.
[0080] Fine-grained 3D scene information refers to the 3D scene information obtained after object-level recognition of sample natural content images. 3D scene information includes object labels, 2D bounding boxes, 3D bounding boxes, object-level point clouds, and mask instances in the sample natural content images.
[0081] Step 102: Determine the setting template corresponding to the fine-grained three-dimensional scene information, and generate multi-step spatial reasoning information for the sample natural content image based on the setting template and the fine-grained three-dimensional scene information.
[0082] Among them, the template setting refers to the question-and-answer template designed by combining the fine-grained three-dimensional scene information characteristics of the sample's natural content image.
[0083] The template styles include various approaches. For example, one template is based on the spatial ambiguity of pseudo-3D scene maps, focusing on qualitative spatial understanding and reasoning to generate question-and-answer objectives. It only includes spatial concepts such as relative positional relationships, relative size comparisons, and quantitative information from 2D image labels or pseudo-3D scene map labels. Relative positional relationships can be used to capture spatial layouts such as left / right, top / bottom, and front / back. Relative size comparisons can be used to describe object attributes inferred from the image plane projection, such as larger / smaller, taller / shorter, wider / narrower. Quantitative information from 2D image labels or pseudo-3D scene map labels may include spatial reasoning based on estimated depth maps, coordinates of objects in the 2D image, and coarse monocular depth approximations.
[0084] Specifically, the configuration template may include:
[0085] 1. Template evaluation of spatial and dimensional relationships:
[0086] Positional relationship: "Is [A] to the left of [B]?"
[0087] Size comparison: "Which object is larger, [A] or [B]?"
[0088] 2. Template query for the two-dimensional point of the unique identifier object '[A]':
[0089] "Where is [A] located? Please provide its two-dimensional coordinates." For example: "Where is the second apple from the left, the red apple on the left? Please provide its two-dimensional coordinates."
[0090] 3. The template queries the attributes at a specific 2D point '[X]', in the format '(x,y)':
[0091] Depth search: "What is the depth at point [X]?", for example: "What is the depth at point 0.528, 0.317?"
[0092] Object identification: "Which object is located at point [X]?", for example: "Which object is located at 0.753, 0.839?"
[0093] Multi-step spatial reasoning information refers to information that achieves step-by-step reasoning by combining multiple fine-grained three-dimensional scene information from sample natural content images. Multi-step spatial reasoning information can take many forms, such as initial question-and-answer pairs, multiple-choice questions, and factual statements, and this invention does not limit these forms.
[0094] Step 103: Obtain the dataset based on the sample natural content image, the fine-grained 3D scene information of the sample natural content image, and the multi-step spatial reasoning information of the sample natural content image.
[0095] In this dataset, a single sample includes a sample natural content image, fine-grained 3D scene information of the sample natural content image, and multi-step spatial reasoning information of the sample natural content image. The sample natural content image can correspond to multiple multi-step spatial reasoning information.
[0096] The method for constructing a dataset for model training provided in this invention constructs a dataset based on fine-grained 3D scene information and multi-step spatial reasoning information from sample natural content images. This dataset can be used to guide the visual language model in stepwise reasoning during training and provides detailed 3D scene information required in the stepwise reasoning process. This enables the trained visual language model to support multi-step spatial reasoning and to be applied to complex tasks that require multi-step reasoning.
[0097] Based on the above embodiments, a sample natural content image set is obtained, including:
[0098] Obtain the first sample image dataset and vectorize the first sample images in the first sample image dataset to obtain the embeddings of the first sample images;
[0099] The similarity is determined based on the embedding of the first sample image and the embedding of each label in the set of labels;
[0100] The label with the highest similarity is obtained, and if the label with the highest similarity belongs to the positive label subset of the set of labels, the first sample image after coarse screening is obtained;
[0101] A set of natural content images of samples is obtained based on the first sample image after coarse screening.
[0102] The first sample image dataset can be a two-dimensional image dataset or a two-dimensional network image dataset. There are many ways to obtain the first sample image dataset, such as through open-source datasets or manual photography, and this invention does not limit this method.
[0103] It is understandable that obtaining a set of sample natural content images based on a two-dimensional image dataset can enable the obtained dataset to endow the visual language model with basic spatial concepts, comprehensive depth perception of indoor and outdoor scenes, etc. during the training process.
[0104] A defined label set refers to a predefined set of labels associated with image features in the spatial semantic dimension. The defined label set includes positive labels and negative labels. In one embodiment, positive labels can be positive text labels representing desired image features; negative labels can be negative text labels representing undesired image features. These labels can serve as semantic anchors for visual-text alignment. Furthermore, the defined label set can be iteratively refined based on user input instructions to achieve manual iterative refinement, balancing recall and precision, ensuring relevance, and eliminating noise.
[0105] There are many ways to vectorize the first sample image in the first sample image dataset to obtain the embedding of the first sample image, such as extraction by model, extraction by image processing algorithm, etc., and this invention does not limit this.
[0106] The embedding of the first sample image is a numerical vector obtained by feature vectorization of the first sample image. Similarly, the embedding of the label is a numerical vector obtained by feature vectorization of the label. There are many ways to determine the similarity between the embedding of the first sample image and the embedding of each label, such as cosine similarity, Euclidean distance, etc., and this invention does not limit this method.
[0107] For example, the Siglip2 model can be used for coarse screening, also known as preliminary screening. Specifically, the cosine similarity between the embedding of the first sample image and the embedding of each label in the set of labels can be calculated using the siglip2-giant-opt-patch16-384 model.
[0108] For example, Positive Labels = ["Mid-distance observation of some objects on a table","Some objects on the desktop","Distant view of some animals","Mid-distance observation of some animals","Distant view of oneobject","Mid-distance observation of one object","Distant view of some objects","Mid-distance observation of some objects","Distant view of aperson","Mid-distance observation of a person","Distant view of somepeople","Mid-distance observation of some people","Distant view of indoorscene","Distant view of outdoor scene","Distant view of traffic","Distantview of Urban architecture"]
[0109] Negative Labels = ["Macro shot of an animal","Macro shot of one object","Macro shot of a person","Macro shot of flowers","A piece of text","A person displayed in front of a white background","A productdisplayed in front of a white background","A screenshot of the graphics userinterface","A dimly lit environment"]
[0110] Based on any of the above embodiments, obtaining a sample natural content image set based on the first sample image after coarse screening further includes:
[0111] The first sample image after coarse screening is input into the large language model, and the large language model is guided by the first prompt word to output the discrimination result of the first sample image after coarse screening.
[0112] Based on the discrimination results, at least a portion of the coarsely screened first sample images are selected from the retained first sample images to obtain finely screened first sample images, and a sample natural content image set is obtained based on the finely screened first sample images.
[0113] The large language model can be any existing model capable of recognizing images; for example, it can be the Qwen2.5-VL-7B model.
[0114] The first cue words are obtained through a structured cue engineering strategy. The first cue words may include system cue words and user cue words corresponding to each first sample image. System cue words define the role of the large language model as an image analysis expert, clarifying the key visual attributes to be evaluated and the negative categories to be detected, and enforcing a rigorous workflow. User cue words guide the large language model in determining whether the first sample image belongs to any predefined negative category.
[0115] If the discrimination result indicates that the first sample image does not belong to the predefined negative category, then the first sample image is retained. For example, the discrimination result "Yes" or "No" of the coarsely screened first sample image can be output via a pipeline. If it is "Yes", the first sample image is discarded; if it is "No", the first sample image is retained. Repeating this process for each coarsely screened first sample image yields the finely screened first sample image.
[0116] like Figure 2 As shown, in this embodiment, fine screening is performed on the basis of efficiently eliminating low-quality or irrelevant images such as those with irrelevant scenes or lacking multiple everyday objects through coarse screening. This can improve the overall screening efficiency while determining that the first sample image in the final sample natural content image set is clear, realistic, and suitable for spatial understanding and reasoning, especially for applications in spatial reference tasks.
[0117] Figure 2 The first row shows the initial sample images after coarse screening, successfully removing images lacking spatial semantics, such as close-ups of animals, people, or flowers, text content, and GUI screenshots. The second row shows the initial sample images after fine screening, further removing unsuitable categories, including artwork, dimly lit scenes, black and white images, geometrically distorted content, and image collages. The third row shows the final retained initial sample images, which possess rich spatial relationships. This verifies that the initial sample images after coarse and fine screening are suitable for spatial understanding and reasoning, and demonstrates the overall quality of the final dataset.
[0118] In one embodiment, 934,000 images are obtained from 1.7 million images through SigLIP2 coarse screening, and 846,000 images are obtained from 934,000 images through Qwen2.5-VL fine screening.
[0119] Based on any of the above embodiments, the natural content images in the sample natural content image set are identified to obtain fine-grained three-dimensional scene information of the sample natural content images, including:
[0120] The sample natural content images in the sample natural content image set are identified to obtain the category description information of objects in the sample natural content images and the image-level point cloud of the sample natural content images;
[0121] Input the sample natural content images from the sample natural content image set into the visual detection model, and guide the bounding boxes of objects in the sample natural content images output by the visual detection model through the category description information corresponding to the sample natural content images.
[0122] The instance mask of the object in the sample natural content image is obtained based on the bounding box of the object in the sample natural content image, so as to obtain the object-level point cloud of the object in the sample natural content image according to the instance mask and the image-level point cloud of the sample natural content image.
[0123] There are many ways to identify the sample natural content images in the sample natural content image set, such as model recognition, algorithm recognition, etc., and this invention does not limit this.
[0124] Here, category description information refers to the semantic category labels identified for objects in the sample natural content image. For example,
[0125] In one embodiment, the natural content images in the sample natural content image set are identified to obtain category description information of objects in the sample natural content images, including:
[0126] The sample natural content images in the sample natural content image set are identified. If it is determined that the sample natural content images do not contain at least two objects of the same category, the sample natural content images in the sample natural content image set are input into the visual perception big model to obtain the category description information output by the visual perception big model.
[0127] Among them, the visual perception big model can be the Recognize Anything Model (RAM).
[0128] Understandably, by using the category description information of objects in the natural content images of the output sample from the trained visual perception model, the extensive recognition capabilities of the trained visual perception model can ensure comprehensive semantic coverage, thereby providing guidance for subsequent localization.
[0129] In another embodiment, the sample natural content images in the sample natural content image set are identified to obtain category description information of objects in the sample natural content images, including:
[0130] If the sample natural content images in the sample natural content image set are identified and it is determined that there are at least two objects of the same category in the sample natural content images, the sample natural content images in the sample natural content image set are input into the category detection model, and the category detection model is guided by the second prompt word to output at least one initial object description information of the sample natural content images.
[0131] Determine the template corresponding to the sample natural content image, and generate a hierarchical object description by expanding the initial object description information based on the template.
[0132] The sample natural content image can be input into the Qwen2.5-VL-7B model, and the Qwen2.5-VL-7B model can be guided to output at least one initial object description information of the sample natural content image by the second prompt word corresponding to the object description information.
[0133] This process involves determining the spatial variations of at least two objects of the same category in the sample natural content image along three principal axes. The axis with the greatest spatial variation is identified as the primary sorting axis. Based on the primary sorting axis, templates are retrieved from a predefined template library to expand the initial object description information. The principal axes refer to the front-back axis, the left-right axis, and the top-bottom axis.
[0134] For example, for samples of natural content images where at least two objects of the same category are chairs, and the left and right axes are the primary ranking axes, the template might be: "{dense_caption}, this is the {ordinal}th {class_name} from left to right", or "{dense_caption}, this is the {ordinal}th {class_name} in the left-to-right sequence". Here, dense_caption represents the object description information generated by the Qwen2.5-VL model, ordinal represents the object's position in the ranking sequence, and class_name is the class description information predicted by RAM.
[0135] The accuracy of the category detection model is greater than that of the well-trained visual perception model. In this embodiment, different models are used to obtain category description information based on whether the sample natural content image includes multiple objects of the same type. Each model can play its own advantages at different stages, and the overall efficiency can be improved through staged division of labor.
[0136] In one embodiment, before inputting the sample natural content images from the sample natural content image set into the category detection model, the method further includes:
[0137] Determine the variance of at least two objects of the same category in the sample natural content image on three principal axes. If the variance on each principal axis is lower than a set variance threshold, discard the sample natural content image.
[0138] The specific value of the variance threshold can be set according to actual needs, and this embodiment does not limit it.
[0139] In this embodiment, by discarding some sample natural content images through variance thresholding, ambiguity in spatial reference tasks can be avoided, ensuring the clarity and effectiveness of the image description information and at least one initial object description information of the obtained sample natural content images.
[0140] Image-level point clouds refer to the set of three-dimensional points extracted and reconstructed from a sample natural content image. Image-level point clouds are used to describe the three-dimensional spatial information of the entire sample natural content image.
[0141] For example, the UniDepth V2 model can be used to perform metric depth estimation on the sample natural content image to extract its 3D perceptual information. The sample natural content image can be input into a WildeCamera to obtain the camera intrinsics of the sample natural content image output by the WildeCamera. The 3D perceptual information and camera intrinsics of the sample natural content image are then combined to reconstruct the image-level point cloud of the sample natural content image.
[0142] Compared to the traditional method of sharing image and depth coding, this embodiment improves the depth perception of objects in complex environments such as indoor scenes by using independent image and depth coding, avoids interference between depth information and image information, and improves the understanding accuracy of processing various types of spatial information.
[0143] The GroundingDINO model can input sample natural content images and use the category descriptions corresponding to these images as prompts (also known as text prompts) to guide the output of bounding boxes for objects in the sample natural content images. Here, "bounding box" refers to a two-dimensional bounding box.
[0144] After obtaining the bounding boxes of objects in the sample natural content image, these bounding boxes can be input into the instance segmentation model to obtain the instance mask of the objects in the sample natural content image, which is output by the instance segmentation model. The instance segmentation model can be SAM (Segment Anything Model).
[0145] For example, by combining the instance mask of the object in the sample natural content image and the image-level point cloud of the sample natural content image, the object-level point cloud of the sample natural content image can be obtained. By combining the bounding box of the object in the sample natural content image, the instance mask, and the object-level point cloud of the sample natural content image, an axis-aligned 3D bounding box can be obtained.
[0146] In this way, a pseudo-3D scene map of the sample natural content image can be generated based on the bounding boxes, instance masks, object-level point clouds, and 3D bounding boxes of objects in the sample natural content image. (See also...) Figure 3 The image shown is a pseudo-3D scene visualization containing sample natural content images, objects in the sample natural content images, and corresponding point clouds.
[0147] In this embodiment, by converting the sample natural content image into a pseudo-3D scene graph, using nodes in the pseudo-3D scene graph to represent object attributes, and using edges in the pseudo-3D scene graph to represent spatial relationships between objects, the amount of 3D spatial information provided by the natural content image can be increased.
[0148] Furthermore, in real-world scenarios, each category typically contains multiple instances, such as multiple tables in a classroom. In this embodiment, by refining the object descriptions in the sample natural content images to generate hierarchical object descriptions, it is possible to achieve a more refined distinction between at least two objects of the same category in the sample natural content images.
[0149] Based on any of the above embodiments, multi-step spatial reasoning information for generating sample natural content images based on a set template and fine-grained 3D scene information includes:
[0150] Simple spatial reasoning information is generated based on a set template and fine-grained 3D scene information to produce sample natural content images; the simple spatial reasoning information includes at least one of preliminary question-and-answer pairs, multiple-choice questions, and factual statements; and image description information is determined for the sample natural content images.
[0151] Simple spatial reasoning information is input into the reasoning big language model. Based on the image description information of the sample natural content image and the category description information of the object in the sample natural content image, the reasoning big language model is guided to output complex spatial reasoning information of the sample natural content image.
[0152] Based on the simple and complex spatial reasoning information of the sample natural content image, multi-step spatial reasoning information of the sample natural content image is obtained.
[0153] This includes the ability to input fine-grained 3D scene information into question-and-answer templates to generate structured question-and-answer pairs for sample natural content images; the ability to input fine-grained 3D scene information into fact templates to generate statements for sample natural content images; and the ability to input fine-grained 3D scene information into multiple-choice templates to generate multiple-choice questions for sample natural content images.
[0154] For example, the fact-setting template includes:
[0155] 1. Approximate Depth: “Distance between point [A] and camera [X]” is based on depth estimation.
[0156] 2. Precise two-dimensional object position: "[A] is located at point [X]", "[A] is to the right of [B]".
[0157] The sample natural content images from the sample natural content image set can be input into the category detection model, and the category detection model can be guided by a third prompt word to output the image description information of the sample natural content images.
[0158] Specifically, the sample natural content image can be input into the Qwen2.5-VL-7B model, and the Qwen2.5-VL-7B model can output the image description information of the sample natural content image through the third prompt word corresponding to the image description information.
[0159] The reasoning language model can use QwQ-32B, which can obtain the fourth prompt word of QwQ-32B based on the image description information of the sample natural content image and the category description information of the object in the sample natural content image.
[0160] In this embodiment, factual statements, initial question-and-answer pairs, and multiple-choice questions are input into the reasoning language model. Based on the global image description and multiple object descriptions of the sample natural content images, prompt words are generated to guide the reasoning language model to output more challenging, conversational, and diverse complex spatial reasoning information, which can transcend the limitations of simple spatial reasoning information.
[0161] Among them, the global image description of the natural content images of the samples can provide a key contextual basis for the reasoning big language model, and improve the relevance and accuracy of the complex spatial reasoning information output by the reasoning big language model.
[0162] In this embodiment, acquiring simple spatial reasoning information makes it easier for the visual language model to understand and utilize it during training. Acquiring complex spatial reasoning information is like a multi-step logical chain, providing questions that guide the visual language model to think in multiple steps to arrive at the answer during training. This can comprehensively cover the range from basic recognition to advanced reasoning, thereby significantly improving the overall ability of the trained visual language model in spatial reference reasoning tasks.
[0163] Based on any of the above embodiments, a sample natural content image set is obtained, including:
[0164] Obtain the second sample image dataset, and perform frame sampling processing on the sample video image data in the second sample image dataset to obtain the second sample image;
[0165] A set of sample natural content images is obtained based on the second sample image.
[0166] The second sample image dataset can be a large-scale 3D video dataset, such as the open-source dataset CA-1M.
[0167] There are many ways to perform frame sampling processing on the sample video image data in the second sample image dataset. For example, a frame can be selected from the 3D video data of the second sample image dataset at a set interval to obtain sample video image data, or a frame number can be selected from the 3D video data of the second sample image dataset at a set interval to obtain sample video image data. This invention does not limit this method. The specific set interval duration and frame number can be set by those skilled in the art according to actual needs, and this invention does not limit this method either. For example, a frame can be selected every 20 frames to obtain sample video image data.
[0168] 3D video datasets typically capture continuous activity. Therefore, the differences between consecutive frames are usually very subtle, resulting in nearly identical visual content and giving the data in 3D video datasets a high degree of temporal redundancy. For example... Figure 4 As shown, the first row displays four frames of images under continuous sampling, with very small differences between frames. The second row displays four frames of images under frame sampling processing, with significantly increased differences between frames. In this embodiment, frame sampling processing can retain frames with obvious differences, retain meaningful scene and perspective transitions, maintain scene diversity, provide effective information gain for the training of visual language models, and reduce redundancy.
[0169] The second sample image includes the original bounding boxes of the objects in the second sample image. The sample natural content images of the sample natural content image set obtained based on the second sample image also include the original bounding boxes of the objects.
[0170] In the case where the second sample image dataset is obtained based on the open-source dataset CA-1M, the original bounding boxes can also be called CA-1M bounding boxes, which are bounding boxes that come with the open-source data.
[0171] Currently, the second sample images in the second sample image set of the 3D video data type are labeled with object bounding boxes. However, it is difficult to accurately understand the space using only bounding boxes. To solve this problem, based on any of the above embodiments, the sample natural content images in the sample natural content image set are identified to obtain fine-grained 3D scene information of the sample natural content images, including:
[0172] The sample natural content images in the sample natural content image set are identified to obtain category description information of objects in the sample natural content images;
[0173] Input the sample natural content images from the sample natural content image set into the visual detection model, and guide the visual detection model to output the predicted bounding boxes of the objects in the sample natural content images through the category description information of the objects in the sample natural content images;
[0174] The predicted bounding box and the original bounding box are matched to obtain the matching result. Based on the matching result, the predicted bounding box that matches at least one original bounding box is retained.
[0175] The matching degree between the predicted bounding box and each matching original bounding box is determined. Based on the matching degree, the original bounding box with the highest matching degree with the predicted bounding box is retained to obtain fine-grained 3D scene information of the sample natural content image from the original bounding box with the highest matching degree.
[0176] The working principle and technical effect of obtaining the category description information and two-dimensional bounding box of the object in the sample natural content image obtained based on the second sample image are basically the same as those of obtaining the category description information and two-dimensional bounding box of the object in the sample natural content image obtained based on the first sample image, and will not be repeated here.
[0177] In one embodiment, to improve recall, the confidence thresholds of the category detection model that generates category description information and the visual detection model that generates two-dimensional bounding boxes can be lowered to retain low-confidence but potentially relevant candidate objects and avoid missing possible objects.
[0178] like Figure 5 As shown, bounding boxes can be matched with the predicted bounding boxes based on the Intersection over Union (IoU). Due to the sparsity, occlusion, and fragmentation of the annotations, multiple original bounding boxes can be obtained that match the predicted bounding boxes. The one with the highest IoU value among these original bounding boxes can be uniquely matched with the predicted bounding box, eliminating redundant or poorly aligned original bounding boxes, resulting in an optimized set of original bounding boxes with strong semantic alignment.
[0179] Figure 5 In the image, the first row shows the original bounding boxes that were successfully matched, which usually correspond to well-labeled objects and include category descriptions; the second row shows the original bounding boxes that were not matched, which usually correspond to objects with ambiguous labels.
[0180] Based on any of the above embodiments, the natural content images in the sample natural content image set are identified to obtain fine-grained three-dimensional scene information of the sample natural content images, including:
[0181] The sample natural content images in the sample natural content image set are identified to obtain the object-level point cloud and 3D bounding box of the objects in the sample natural content images;
[0182] Based on the gravity alignment matrix of the second sample image dataset, the object-level point cloud and 3D bounding box of the object are transformed to obtain the gravity-aligned object-level point cloud and 3D bounding box.
[0183] Based on gravity-aligned object-level point clouds and 3D bounding boxes, free space for object placement is identified.
[0184] The working principle and technical effect of obtaining the object-level point cloud and 3D bounding box of the object in the sample natural content image based on the second sample image set are basically the same as those of obtaining the object-level point cloud and 3D bounding box of the object in the sample natural content image based on the first sample image set. Therefore, this embodiment does not limit the specific method used.
[0185] The gravity alignment matrix is carried by the second set of sample images, and could be, for example, a gravity alignment matrix from CA-1M. Figure 6 As shown, the gravity alignment matrix can transform the object-level point cloud and 3D bounding box of an object into a coordinate system where gravity is always vertically downward, resulting in gravity-aligned object-level point cloud and 3D bounding box.
[0186] In this embodiment, gravity alignment can project the scene in the sample natural content image onto a plane perpendicular to the gravity vector, thereby more clearly showing the layout and spatial relationships of objects in the sample natural content image.
[0187] like Figure 7 As shown, in one embodiment, before recognizing the sample natural content images in the sample natural content image set, the method further includes: inputting the sample natural content images into a large language model, guiding the large language model to initialize through system prompts, and guiding the large language model to output a judgment result on whether the sample natural content images have a planar candidate image layer through user prompts. Scenes such as walls or ceilings that do not have a planar candidate image layer can be excluded based on the judgment result.
[0188] Among them, the large language model can be Qwen2.5-VL.
[0189] In one embodiment, after obtaining the gravity-aligned object-level point cloud and 3D bounding box, the method further includes: determining candidate platforms for objects in the sample natural content image; determining supporting platforms from the candidate platforms based on the distance between the top surface of the candidate platform and the bottom surface of the object in the sample natural content image and the overlapping area in the clothing direction; and determining relationships such as "front / back / left / right" between objects on the same supporting platform. These relationships can be defined as distances within 0.05 meters and overlapping areas of not less than 70%.
[0190] When determining the "below" relationship, if the object is suspended, the candidate platform directly below the object is selected as the support platform. The plane closest to the bottom of the object is considered the reference plane, and the area of its intersection with the object from top to bottom must exceed 70% of the object's area. When determining the "above" relationship, the top surface of the object is used as the reference plane.
[0191] When determining the relationship between two objects, the supporting plane of each object is determined independently using the same steps as when determining the "front / back / left / right" relationship. Only when the supporting planes of the two objects are the same is the space between them considered valid, thus ensuring that spatial reasoning is carried out in a unified physical environment.
[0192] Based on any of the foregoing embodiments, the identification of free space for object placement based on the gravity-aligned object-level point cloud and 3D bounding box includes:
[0193] Given a set volume range for objects in sample natural content images, the free space below the objects in the sample natural content images is identified based on the gravity-aligned object-level point cloud and 3D bounding box for object placement.
[0194] The defined volume range is the range of volumes that are relatively large and may be hollow inside.
[0195] There are many ways to determine whether the volume of an object in a sample natural content image is within a set volume range. For example, one can determine whether the volume of the object in the sample natural content image is within a set absolute volume value range, or one can determine whether the volume of the object in the sample natural content image is within a multiple range relative to the target volume, etc. This invention does not limit this. For example, the set volume range can be 4.236 times the target object volume range.
[0196] When selecting visible points using depth matching, solid objects are typically filtered by bounding box volume. However, this can lead to larger, potentially hollow objects being incorrectly classified as completely occluded, thus misleading placement reasoning on the support platform. Figure 8 As shown, in this embodiment, by introducing volume constraints, large tables with volumes exceeding the threshold can be filtered out to obtain objects of small to medium size. The bounding box volume can accurately reflect the actual occupancy of the object and does not affect the placement reasoning on the support platform.
[0197] Furthermore, for relationships such as "front / back / left / right / between", after identifying the target object and its supporting platform, candidate objects on the supporting platform can be determined based on the following criteria:
[0198] 1. Their bottoms must not be significantly higher than the top of the target object.
[0199] 2. Their heads must remain above the platform surface.
[0200] 3. Their projections on the XZ plane must intersect with the projection of the target object.
[0201] 4. Their volume must not exceed 4.236 times the volume of the target object.
[0202] For the relationship of "below" or "below", candidate objects below the target object can be determined based on the following criteria:
[0203] 1. Its projection onto the XZ plane intersects with the projection of the target object;
[0204] 2. Its bottom must not be higher than the top of the target object;
[0205] 3. Its top is not lower than the top surface of its supporting platform.
[0206] For the relationship of "above" or "above", candidate objects above the target object can be determined based on the following criteria:
[0207] 1. Its bottom should be no more than 20 centimeters from the top surface of the target object;
[0208] 2. Its top surface must not be lower than the top surface of the target object;
[0209] 3. Its projection onto the XZ plane overlaps with the projection of the target object.
[0210] In one embodiment, after determining the supporting platform of the target object and adjacent objects, the areas in front of, behind, to the left, to the right, above, or below the target object can be determined. For example, a 90° sector-shaped area can be defined centered on the target object and oriented in the corresponding direction. The radius of the sector-shaped area can be set as the larger of the object's diagonal length and a fixed 20 cm, thus ensuring sufficient coverage. Figure 8 As shown, Figure 8The target is the object in the middle, the platform is the supporting platform for the target, and the empty point on the right is the fan-shaped area between the target and other obstacles.
[0211] For the "up / down" direction, the top or bottom surface of an object can be projected onto the support platform, and the projection can be scaled down to 80% of its original size with the center as the reference. This reduces overestimation caused by the coarse 3D bounding box, thereby reducing overlap with nearby objects and better approximating the available space. Figure 9 As shown, the target is the target object, the platform is the supporting platform for the target object, and the empty points at the bottom are the scaled-down projection areas.
[0212] Free space can be calculated by analyzing the object outlines in the gravity-aligned top view, and a minimum free area constraint can be enforced in all spatial scenarios "above / below / between": the unoccupied area in the top-view direction must exceed 0.036 meters to filter out small spaces and ensure that the unoccupied area can accommodate objects such as books or cups. Figure 10 As shown, Target 1 is Target Object 1, Target 2 is Target Object 2, the platform is the common support platform for Target Object 1 and Target Object 2, and the intermediate idle point is the intermediate space area obtained from Target Object 1 and Target Object 2.
[0213] like Figure 8-10 In the diagram, the blue shaded area represents the final sampled area after considering bounding box scaling and occlusion, highlighting the feasible unoccupied area for subsequent placement analysis. The areas are arranged from left to right. Figure 8 The first part is the top view map of the search area on the right, the second part is the sampling points projected onto the two-dimensional image plane, and the third part is the final visible point. Figure 9 The first part is the top-down view of the object's bottom surface, the second part is the sampling points projected onto the two-dimensional image plane, and the third part is its final visible points; Figure 10 The first part is a top-view occupancy map with a central search area, the second part is the sampled points projected onto the two-dimensional image plane, and the third part is the final visible point.
[0214] In one embodiment, for view from above Figure X Candidate points sampled in the Z-plane can be assigned a y-coordinate to the top surface of the platform using the camera's intrinsic and extrinsic parameters and gravity alignment, and then projected onto the original 2D image.
[0215] The Z coordinate of each point is compared with the corresponding depth value in the aligned depth image. If the difference exceeds 2.5 cm, the point is considered occluded and discarded.
[0216] For the front, back, left, and right directions, 9000 points are sampled for each direction. If at least 2000 points are visible, the direction is retained. For the top, bottom, and between directions, 10000 points are sampled. If at least 6000 points are visible, the direction is retained.
[0217] The average position of the remaining visible points can be calculated to obtain a representative target location. If the depth deviation from the depth image exceeds 2.5 cm, the nearest point within that threshold is selected. Figure 8-10 The blue circle shown.
[0218] Based on any of the above embodiments, the working principle and technical effect of generating multi-step spatial reasoning information of sample natural content images based on the first sample image set based on the set template and fine-grained three-dimensional scene information are basically the same as the working principle and technical effect of generating multi-step spatial reasoning information of sample natural content images based on the second sample image set based on the set template and fine-grained three-dimensional scene information, and will not be repeated here.
[0219] The difference lies in the fact that the second sample image set can provide richer and more accurate fine-grained 3D scene information, such as depth maps, camera poses, and 3D bounding boxes for each object, enabling the construction of more complex multi-step spatial reasoning information that includes spatial references and inferences.
[0220] In addition, the second sample image set can provide a world coordinate system that reflects gravity and a camera coordinate system that reflects the vertical direction of the image. Based on the sample natural content image obtained from the second sample image set, the set template corresponding to the fine-grained 3D scene information can include relative positional relationships, direction and rotation reasoning, geometric attribute comparison, quantitative spatial reasoning, free space reasoning, spatial reference position, and spatial reference layout prediction.
[0221] Relative positional relationships can be used to capture spatial relationships such as left / right, top / bottom, front / back, inside / outside, contact / separation, and near / far. Orientation and rotation inference can infer orientation and viewpoint changes using 3D objects or camera intrinsics such as orientation vectors and rotation matrices. Geometric attribute comparison is used to compare attributes such as size (large / small), height (high / low), and width (wide / narrow) based on the object's true 3D dimensions to reduce distortion caused by 2D projection. Quantitative spatial inference is used to calculate depth, distance, relative angles, and spatial intermolecular properties using precise 3D coordinates and measurements. Free space inference is used to identify free space above, below, or between objects. Spatial reference position and spatial reference layout prediction are used to predict precise 2D coordinates from verbal descriptions, such as "pointing to the second chair from the left"—identifying the target object, or "indicating an empty space to the right of the white box on the second shelf"—selecting an effective placement location. Spatial reference position and spatial reference layout prediction enable precise 2D-to-3D projection and refined spatial understanding, forming an important bridge between visual perception and physical interaction and execution.
[0222] Based on any of the above embodiments, a sample natural content image set is obtained, including:
[0223] Generate a set number of indoor scenes; wherein, no objects exist on the target surface in the generated indoor scenes;
[0224] Obtain the category description information of the sample objects in the sample object dataset, and filter the sample object dataset based on the category description information to obtain coarsely screened sample objects;
[0225] Obtain object description information of sample objects in the sample object dataset, and filter the coarse-screened sample objects based on the object description information to obtain fine-screened sample objects;
[0226] Obtain the user's input selection instructions, and filter the fine-screen sample objects based on the user's input selection instructions to obtain the target sample objects;
[0227] The target sample pair is placed in an indoor scene and rendered to obtain a third sample image. The sample natural content image set is then obtained based on the third sample image.
[0228] The specific value of the set quantity can be set according to actual needs, and this invention does not limit it. For example, more than 3,000 unique indoor scenes can be generated.
[0229] There are many ways to generate a set number of indoor scenes, such as through a 3D generator or model drawing, and this invention does not limit this method. When generating a set number of indoor scenes using a 3D generator, the 3D generator can be Infinigen, and compose_indoors.solve_steps_small can be set to 0 to reduce the probability of the generated indoor scenes containing tiny objects, thus reserving space for the subsequent placement of target sample objects.
[0230] like Figure 11 As shown, in one embodiment, the generated indoor scenes can be filtered based on the following set criteria:
[0231] • Sufficient desktop area: The selected scene must include at least one sufficiently large and continuous desktop surface, such as a desk, dining table, or counter, suitable for placing items. Scenes without a desktop or with a desktop that is too small to be practical will be excluded.
[0232] • Acceptable lighting conditions: Scenes with extreme lighting problems, such as being too dark, overexposed, or having unnatural tones, will be discarded to ensure that subsequent lighting adjustments have a feasible basis.
[0233] • Realism and coherence of the scene: Remove scenes with serious geometric inconsistencies or unreasonable layouts to maintain physical plausibility.
[0234] • Camera Accessibility: The scene must allow for proper camera placement and ensure a clear view of the target surface. Highly cluttered or confined environments have lower priority.
[0235] After obtaining the selected indoor scenes, the scenes can be automatically modified to enhance scene diversity and control experimental variables:
[0236] • Lighting randomization: The intensity of the light source, such as ceiling lights or table lamps, is uniformly scaled within the range of [0.6I, 1.4I], where I represents the original intensity.
[0237] • Camera pose adjustment: For each desktop, the camera angle is defined as a pitch angle randomly selected from the range of [−60°, −30°], relative to the desktop plane, facing the center of the area.
[0238] • Camera height and distance variations: The camera is uniformly sampled at a height between 0.3 and 0.8 meters above the tabletop. The distance between the camera and the target area is adjusted according to the surface size and field of view to ensure the target area is fully visible.
[0239] The sample object dataset is a dataset containing textual descriptions of three-dimensional objects. It can be obtained from open-source datasets or artificially constructed; this invention does not limit this. For example, the sample object dataset can be obtained from the Objaverse LVIS dataset, which includes object descriptions and corresponding category information.
[0240] The sample object dataset can be coarsely screened based on the following set criteria and category description information of the sample objects:
[0241] • It can usually be placed on a flat surface.
[0242] • Maximum size not exceeding 1 meter, suitable for desktop use.
[0243] Fine-tuning can be performed based on the attributes of the reference dataset and the object description information of the sample objects:
[0244] • Axis alignment: Key features, such as edges and handles, are aligned with the standard camera coordinate axes.
[0245] • Single object: refers to a single, independent object, rather than a scene or a collection of objects.
[0246] • Color diversity: Includes colors other than white or gray.
[0247] • No ground plane: Does not include a ground plane for visualization.
[0248] • High quality: Clean geometry, well-constructed, and flawless.
[0249] • Distinguishing perspectives: Standard perspectives of the front, back, top, bottom, left, and right present meaningful visual or semantic differences.
[0250] • A reasonable object: represents a common, recognizable object, rather than an abstract shape or an unrecognizable entity.
[0251] The reference dataset can be the OrienText300K dataset.
[0252] Since object size estimation relies on category description information and object description information, with a tolerance of ±30%, further manual review can be performed based on the aforementioned seven established rules to reduce the impact of the accuracy of the sample object's object description information on the recognition results. During this process, irregular geometric shapes such as tangled wires from wired mice, and objects that cause bounding box deformation and thus affect reliable scaling can be removed to obtain the target objects. Some target objects, such as... Figure 12 As shown. In one embodiment, after manual review, approximately 3,000 target objects that meet the requirements can be retained from approximately 9,000 finely screened objects.
[0253] The target object is a 3D asset, or 3D resource, intended for placement in an indoor scene.
[0254] In one embodiment, obtaining object description information of sample objects in the sample object dataset includes:
[0255] The large language model is guided to output the structured text attributes of each sample in the reference dataset based on data such as orientation and description in the reference dataset.
[0256] Based on the structured text attributes of the samples and the original object description information of the sample objects in the sample object dataset, obtain the object description information of the sample objects in the sample object dataset.
[0257] Structured text attributes may include:
[0258] • Location description: Prepositional phrases that indicate the standard front, prominent part, or inherent orientation of an object, such as "on the front of" or "on the handle side of," suitable for insertion into sentence templates.
[0259] • Color label: A single-word description of the object's primary color. If the object has multiple prominent colors, this attribute is labeled "None", for example, "Blue", "None".
[0260] • Object tag: A concise noun phrase used to specify the category of the object, such as "coffee cup" or "computer mouse", which can be used as the subject or object in the template.
[0261] • Category Consistency: A Boolean flag indicating whether an object's visual category is consistent with its textual description.
[0262] Among them, the large language model can be GPT-4o.
[0263] Based on data such as direction and description in the reference dataset, prompt words can be generated. These prompt words guide the large language model to output the structured text attributes of each sample in the reference dataset. The specific prompt words can be determined by the large language model according to actual needs; this invention does not impose any limitations on this.
[0264] The original object description information of a sample object is the object description information that is inherent in the sample object dataset. The original object description information of the sample objects in the sample object dataset can be supplemented and / or replaced based on the structured text attributes of the sample to obtain the new object description information.
[0265] There are many ways to place the target sample pair into an indoor scene and render it to obtain a third sample image, such as through programming or user commands. This invention does not limit the methods used.
[0266] When placing target sample pairs into an indoor scene rendering to obtain a third sample image through programmatic methods, the following placement strategies can be used:
[0267] - In many scenarios, increase the scale of objects with directional vectors in the XY plane, such as laptops, teddy bears, and mugs.
[0268] -Increase the co-occurrence of objects from the same category but with significantly different characteristics, such as a ceramic mug with a handle and a paper Coke cup.
[0269] - Ensure a reasonable physical layout, for example, avoid excessive interweaving and place objects upright.
[0270] The number of assets in each scenario ranges from 3 to 9.
[0271] Indoor scenes can be rendered based on the following rendering rules:
[0272] • Renderer: Use Blender Cycles to render high-fidelity, physically based images.
[0273] • Image resolution: The output image is rendered at 960×540 pixels.
[0274] • Rendering quality: Set the configure_render_cycles.num_samples parameter to 2048 to achieve high-quality rendering while keeping noise within a reasonable range.
[0275] Based on any of the above embodiments, the template setting includes a simulation template that formalizes strongly dependent spatial relationships.
[0276] For example, the simulation specification template may include the following formalized relationships:
[0277] • Position: The location of one object relative to another object, for example, "to the left of the green bottle".
[0278] • Orientation: Questions involving the inherent orientation of an object, such as "on the handle side of the red mug".
[0279] • Distance query: Precise distance, for example, "0.2 meters to the left of the plate".
[0280] • Between two: Identify an object located between two other objects, such as "between the stapler and the telephone".
[0281] • Specific surface location: Positioning an object relative to a certain part of a surface, for example, "at the bottom left corner of the table".
[0282] Multi-step spatial reasoning information is generated based on a set template and fine-grained 3D scene information to produce sample natural content images, including:
[0283] Based on the category description information and object description information of the sample objects, generate the reference expression of the sample objects;
[0284] Multi-step spatial reasoning information is generated based on simulation-defined templates and the referential expressions of sample objects to produce sample natural content images.
[0285] Among them, the category description information and object description information of the sample object can be combined to generate an explicit referential expression for the sample object, for example:
[0286] • Feature category: The semantic category of the object, such as "muzzle" or "laptop".
[0287] • Color: The primary color of an object, for example, "red mug".
[0288] • Left-right order: The ordinal position from left to right, for example, "the third bottle from the left".
[0289] • Row order: The ordinal position from front to back, for example, "the last LEGO minifigure".
[0290] • Distance ranking from a reference point: The ordinal position is determined based on the distance from a prominent reference point, such as "the plate closest to the blue mug".
[0291] • Height ranking: Ordinal position determined by height, for example, "the second tallest teddy bear".
[0292] The apparatus for constructing a dataset for model training provided by the present invention will be described below. The apparatus for constructing a dataset for model training described below can be referred to in correspondence with the method for constructing a dataset for model training described above.
[0293] Figure 13 This is a schematic diagram of the structure of the simulation visualization device for subway train operation diagrams provided by the present invention, as shown below. Figure 13 As shown, the device includes:
[0294] The scene information acquisition module 1301 is used to acquire a sample natural content image set and to identify the sample natural content images in the sample natural content image set in order to obtain fine-grained three-dimensional scene information of the sample natural content images.
[0295] The reasoning information generation module 1302 is used to determine the set template corresponding to the fine-grained three-dimensional scene information, and generate multi-step spatial reasoning information of the sample natural content image based on the set template and the fine-grained three-dimensional scene information;
[0296] The dataset acquisition module 1303 is used to obtain a dataset based on the sample natural content image, the fine-grained three-dimensional scene information of the sample natural content image, and the multi-step spatial reasoning information of the sample natural content image.
[0297] Based on any of the above embodiments, the scene information acquisition module 1301 is used for:
[0298] Obtain a first sample image dataset and perform vectorization processing on the first sample images in the first sample image dataset to obtain the embedding of the first sample images;
[0299] The similarity is determined based on the embedding of the first sample image and the embedding of each label in the set of labels;
[0300] The label with the highest similarity is obtained, and if the label with the highest similarity belongs to the positive label subset of the set label set, the first sample image after coarse screening is obtained;
[0301] A set of natural content images of the samples is obtained based on the first sample image after coarse screening.
[0302] Based on any of the above embodiments, the scene information acquisition module 1301 is used for:
[0303] The first sample image after coarse screening is input into the large language model, and the large language model is guided by the first prompt word to output the discrimination result of the first sample image after coarse screening;
[0304] Based on the discrimination result, at least a portion of the coarsely screened first sample images are selected from the retained first sample images to obtain finely screened first sample images, so as to obtain a sample natural content image set based on the finely screened first sample images.
[0305] Based on any of the above embodiments, the scene information acquisition module 1301 is used for:
[0306] The sample natural content images in the sample natural content image set are identified to obtain category description information of objects in the sample natural content images and image-level point clouds of the sample natural content images;
[0307] The sample natural content images in the sample natural content image set are input into the visual detection model, and the category description information corresponding to the sample natural content images guides the visual detection model to output the bounding boxes of objects in the sample natural content images;
[0308] Based on the bounding boxes of objects in the sample natural content image, an instance mask of the objects in the sample natural content image is obtained, so as to obtain the object-level point cloud of the objects in the sample natural content image according to the instance mask and the image-level point cloud of the sample natural content image.
[0309] Based on any of the above embodiments, the scene information acquisition module 1301 is used for:
[0310] If the sample natural content images in the sample natural content image set are identified and it is determined that the sample natural content images include at least two objects of the same category, the sample natural content images in the sample natural content image set are input into the category detection model, and the category detection model is guided by a second prompt word to output at least one initial object description information of the sample natural content images.
[0311] A set template corresponding to the sample natural content image is determined, and a hierarchical object description is generated by expanding the initial object description information based on the set template.
[0312] Based on any of the above embodiments, the reasoning information generation module 1302 is used for:
[0313] Based on the set template and the fine-grained 3D scene information, simple spatial reasoning information is generated for the sample natural content image; the simple spatial reasoning information includes at least one of preliminary question-answer pairs, multiple-choice questions, and factual statements; image description information of the sample natural content image is determined;
[0314] The simple spatial reasoning information is input into the reasoning big language model. Based on the image description information of the sample natural content image and the category description information of the objects in the sample natural content image, the reasoning big language model is guided to output the complex spatial reasoning information of the sample natural content image.
[0315] Based on the simple spatial reasoning information and complex spatial reasoning information of the sample natural content image, multi-step spatial reasoning information of the sample natural content image is obtained.
[0316] Based on any of the above embodiments, the scene information acquisition module 1301 is used for:
[0317] Obtain the second sample image dataset, and perform frame sampling processing on the sample video image data in the second sample image dataset to obtain the second sample image;
[0318] The sample natural content image set is obtained based on the second sample image.
[0319] Based on any of the above embodiments, the scene information acquisition module 1301 is used for:
[0320] The sample natural content images in the sample natural content image set are identified to obtain category description information of objects in the sample natural content images;
[0321] The sample natural content images in the sample natural content image set are input into the visual detection model. The visual detection model is guided to output the predicted bounding boxes of the objects in the sample natural content images by the category description information of the objects in the sample natural content images.
[0322] The predicted bounding box and the original bounding box are matched to obtain a matching result, and the predicted bounding box that matches at least one of the original bounding boxes is retained based on the matching result.
[0323] The matching degree between the predicted bounding box and each of the original bounding boxes that match it is determined. Based on the matching degree, the original bounding box with the highest matching degree with the predicted bounding box is retained, so as to obtain fine-grained 3D scene information of the sample natural content image according to the original bounding box with the highest matching degree.
[0324] Based on any of the above embodiments, the scene information acquisition module 1301 is used for:
[0325] The sample natural content images in the sample natural content image set are identified to obtain object-level point clouds and 3D bounding boxes of objects in the sample natural content images;
[0326] Based on the gravity alignment matrix of the second sample image dataset, the object-level point cloud and 3D bounding box of the object are transformed to obtain the gravity-aligned object-level point cloud and 3D bounding box.
[0327] Based on gravity-aligned object-level point clouds and 3D bounding boxes, free space for object placement is identified.
[0328] Based on any of the above embodiments, the scene information acquisition module 1301 is used for:
[0329] When the volume of an object in the sample natural content image is within a set volume range, the free space below the object in the sample natural content image for object placement is identified based on the gravity-aligned object-level point cloud and 3D bounding box.
[0330] Based on any of the above embodiments, the scene information acquisition module 1301 is used for:
[0331] Generate a set number of indoor scenes; wherein, no objects exist on the target surface in the generated indoor scenes;
[0332] Obtain category description information of sample objects in the sample object dataset, and filter the sample object dataset based on the category description information to obtain coarsely screened sample objects;
[0333] Obtain object description information of sample objects in the sample object dataset, and filter the coarse-screened sample objects based on the object description information to obtain fine-screened sample objects;
[0334] Obtain the selection instruction input by the user, and filter the fine-screen sample objects based on the selection instruction input by the user to obtain the target sample object.
[0335] The target sample pair is placed in the indoor scene and rendered to obtain a third sample image, and a sample natural content image set is obtained based on the third sample image.
[0336] Based on any of the above embodiments, the setting template includes a simulation setting template that formalizes strongly dependent spatial relationships;
[0337] Inference information generation module 1302 is used for:
[0338] Based on the category description information and object description information of the sample object, generate the reference expression of the sample object;
[0339] Multi-step spatial reasoning information for generating the natural content image of the sample is generated based on the simulation setting template and the referential expression of the sample object.
[0340] Figure 14 An example is a schematic diagram of the structure of an electronic device, such as... Figure 14 As shown, the electronic device may include: a processor 1410, a communications interface 1420, a memory 1430, and a communication bus 1440, wherein the processor 1410, the communications interface 1420, and the memory 1430 communicate with each other through the communication bus 1440. The processor 1410 can call logical instructions in the memory 1430 to execute a method for constructing a dataset for model training. This method includes: acquiring a set of sample natural content images and recognizing the sample natural content images in the set to obtain fine-grained three-dimensional scene information of the sample natural content images; determining a set template corresponding to the fine-grained three-dimensional scene information; generating multi-step spatial inference information of the sample natural content images based on the set template and the fine-grained three-dimensional scene information; and obtaining a dataset based on the sample natural content images, the fine-grained three-dimensional scene information of the sample natural content images, and the multi-step spatial inference information of the sample natural content images.
[0341] Furthermore, the logical instructions in the aforementioned memory 1430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device, such as a personal computer, server, or network device, to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0342] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the method for constructing a dataset for model training provided by the above methods. The method includes: acquiring a set of sample natural content images and recognizing the sample natural content images in the set to obtain fine-grained three-dimensional scene information of the sample natural content images; determining a set template corresponding to the fine-grained three-dimensional scene information; generating multi-step spatial reasoning information of the sample natural content images based on the set template and the fine-grained three-dimensional scene information; and obtaining a dataset based on the sample natural content images, the fine-grained three-dimensional scene information of the sample natural content images, and the multi-step spatial reasoning information of the sample natural content images.
[0343] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for constructing a dataset for model training provided by the methods described above. This method includes: acquiring a set of sample natural content images and recognizing the sample natural content images in the set to obtain fine-grained three-dimensional scene information of the sample natural content images; determining a set template corresponding to the fine-grained three-dimensional scene information; generating multi-step spatial reasoning information of the sample natural content images based on the set template and the fine-grained three-dimensional scene information; and obtaining a dataset based on the sample natural content images, the fine-grained three-dimensional scene information of the sample natural content images, and the multi-step spatial reasoning information of the sample natural content images.
[0344] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0345] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device, such as a personal computer, server, or network device, to execute the methods described in the various embodiments or some parts of the embodiments.
[0346] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a dataset for model training, characterized in that, include: A set of sample natural content images is acquired, and the sample natural content images in the set are identified to obtain fine-grained three-dimensional scene information of the sample natural content images. A set template corresponding to the fine-grained three-dimensional scene information is determined, and multi-step spatial reasoning information of the sample natural content image is generated based on the set template and the fine-grained three-dimensional scene information; wherein, the set template refers to a question-and-answer template designed in combination with the characteristics of the fine-grained three-dimensional scene information of the sample natural content image, and the multi-step spatial reasoning information refers to information that achieves step-by-step reasoning by combining multiple fine-grained three-dimensional scene information in the sample natural content image; A dataset is obtained based on sample natural content images, fine-grained 3D scene information of the sample natural content images, and multi-step spatial reasoning information of the sample natural content images; The multi-step spatial reasoning information for generating the sample natural content image based on the set template and the fine-grained 3D scene information includes: Based on the set template and the fine-grained 3D scene information, simple spatial reasoning information is generated for the sample natural content image; the simple spatial reasoning information includes at least one of preliminary question-answer pairs, multiple-choice questions, and factual statements; image description information of the sample natural content image is determined; The simple spatial reasoning information is input into the reasoning language model. Based on the image description information of the sample natural content image and the category description information of the objects in the sample natural content image, the reasoning language model is guided to output the complex spatial reasoning information of the sample natural content image. The category description information of the objects in the sample natural content image is obtained by recognizing the sample natural content images in the sample natural content image set. Based on the simple spatial reasoning information and complex spatial reasoning information of the sample natural content image, multi-step spatial reasoning information of the sample natural content image is obtained.
2. The method for constructing a dataset for model training according to claim 1, characterized in that, The acquisition of the sample natural content image set includes: Obtain a first sample image dataset and perform vectorization processing on the first sample images in the first sample image dataset to obtain the embedding of the first sample images; The similarity is determined based on the embedding of the first sample image and the embedding of each label in the set of labels; The label with the highest similarity is obtained, and if the label with the highest similarity belongs to the positive label subset of the set label set, the first sample image after coarse screening is obtained; A set of natural content images of the samples is obtained based on the first sample image after coarse screening.
3. The method for constructing a dataset for model training according to claim 2, characterized in that, The process of obtaining a set of natural content images of samples based on the first sample images after coarse screening also includes: The first sample image after coarse screening is input into the large language model, and the large language model is guided by the first prompt word to output the discrimination result of the first sample image after coarse screening; Based on the discrimination result, at least a portion of the coarsely screened first sample images are selected from the retained first sample images to obtain finely screened first sample images, so as to obtain a sample natural content image set based on the finely screened first sample images.
4. The method for constructing a dataset for model training according to claim 1, characterized in that, The acquisition of the sample natural content image set includes: Obtain the second sample image dataset, and perform frame sampling processing on the sample video image data in the second sample image dataset to obtain the second sample image; The sample natural content image set is obtained based on the second sample image.
5. The method for constructing a dataset for model training according to any one of claims 1-4, characterized in that, Identify the sample natural content images in the sample natural content image set to obtain fine-grained three-dimensional scene information of the sample natural content images, including: The sample natural content images in the sample natural content image set are identified to obtain category description information of objects in the sample natural content images and image-level point clouds of the sample natural content images; The sample natural content images in the sample natural content image set are input into the visual detection model, and the category description information corresponding to the sample natural content images guides the visual detection model to output the bounding boxes of objects in the sample natural content images; Based on the bounding boxes of objects in the sample natural content image, an instance mask of the objects in the sample natural content image is obtained, so as to obtain the object-level point cloud of the objects in the sample natural content image according to the instance mask and the image-level point cloud of the sample natural content image.
6. The method for constructing a dataset for model training according to claim 5, characterized in that, Identify sample natural content images in the sample natural content image set to obtain category description information of objects in the sample natural content images, including: If the sample natural content images in the sample natural content image set are identified and it is determined that the sample natural content images include at least two objects of the same category, the sample natural content images in the sample natural content image set are input into the category detection model, and the category detection model is guided by a second prompt word to output at least one initial object description information of the sample natural content images. A set template corresponding to the sample natural content image is determined, and a hierarchical object description is generated by expanding the initial object description information based on the set template.
7. The method for constructing a dataset for model training according to claim 4, characterized in that, The second sample image includes the original bounding boxes of the objects in the second sample image; Identify the sample natural content images in the sample natural content image set to obtain fine-grained three-dimensional scene information of the sample natural content images, including: The sample natural content images in the sample natural content image set are identified to obtain category description information of objects in the sample natural content images; The sample natural content images in the sample natural content image set are input into the visual detection model. The visual detection model is guided to output the predicted bounding boxes of the objects in the sample natural content images by the category description information of the objects in the sample natural content images. The predicted bounding box and the original bounding box are matched to obtain a matching result, and the predicted bounding box that matches at least one of the original bounding boxes is retained based on the matching result. The matching degree between the predicted bounding box and each of the original bounding boxes that match it is determined. Based on the matching degree, the original bounding box with the highest matching degree with the predicted bounding box is retained, so as to obtain fine-grained 3D scene information of the sample natural content image according to the original bounding box with the highest matching degree.
8. The method for constructing a dataset for model training according to claim 4, characterized in that, Identify the sample natural content images in the sample natural content image set to obtain fine-grained three-dimensional scene information of the sample natural content images, including: The sample natural content images in the sample natural content image set are identified to obtain object-level point clouds and 3D bounding boxes of objects in the sample natural content images; Based on the gravity alignment matrix of the second sample image dataset, the object-level point cloud and 3D bounding box of the object are transformed to obtain the gravity-aligned object-level point cloud and 3D bounding box. Based on gravity-aligned object-level point clouds and 3D bounding boxes, free space for object placement is identified.
9. The method for constructing a dataset for model training according to claim 8, characterized in that, The method for identifying free space for object placement based on gravity-aligned object-level point clouds and 3D bounding boxes includes: If the volume of an object in the sample natural content image is within a set volume range, the free space below the object in the sample natural content image for object placement is identified based on the gravity-aligned object-level point cloud and 3D bounding box.
10. The method for constructing a dataset for model training according to claim 1, characterized in that, The acquisition of the sample natural content image set includes: Generate a set number of indoor scenes; wherein, no objects exist on the target surface in the generated indoor scenes; Obtain category description information of sample objects in the sample object dataset, and filter the sample object dataset based on the category description information to obtain coarsely screened sample objects; Obtain object description information of sample objects in the sample object dataset, and filter the coarse-screened sample objects based on the object description information to obtain fine-screened sample objects; Obtain the selection command input by the user, and filter the fine-screen sample object based on the selection command input by the user to obtain the target sample object; The target sample pair is placed in the indoor scene and rendered to obtain a third sample image, and a sample natural content image set is obtained based on the third sample image.
11. The method for constructing a dataset for model training according to claim 10, characterized in that, The setting template includes a simulation setting template that formalizes strongly dependent spatial relationships; Multi-step spatial reasoning information for generating the sample natural content image based on the set template and the fine-grained 3D scene information includes: Based on the category description information and object description information of the sample object, generate the reference expression of the sample object; Multi-step spatial reasoning information for generating the natural content image of the sample is generated based on the simulation setting template and the referential expression of the sample object.
12. An apparatus for constructing a dataset for model training, characterized in that, include: The scene information acquisition module is used to acquire a set of sample natural content images and to identify the sample natural content images in the set of sample natural content images in order to obtain fine-grained three-dimensional scene information of the sample natural content images. The reasoning information generation module is used to determine the set template corresponding to the fine-grained three-dimensional scene information, and generate multi-step spatial reasoning information of the sample natural content image based on the set template and the fine-grained three-dimensional scene information; wherein, the set template refers to a question-and-answer template designed in combination with the characteristics of the fine-grained three-dimensional scene information of the sample natural content image, and the multi-step spatial reasoning information refers to information that achieves step-by-step reasoning by combining multiple fine-grained three-dimensional scene information in the sample natural content image; The dataset acquisition module is used to obtain a dataset based on sample natural content images, fine-grained three-dimensional scene information of the sample natural content images, and multi-step spatial reasoning information of the sample natural content images; The reasoning information generation module is used for: Based on the set template and the fine-grained 3D scene information, simple spatial reasoning information is generated for the sample natural content image; the simple spatial reasoning information includes at least one of preliminary question-answer pairs, multiple-choice questions, and factual statements; image description information of the sample natural content image is determined; The simple spatial reasoning information is input into the reasoning language model. Based on the image description information of the sample natural content image and the category description information of the objects in the sample natural content image, the reasoning language model is guided to output the complex spatial reasoning information of the sample natural content image. The category description information of the objects in the sample natural content image is obtained by recognizing the sample natural content images in the sample natural content image set. Based on the simple spatial reasoning information and complex spatial reasoning information of the sample natural content image, multi-step spatial reasoning information of the sample natural content image is obtained.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method for constructing a dataset for model training as described in any one of claims 1 to 11.
14. A computer storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for constructing a dataset for model training as described in any one of claims 1 to 11.