A model training method and device
By generating a training dataset corresponding to the target industrial inspection task and performing pre-training and adaptation, the problem of the scarcity of high-quality labeled data in industrial visual inspection is solved, and model training that meets the industrial inspection accuracy is achieved with a very small number of real labeled samples.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHINING 3D TECH CO LTD
- Filing Date
- 2026-04-22
- Publication Date
- 2026-06-12
AI Technical Summary
In existing technologies, high-quality labeled data is scarce in industrial visual inspection scenarios, resulting in insufficient model accuracy of deep learning methods under extremely small sample conditions, which cannot meet the requirements of industrial precision inspection.
A training dataset corresponding to the target industrial inspection task is generated. Synthetic samples are generated using geometric information and rendering parameters. The model is pre-trained using the synthetic dataset. Finally, the model is adapted based on the differences in data feature distribution between the synthetic samples and the real labeled samples.
With a very small number of real labeled samples, a reliable model that meets the accuracy requirements of industrial detection was trained, which solved the problem of scarcity of high-quality labeled data and improved the detection accuracy and robustness of the model in real-world scenarios.
Smart Images

Figure CN122196551A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of model training technology, and in particular to a model training method, apparatus, device and storage medium. Background Technology
[0002] In some modern industrial sectors, industrial visual inspection is a core technological means to ensure product quality, improve production efficiency, and ensure operational safety. Inspection errors often need to be controlled within micrometers or even sub-micrometers to meet manufacturing tolerances and industry safety standards.
[0003] Currently, the above tasks can be detected using deep learning-based methods. These methods typically rely on massive amounts of high-quality labeled data for model training to learn the complex mapping relationship from images to the detection targets.
[0004] Due to the extreme scarcity of high-quality labeled data in industrial visual inspection scenarios in current technological solutions, the existing deep learning methods cannot meet the dependence on large-scale training data, resulting in severe model inaccuracies under extremely small sample conditions, which cannot meet the requirements of industrial precision inspection. Summary of the Invention
[0005] To solve the above-mentioned technical problems, or at least partially solve them, this disclosure provides a model training method, apparatus, device, and storage medium.
[0006] This disclosure provides a model training method, the method comprising: First, a training dataset corresponding to the target industrial inspection task can be generated. The geometric annotations of the synthetic samples in the training dataset are generated based on the geometric information and rendering parameters on which the synthetic samples are generated. Then, the detection model is pre-trained using the training dataset. Finally, real labeled samples corresponding to the target industrial detection task are obtained. Based on the differences in data feature distribution between the synthetic samples in the training dataset and the real labeled samples, the first detection model obtained after pre-training is adapted to obtain the target detection model corresponding to the target industrial detection task. The entire scheme enables the training of a reliable model that meets the accuracy requirements of industrial detection even when there are only a few real labeled samples in each class.
[0007] This disclosure also provides a model training apparatus, the apparatus comprising: The generation unit is used to generate a training dataset corresponding to the target industrial inspection task, wherein the geometric annotations of the synthetic samples in the training dataset are generated based on the geometric information and rendering parameters on which the synthetic samples are generated; A pre-training unit is used to pre-train the detection model using the training dataset; An adaptation unit is used to obtain real labeled samples corresponding to the target industrial inspection task, and adapt the first detection model obtained after pre-training based on the data feature distribution difference between the synthetic samples in the training dataset and the real labeled samples to obtain the target detection model corresponding to the target industrial inspection task.
[0008] This disclosure also provides a computing device, the computing device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the model training method provided in this disclosure.
[0009] This disclosure also provides a computer-readable storage medium storing a computer program for executing the model training method provided in this disclosure. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0011] Figure 1 A schematic flowchart of a model training method provided in an embodiment of this disclosure; Figure 2 This is a schematic diagram of the structure of a model training device provided in an embodiment of the present disclosure; Figure 3 This is a schematic diagram of the structure of a computing device provided in an embodiment of the present disclosure. Detailed Implementation
[0012] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0013] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0014] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0015] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0016] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0017] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0018] In some modern industrial sectors, industrial vision inspection is a core technological means to ensure product quality, improve production efficiency, and ensure operational safety. Its tasks, such as the dimensional measurement and classification of holes in precision parts and the location and identification of surface defects on engine blades, place extremely stringent requirements on the accuracy, reliability, and generalization ability of the inspection. The inspection error often needs to be controlled at the micrometer or even sub-micrometer level to meet manufacturing tolerances and industry safety standards.
[0019] Currently, deep learning-based object detection methods have shown great potential in the aforementioned tasks and have become the mainstream technical approach. These methods typically rely on massive amounts of high-quality labeled data for model training to learn the complex mapping relationship from images to the objects to be detected.
[0020] Due to the extreme scarcity of high-quality labeled data in industrial visual inspection scenarios in current technological solutions, the existing deep learning methods cannot meet the dependence on large-scale training data, resulting in severe model inaccuracies under extremely small sample conditions, which cannot meet the requirements of industrial precision inspection.
[0021] This application provides a model training method, including generating a training dataset corresponding to a target industrial inspection task. The geometric annotations of synthetic samples in the training dataset are generated based on the geometric information and rendering parameters used to generate the synthetic samples. This directly produces a large amount of accurately annotated training data, fundamentally replacing the reliance on scarce manually annotated data and providing a sufficient data foundation for model training. Then, the model is pre-trained using the synthetic dataset, enabling it to establish basic feature representations of the inspected object. Finally, the model is adapted based on the differences in data feature distribution between the synthetic data and a very small amount of real data to correct inter-domain differences between simulation data and real-world scenarios. This ensures that the knowledge learned by the model on abundant synthetic data can be effectively calibrated and transferred to a very small number of real samples. Therefore, the entire scheme enables the training of a reliable model that meets the accuracy requirements of industrial inspection even with only a small number of real annotated samples in each class.
[0022] The method will be described below with reference to specific embodiments.
[0023] Figure 1 This is a flowchart illustrating a model training method provided in an embodiment of the present disclosure. The method can be executed by a model training device, which can be implemented in software and / or hardware, and is generally integrated into a computing device. Figure 1 As shown, the method includes: S101: Generate a training dataset corresponding to the target industrial inspection task.
[0024] The computing device can generate a training dataset corresponding to the target industrial inspection task. The geometric annotations of the synthetic samples in the training dataset are generated based on the geometric information and rendering parameters used to generate the synthetic samples. The training dataset refers to the data set generated to train the inspection model. Synthetic samples refer to images generated through computation, rather than real-world images directly captured by a camera. Geometric annotations are label information related to the geometric attributes of the target object in the image, such as shape, position, and size, used to guide model learning. Examples include bounding box coordinates, pixel-level segmentation masks, or keypoint coordinates. The target industrial inspection task refers to the specific industrial vision inspection work that needs to be completed using the trained model, such as hole recognition and classification in a flash inspection scenario or industrial defect detection.
[0025] The geometric information defines the shape, structure, and spatial relationships of the features to be detected (such as various pore types, or defects like solid cracks and corrosion pits) in three-dimensional space, forming the basis for generating a synthetic image with correct geometric morphology. Rendering parameters define the visual conditions during image generation, such as simulated camera angles, lighting settings, and the appearance properties of the object's surface material. During the generation of each synthetic sample image, a corresponding two-dimensional image is generated through calculation (e.g., using a physically based rendering engine or neural rendering technology) based on the preset geometric information and a set of rendering parameters. Geometric annotations are not manually added or estimated afterward, but are automatically generated during the same calculation process of generating the two-dimensional image, based on the geometric information and rendering parameters, through deterministic projection, calculation, or derivation relationships. Therefore, the geometric annotations and the synthetic sample image have inherent consistency at the geometric level, and their accuracy is guaranteed by the geometric information and the rendering calculation process itself.
[0026] In some possible implementations, generating the training dataset can specifically be as follows: Obtain material sample images of the objects to be inspected corresponding to the target industrial inspection task, fit the material sample images to obtain the material optical parameters of the objects to be inspected, construct a rendering data model of the objects to be inspected based on the material optical parameters and geometric information, drive the physical rendering engine to generate the synthetic samples through the rendering data model and the rendering parameters, and use the synthetic samples to generate the training dataset.
[0027] For example, this step involves building a digital simulation model (i.e., a rendered data model) for the target industrial inspection object (e.g., an industrial part to be inspected, such as a precision part or blade, rail wheel, etc.), and using this digital simulation model to generate synthetic samples with reliable geometric annotations in batches, which are then collected to form a training dataset.
[0028] Specifically, the process begins by acquiring a small number of material sample images of the object being inspected. These images reflect the visual appearance of the object's surface under realistic lighting. Through an optimization calculation process involving inverse rendering fitting of these images, material optical parameters (such as the bidirectional reflectance distribution function (BRDF) parameters) describing the object's surface's physical behavior regarding light reflection and scattering can be obtained. Subsequently, these material optical parameters are combined with geometric information to construct a rendering data model of the object being inspected. This rendering data model can be a complete digital representation integrating accurate geometry and realistic materials. When generating synthetic samples, this rendering data model, along with rendering parameters defining imaging conditions (such as camera pose and light source settings), is input into the physical rendering engine, which then generates high-fidelity synthetic sample images following the laws of physical optics. It should be noted that the geometric annotations of the synthetic samples (such as the boundaries of apertures and the contours of defects) can be calculated by the rendering engine using the same set of geometric information and rendering parameters while rendering the image.
[0029] This approach, by generating synthetic samples based on a physically based rendering engine, ensures the accuracy and reliability of the geometric annotations in the synthetic data. Unlike traditional methods such as GANs, which may produce synthetic samples with annotation noise, this embodiment constructs a rendering data model based on geometric information and optical parameters of the real material obtained through inverse rendering fitting, and drives the physically based rendering engine to generate images. During this process, the geometric annotations of the synthetic samples can be calculated by the rendering engine based on the geometric information and rendering parameters; theoretically, the annotation error is only limited by the accuracy of the rendering engine itself. Therefore, this method directly provides large-scale, high-quality training data with inherently accurate annotations for industrial inspection scenarios with stringent geometric accuracy requirements, such as hole size measurement and defect size determination, effectively solving the fundamental problem of the scarcity of high-quality annotation data.
[0030] In some possible implementations, generating the training dataset can specifically be as follows: A real scene image corresponding to the target industrial inspection task is obtained. Based on the real scene image, a neural radiation field corresponding to the target industrial inspection task is constructed. Multiple synthetic images from different perspectives are rendered and generated as synthetic samples through the neural radiation field. The geometric annotation of the synthetic samples is determined based on the geometric information included in the neural radiation field.
[0031] For example, a computing device can acquire a small number of real-world scene images (such as images of perforated parts or defective blades taken from different angles) related to a target industrial inspection task. Then, using these real-world scene images as supervision, a neural network is trained to construct a neural radiation field representing that specific scene. This neural radiation field includes rich geometric information. During the data generation phase, by setting rendering parameters (e.g., including new, desired camera viewpoints) to the trained neural radiation field, high-fidelity synthetic images observed from those different viewpoints can be rendered; these synthetic images serve as synthetic samples. It should be noted that because the neural radiation field encodes precise 3D geometric information (such as surface density distribution), while rendering the synthetic sample images, information such as the depth, surface normal, or segmentation mask corresponding to the target object in the image can be determined based on this neural radiation field; this information constitutes the geometric annotation of the synthetic sample.
[0032] This approach learns and reconstructs a resolvable 3D scene representation from a small number of real-world scene images, simultaneously determining geometric annotations while generating sample images. It provides an effective means of augmenting data using real-world scene images collected on-site, without relying on pre-set CAD models. It also addresses the challenges of extremely scarce annotation data and the reliability of geometric annotations in synthetic data during industrial visual inspection.
[0033] In some possible implementations, the synthesized samples are generated by driving a physical rendering engine through the rendering data model and the rendering parameters. Specifically, this can be: Based on the target industrial inspection task, a structured cue word template describing the rendering conditions is constructed. The structured cue word template is input into a large language model to generate various different rendering parameters that conform to the constraints of the structured cue word template. Through the rendering data model and various different rendering parameters, the physical rendering engine is driven to generate a first number of synthetic samples.
[0034] The structured cue word template is a pre-designed text framework with specific fields and formats used to describe various conditional parameters required to generate synthetic samples. For example, in a flash hole type recognition scenario, this structured cue word template can be used to describe parameters such as hole type, part material, surface roughness, illumination angle, and camera pose; in an industrial defect detection scenario, it can be used to describe parameters such as the physical morphology and size of defects, part surface areas, defect age, and detection illumination conditions. The constructed structured cue word template is then input into a large language model. The large language model understands the semantics and constraints of the template and generates various rendering parameters that conform to the constraints of the structured cue word template. This step realizes the transformation from text description to a specific, executable set of rendering parameters, enabling the generation of a large number of parameter combinations with differences in dimensions such as illumination, viewpoint, and surface state, thereby greatly expanding the diversity of synthetic samples.
[0035] Finally, the constructed rendering data model is combined with various rendering parameters generated by the large language model to jointly drive the physically based rendering engine to perform rendering calculations. Based on each different set of parameter configurations, the physically based rendering engine simulates the corresponding physical imaging process, thereby automatically and continuously generating the first batch of synthetic samples. This method replaces manual, tedious parameter design and configuration with a large language model, enabling the batch generation of diverse training samples, significantly improving data synthesis efficiency, covering the multi-dimensional scene changes required for detection tasks, and providing a sufficient data foundation for model training under small sample conditions.
[0036] S102: Use the training dataset to pre-train the detection model.
[0037] The computing device can use the training dataset to pre-train the detection model.
[0038] For example, a computing device can use a training dataset to pre-train a detection model.
[0039] This means that the detection model can be progressively pre-trained using the training dataset, following an order of increasing complexity of the synthetic training samples. This allows the detection model to follow a gradual learning process from simple to complex.
[0040] The complexity here can refer to the difficulty of the industrial testing conditions simulated by the synthetic sample. For example, simple samples with a normal viewing angle, uniform lighting, and no obstruction may be used for training first, and then difficult samples with tilted viewing angle, complex reflection, simulated clamping deviation, or background interference may be gradually introduced.
[0041] For example, the detection model can first learn on a large number of relatively simple synthetic training samples to establish a basic understanding of the target's fundamental shape and features. Subsequently, the complexity of the input synthetic samples is gradually increased, guiding the detection model to adapt to and solve more challenging visual situations. Through this progressive training approach, the detection model can learn feature representations applicable to the target industrial detection domain more robustly and systematically, thus laying a good initial foundation for rapid adaptation on a very small number of real labeled samples.
[0042] S103: Obtain the real labeled samples corresponding to the target industrial inspection task, and adapt the first detection model obtained after pre-training based on the data feature distribution differences between the synthetic samples in the training dataset and the real labeled samples.
[0043] The computing device can acquire real labeled samples corresponding to the target industrial inspection task. Based on the data feature distribution differences between the synthetic samples in the training dataset and the real labeled samples, it adapts the first detection model obtained after pre-training to obtain the target detection model corresponding to the target industrial inspection task. Thus, the target detection model can be used to detect the inspection object.
[0044] To address the discrepancy (or domain difference) in data feature distribution between the synthetic domain (composed of synthetic samples) and the real domain (represented by real labeled samples), which stems from differences in texture, noise, and lighting details between simulated generation and real imaging, directly applying a first detection model pre-trained only on synthetic data to a real scene may result in decreased performance due to the lack of learning of real data features.
[0045] Therefore, the first detection model needs to be adapted. Adaptation refers to the process of specifically adjusting the model using a very small number of real labeled samples corresponding to the target industrial detection task.
[0046] For example, this step is driven by differences in data feature distribution. Through a specific learning mechanism (such as meta-learning methods), the first detection model can quickly transfer and adapt the general knowledge learned on rich synthetic data to the real data distribution represented by a small number of real labeled samples. For instance, in the flash hole detection scenario, a metric-based meta-learning method can be used for adaptation. In the industrial defect detection scenario, an optimization-based meta-learning method can be used. Through this adaptation step, the detection model can correct its over-reliance on synthetic data features, learn to focus on key visual patterns in the real scene, and ultimately obtain a target detection model specifically designed for the target industrial inspection task.
[0047] In some possible implementations, adapting the pre-trained first detection model based on the differences in the data feature distribution can specifically be as follows: Computational devices can perform geometrically constrained data augmentation on real-world labeled samples.
[0048] Data augmentation that preserves geometric constraints includes at least one of the following operations: While keeping the geometry of the target object unchanged in the real annotated sample, simulate the clamping deviation of the part and apply perspective transformation within a preset range to the sample image; Simulate various light source conditions in an industrial setting, and overlay simulated highlight areas and / or shadow areas onto sample images; For reflective material surfaces, apply brightness and / or morphological perturbations to the highlight areas in the sample image; When the real labeled samples are binocular vision image pairs, apply the same image transformation to the left and right eye images; The first detection model is adapted based on the differences in data feature distribution between the synthetic samples and the real labeled samples after data augmentation.
[0049] Data augmentation that preserves geometric constraints refers to maintaining the key geometric characteristics (such as shape and size) of the target object (such as the hole type to be identified or the defect to be detected) in the real-world labeled sample while applying any image transformation. This ensures that the geometric annotation of the augmented sample remains accurate and reliable. This data augmentation can include at least one of the following specially designed operations: First, simulating part clamping deviations by applying a slight perspective transformation to the sample image within a preset allowable range to simulate minute angular changes in part placement during actual inspection. Second, simulating various light source conditions in an industrial setting by overlaying simulated highlight and / or shadow areas on the image to reproduce complex lighting effects under multi-light source environments. Third, for reflective material surfaces (such as polished metal), applying brightness or shape perturbations to the highlight areas in the image to improve the model's tolerance to metal reflection interference. Fourth, when the real-world labeled sample is a pair of binocular visual images, the same image transformation must be applied to the left and right visual images to ensure parallax geometric consistency.
[0050] After completing the above data augmentation, the first detection model is adapted based on the differences in data feature distribution between the synthetic samples and the augmented real-labeled samples. The data augmentation operation, by enriching the visual representation of the real-labeled samples, indirectly helps the detection model learn more effectively during the adaptation phase how to correct the data distribution differences between the synthetic samples and the (augmented) real samples, thereby improving the accuracy and robustness of the final adapted target detection model.
[0051] In some possible implementations, computing devices can construct multiple few-shot learning tasks from real-labeled samples.
[0052] With the goal of correcting the differences in data feature distribution, the first detection model is adapted by iteratively learning and updating on the multiple small sample learning tasks.
[0053] For example, the computing device can first construct multiple few-shot learning tasks from real-labeled samples. Each few-shot learning task simulates a complete, data-scarce new detection scenario. For example, it can include a support set (for quickly learning the new task) and a query set (for evaluation and updating). By sampling and combining from a limited number of real-labeled samples, multiple such tasks can be constructed, each designed to teach the detection model how to identify a specific target from a small number of real-labeled samples. This allows the first detection model to learn the ability to quickly narrow down or adapt to different data distributions (referring to the learning ability from synthetic data distributions to real data distributions).
[0054] For example, the first detection model iteratively learns and updates on multiple few-shot learning tasks. In each iteration, the first detection model performs internal learning on the support set of a few-shot learning task (meaning the first detection model quickly learns and adjusts parameters using its support set in a single few-shot learning task) and evaluates its adaptation on the query set of that task. Then, based on the adaptation experience accumulated on multiple such tasks, the first detection model performs iterative parameter updates (meaning the first detection model updates and optimizes its initial parameters (i.e., the starting parameters before facing a new task) based on the accumulated performance feedback (such as the loss on each query set) after undergoing internal learning adaptation and evaluation on multiple such tasks. The direction of this update is to optimize the model's initial parameters so that the first detection model can achieve good performance with minimal adjustments when facing any new few-shot task defined by a very small number of real labeled samples. Through this iterative learning on a large number of simulated tasks, the first detection model ultimately efficiently transfers and adapts the knowledge it has pre-trained on rich synthetic data to real-world, few-shot detection scenarios, thereby obtaining a high-precision model that can be used for practical deployment.
[0055] In some possible implementations, the computing device can also decompose the target industrial inspection task into multiple inspection sub-tasks based on the domain prior information of the target industrial inspection task, and then adapt the detection model corresponding to each pre-trained inspection sub-task to each sub-task. The training dataset includes the synthetic training datasets generated for each of the inspection sub-tasks, and the training datasets of different inspection sub-tasks are used to pre-train their respective corresponding detection models.
[0056] For example, in the flash hole shape recognition scenario, hole shapes of different sizes can be divided into different detection sub-task groups based on prior domain information (such as the hole diameter size range parsed from the part's CAD model). (For example, the recognition of small, medium, and large hole shapes can be treated as independent sub-tasks.) In the industrial defect detection scenario, defect categories can be decomposed into two main detection sub-tasks based on the physical morphology of the defects: linear defect groups (such as scratches, cracks, etc.) and regional defect groups (such as corrosion pits, peeling, etc.).
[0057] This allows for the independent generation of corresponding synthetic training datasets for each detection subtask. This means the training data is specifically tailored to the features of each subtask (such as holes within a specific size range or defects of a specific shape), thus forming training datasets for different detection subtasks. This training data can then be used to pre-train their respective detection models; that is, each subtask has a specialized model to learn, effectively reducing the complexity of a single model needing to distinguish all categories simultaneously. During the adaptation phase, the pre-trained detection models for each subtask can be adapted separately, i.e., using a small number of real-labeled samples for each subtask to independently adjust their respective detection models. Finally, during inference, the results of these parallel-working sub-detection models can be integrated to output a complete detection result. This significantly improves the overall learning efficiency and final detection accuracy of the model under conditions of limited real-labeled samples. By leveraging prior domain information (such as geometric dimensions and defect physical morphology), complex industrial target detection tasks are decomposed into multiple detection subtasks. The divide-and-conquer strategy significantly reduces the learning difficulty of each subtask under small sample conditions. Based on this, targeted synthetic training data is generated independently for each subtask and combined with subsequent pre-training and adaptation. Together, these data enable the detection model to achieve high-precision target detection even with very few real labeled samples, effectively solving the problem of insufficient model generalization ability caused by the scarcity of high-quality labeled data in industrial scenarios.
[0058] In some possible implementations, the computing device can also parse the domain prior information, wherein parsing the domain prior information includes at least one of the following methods: Analyze the geometric category and scale information of the features to be detected in the preset model of the object to be detected; Based on the physical morphology information of defects in the target industrial inspection task, the defect categories are determined as linear defect group and regional defect group; Then, based on geometric category, scale information, or physical morphology information of defects, the target industrial inspection task is decomposed into multiple inspection sub-tasks.
[0059] Here, the preset model can be, for example, a computer-aided design model, and the feature to be detected can be various hole types. By analyzing this preset model, the geometric category of the hole type (such as through hole, blind hole, or threaded hole) and its precise dimensional information (such as the numerical range of the hole diameter) can be extracted. Secondly, based on the physical morphological information of defects in the target industrial inspection task, the defect categories are determined into linear defect groups and regional defect groups. This corresponds to defect inspection scenarios for objects such as blades and wheels. Based on the physical morphological information of defects (such as shape, ductility, etc.), all defect categories can be definitively divided into two major groups: linear defect groups (such as scratches, cracks, etc., defects with obvious length characteristics) and regional defect groups (such as corrosion pits, peeling, etc., defects distributed in a planar or blocky manner).
[0060] After completing the above analysis, the target industrial inspection task can be decomposed into multiple sub-tasks based on geometric categories, scale information, or the physical morphology information of defects. For example, in the aperture recognition scenario, the recognition task can be decomposed into different sub-tasks targeting small, medium, and large apertures based on the different scale information obtained from the analysis (e.g., aperture range). In the defect detection scenario, the overall defect detection task can be decomposed into two independent sub-tasks directly based on the division of the physical morphology information of the defects (e.g., dividing them into linear groups and regional groups). This analysis step lays the foundation for subsequently generating data and training models independently for each sub-task, effectively reducing the learning difficulty of each sub-task under small sample conditions.
[0061] In some possible implementations, the computing device can initially train a first detection model based on an industrial image dataset, with the goal of correcting the differences in data feature distribution between synthetic samples and the real labeled samples. Based on the real labeled samples, the first detection model can be adapted by fine-tuning the network parameters of the initially trained first detection model.
[0062] For example, this step may include two stages. The first stage is initial training, which trains the first detection model based on an industrial image dataset (e.g., a pre-existing, large-scale collection of real industrial scene images). The purpose of this step is to enable the first detection model to learn general visual feature representations from massive and diverse real industrial images, thereby obtaining a good initial state oriented towards the industrial field, which forms the basis for subsequent adaptation. The second stage is fine-tuning and adaptation, which aims to correct the differences in data feature distribution between the synthetic samples and the real labeled samples, using a very small number of real labeled samples corresponding to the current target industrial detection task as training data. By making targeted and small-scale adjustments (i.e., fine-tuning) to the network parameters of the first detection model that has completed initial training, the feature representation of the first detection model is aligned from the general industrial image distribution to the data distribution defined by a small number of real labeled samples specific to the current task. This fine-tuning process is the core action of adapting the model. It should be noted that this initial training is different from pre-training; it is an operation performed after pre-training, and their functions and purposes are also different.
[0063] By combining large-scale industrial image pre-training with small-sample real data fine-tuning of transfer learning strategies, it is also possible to effectively utilize prior knowledge and correct for inter-domain distribution differences, ultimately enabling the first detection model to adapt to specific real detection scenarios.
[0064] This application generates a training dataset corresponding to the target industrial inspection task. The geometric annotations of the synthetic samples in the training dataset are generated based on the geometric information and rendering parameters used to generate the synthetic samples. This directly produces a massive amount of accurately annotated training data, fundamentally replacing the reliance on sparse manually annotated data and providing a sufficient data foundation for model training. Then, the model is pre-trained using the synthetic dataset, enabling it to establish basic feature representations of the inspected objects. Finally, the model is adapted based on the differences in data feature distribution between the synthetic data and a very small amount of real data to correct the inter-domain differences between the simulation data and the real scene. This ensures that the knowledge learned by the model on the rich synthetic data can be effectively calibrated and transferred to a very small number of real samples. Therefore, the entire scheme enables the training of a reliable model that meets the accuracy requirements of industrial inspection even with only a small number of real annotated samples in each class.
[0065] To implement the above embodiments, this disclosure also proposes a model training apparatus.
[0066] Figure 2 This is a schematic diagram of a model training device provided in an embodiment of the present disclosure. The device can be implemented by software and / or hardware, and is generally integrated into a computing device. Figure 2 As shown, the device includes: The generation unit 210 is used to generate a training dataset corresponding to the target industrial inspection task, wherein the geometric annotations of the synthetic samples in the training dataset are generated based on the geometric information and rendering parameters on which the synthetic samples are generated. The pre-training unit 220 is used to pre-train the detection model using the training dataset; The adaptation unit 230 is used to obtain real labeled samples corresponding to the target industrial inspection task, and adapt the first detection model obtained after pre-training based on the data feature distribution difference between the synthetic samples in the training dataset and the real labeled samples to obtain the target detection model corresponding to the target industrial inspection task.
[0067] Optionally, the generating unit is specifically used for: Obtain material sample images of the objects to be inspected corresponding to the target industrial inspection task; The material sample image is fitted to obtain the material optical parameters of the detected object; Based on the material's optical parameters and the geometric information, a rendering data model of the detected object is constructed; The physical rendering engine is driven to generate the synthetic samples using the rendering data model and the rendering parameters, and the training dataset is generated using the synthetic samples.
[0068] Optionally, the generating unit is specifically used for: Based on the target industrial inspection task, construct a structured prompt word template describing the rendering conditions; The structured prompt word template is input into the large language model to generate various different rendering parameters that conform to the constraints of the structured prompt word template; The rendering data model, along with various rendering parameters, drives the physical rendering engine to generate a first number of synthetic samples.
[0069] Optionally, the adapter unit is specifically used for: Geometrically constrained data augmentation is applied to the ground-labeled samples. Data augmentation that preserves geometric constraints includes at least one of the following operations: While keeping the geometry of the target object in the real labeled sample unchanged, simulate the clamping deviation of the part and apply perspective transformation within a preset range to the sample image; Simulate various light source conditions in an industrial setting, and overlay simulated highlight areas and / or shadow areas onto sample images; For reflective material surfaces, apply brightness and / or morphological perturbations to the highlight areas in the sample image; When the real labeled sample is a pair of binocular visual images, the same image transformation is applied to the left and right images; The first detection model is adapted based on the differences in data feature distribution between the synthetic sample and the real labeled sample after data augmentation.
[0070] Optional, pre-trained units, specifically used for: Using the training dataset, the detection model is progressively pre-trained in order of increasing complexity of the synthetic training samples; Based on the differences in the data feature distribution, the first detection model is adapted, including: The real labeled samples are constructed into multiple few-sample learning tasks; With the goal of correcting the differences in the distribution of the data features, the first detection model is adapted by iteratively learning and updating on the multiple small sample learning tasks.
[0071] Optionally, the device further includes: The decomposition unit is used to decompose the target industrial inspection task into multiple inspection sub-tasks based on the domain prior information of the target industrial inspection task. The adapter unit is specifically used for: The detection models corresponding to each of the pre-trained detection sub-tasks are adapted accordingly; the training dataset includes the synthetic training datasets generated for each of the detection sub-tasks, and the training datasets for different detection sub-tasks are used to pre-train their respective corresponding detection models.
[0072] Optionally, the device further includes: A parsing unit is used to parse the domain prior information, wherein parsing the domain prior information includes at least one of the following methods: Analyze the geometric category and scale information of the features to be detected in the preset model of the object to be detected; Based on the physical morphology information of defects in the target industrial inspection task, the defect categories are determined as linear defect group and regional defect group; The decomposition unit is specifically used to decompose the target industrial inspection task into multiple inspection sub-tasks based on the geometric category, scale information, or physical morphology information of the defects.
[0073] Optional, generating unit, specifically used for: Acquire real-world scene images corresponding to the target industrial inspection task; Based on the real scene images, construct the neural radiation field corresponding to the target industrial inspection task; Multiple synthetic images from different perspectives are generated using the neural radiation field as synthetic samples, wherein the geometric annotations of the synthetic samples are determined based on the geometric information included in the neural radiation field.
[0074] Optional, adapter unit, specifically used for: The first detection model was initially trained based on an industrial image dataset. With the goal of correcting the difference in data feature distribution between the synthetic samples and the real labeled samples, the first detection model is adapted by fine-tuning the network parameters of the first detection model trained for the first time based on the real labeled samples.
[0075] The model training apparatus provided in this disclosure can execute the model training method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0076] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program / instructions, which, when executed by a processor, implements the model training method in the above embodiments.
[0077] Figure 3 This is a schematic diagram of the structure of a computing device provided in an embodiment of the present disclosure.
[0078] The following is a detailed reference. Figure 3 The diagram illustrates a structural schematic suitable for implementing the computing device 300 in the embodiments of this disclosure. The computing device 300 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 3 The computing device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0079] like Figure 3 As shown, the computing device 300 may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 302 or a program loaded from memory 308 into random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the computing device 300. The processor 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0080] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows computing device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 A computing device 300 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0081] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from memory 308, or installed from ROM 302. When the computer program is executed by processor 301, it performs the functions defined above in the model training method of embodiments of this disclosure.
[0082] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0083] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0084] The aforementioned computer-readable medium may be included in the aforementioned computing device; or it may exist independently and not assembled into the computing device.
[0085] The aforementioned computer-readable medium carries one or more programs, which, when executed by the computing device, cause the computing device to perform the aforementioned model training method.
[0086] The computing device can be programmed with computer program code in one or more programming languages or a combination thereof to perform the operations of this disclosure. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0087] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0088] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0089] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0090] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0091] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0092] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0093] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A model training method, characterized in that, The method includes: A training dataset corresponding to the target industrial inspection task is generated, wherein the geometric annotations of the synthetic samples in the training dataset are generated based on the geometric information and rendering parameters on which the synthetic samples are generated; The detection model is pre-trained using the training dataset. Obtain real labeled samples corresponding to the target industrial inspection task. Based on the data feature distribution differences between the synthetic samples in the training dataset and the real labeled samples, adapt the pre-trained first detection model to obtain the target detection model corresponding to the target industrial inspection task.
2. The method according to claim 1, characterized in that, The training dataset is generated by: Obtain material sample images of the objects to be inspected corresponding to the target industrial inspection task; The material sample image is fitted to obtain the material optical parameters of the detected object; Based on the material's optical parameters and the geometric information, a rendering data model of the detected object is constructed; The physical rendering engine is driven to generate the synthetic samples using the rendering data model and the rendering parameters, and the training dataset is generated using the synthetic samples.
3. The method according to claim 2, characterized in that, The step of driving the physical rendering engine to generate the synthesized sample using the rendering data model and the rendering parameters includes: Based on the target industrial inspection task, construct a structured prompt word template describing the rendering conditions; The structured prompt word template is input into the large language model to generate various different rendering parameters that conform to the constraints of the structured prompt word template; The rendering data model, along with various rendering parameters, drives the physical rendering engine to generate a first number of synthetic samples.
4. The method according to any one of claims 1-3, characterized in that, Based on the differences in the distribution of the data features, the first detection model obtained through pre-training is adapted, including: Geometrically constrained data augmentation is applied to the ground-labeled samples. Data augmentation that preserves geometric constraints includes at least one of the following operations: While keeping the geometry of the target object in the real labeled sample unchanged, simulate the clamping deviation of the part and apply perspective transformation within a preset range to the sample image; Simulate various light source conditions in an industrial setting, and overlay simulated highlight areas and / or shadow areas onto sample images; For reflective material surfaces, apply brightness and / or morphological perturbations to the highlight areas in the sample image; When the real labeled sample is a pair of binocular visual images, the same image transformation is applied to the left and right images; The first detection model is adapted based on the differences in data feature distribution between the synthetic sample and the real labeled sample after data augmentation.
5. The method according to any one of claims 1-3, characterized in that, The step of pre-training the detection model using the training dataset includes: Using the training dataset, the detection model is progressively pre-trained in order of increasing complexity of the synthetic samples; Based on the differences in the data feature distribution, the first detection model is adapted, including: The real labeled samples are constructed into multiple few-sample learning tasks; With the goal of correcting the differences in the distribution of the data features, the first detection model is adapted by iteratively learning and updating on the multiple small sample learning tasks.
6. The method according to claim 1, characterized in that, The method further includes: Based on the domain prior information of the target industrial inspection task, the target industrial inspection task is decomposed into multiple inspection sub-tasks; The adaptation of the first detection model includes: The detection models corresponding to each of the pre-trained detection sub-tasks are adapted accordingly; the training dataset includes the synthetic training datasets generated for each of the detection sub-tasks, and the training datasets for different detection sub-tasks are used to pre-train their respective corresponding detection models.
7. The method according to claim 6, characterized in that, The method further includes: The domain prior information is parsed, wherein parsing the domain prior information includes at least one of the following methods: Analyze the geometric category and scale information of the features to be detected in the preset model corresponding to the detection object; Based on the physical morphology information of defects in the target industrial inspection task, the defect categories are determined as linear defect group and regional defect group; Based on the aforementioned prior information in the domain, the target industrial inspection task is decomposed into multiple inspection sub-tasks, including: Based on the geometric category, scale information, or physical morphology information of the defects, the target industrial inspection task is decomposed into multiple inspection sub-tasks.
8. The method according to claim 1, characterized in that, The training dataset is generated by: Acquire real-world scene images corresponding to the target industrial inspection task; Based on the real scene images, construct the neural radiation field corresponding to the target industrial inspection task; Multiple synthetic images from different perspectives are generated using the neural radiation field as synthetic samples, wherein the geometric annotations of the synthetic samples are determined based on the geometric information included in the neural radiation field.
9. The method according to claim 1, characterized in that, Based on the differences in the data feature distribution, the first detection model is adapted, including: The first detection model was initially trained based on an industrial image dataset. With the goal of correcting the difference in data feature distribution between the synthetic samples and the real labeled samples, the first detection model is adapted by fine-tuning the network parameters of the first detection model trained for the first time based on the real labeled samples.
10. A model training device, characterized in that, The device includes: The generation unit is used to generate a training dataset corresponding to the target industrial inspection task, wherein the geometric annotations of the synthetic samples in the training dataset are generated based on the geometric information and rendering parameters on which the synthetic samples are generated; A pre-training unit is used to pre-train the detection model using the training dataset; An adaptation unit is used to obtain real labeled samples corresponding to the target industrial inspection task, and adapt the first detection model obtained after pre-training based on the data feature distribution difference between the synthetic samples in the training dataset and the real labeled samples to obtain the target detection model corresponding to the target industrial inspection task.