A data processing method, device, apparatus, storage medium, and program product

CN122618246BActive Publication Date: 2026-10-09DEXFORCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611107476.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-24
Publication Date
2026-10-09
Estimated Expiration
2046-07-24

AI Technical Summary

Technical Problem

但在该人工标注方案中,单帧掩码绘制耗时数分钟至数十分钟,加之训练评估部署流程,单次迭代周期常达数天,在严重影响模型迭代效率的同时,也约束了产线效率

Benefits of technology

在本申请实施例中,获取业务场景中场景实例分割模型运行时所产生的业务分割数据集;若业务分割数据集中存在可触发场景实例分割模型进行升级的异常分割数据,则对业务分割数据集中的每个业务分割数据中的候选实例进行零样本姿态估计,得到每个业务分割数据中的候选实例的位姿参数,根据预设筛选策略对每个业务分割数据中的候选实例进行筛选,得到每个业务分割数据中的目标实例;根据每个目标实例分别对应的三维模型,对每个目标实例的位姿参数进行坐标系转换,得到每个目标实例分别位于世界坐标系的目标位姿参数;对每个业务分割数据中的目标实例分别对应的目标位姿参数和三维模型进行虚拟渲染,得到H个仿真数据;一个仿真数据中的仿真实例是基于该仿真数据所对应的业务分割数据所对应的目标实例所生成的;H为正整数;根据H个仿真数据对场景实例分割模型进行模型训练,得到第一实例分割模型,根据H个仿真数据对初始实例分割模型进行模型训练,得到第二实例分割模型;初始实例分割模型是指加载通用预训练权重的实例分割模型;根据业务分割数据集分别对场景实例分割模型、第一实例分割模型和第二实例分割模型进行模型验证,得到最优实例分割模型;最优实例分割模型是基于第一实例分割模型或第二实例分割模型确定的;最优实例分割模型用于替换部署在业务场景中的场景实例分割模型。通过以上过程,采集了业务场景中场景实例分割模型运行产生的业务分割数据集,精准捕捉导致模型性能衰减的异常分割数据,作为模型迭代优化的核心数据基础,依托零样本姿态估计、实例筛选、坐标系转换及虚拟渲染等自动化流程,实现仿真标注数据全自动生成,无需人工逐像素标注,大幅缩短模型迭代周期、提升了产线运行效率,同时规避了人工标注的主观误差,统一数据生成标准,稳定保障训练数据质量,有效提升模型优化效果。且可自动完成候选实例姿态解算、精准筛选、坐标转换及大批量仿真数据渲染生成,实现仿真数据生成、参数适配与质量校验的全自动化闭环,无需人工调试干预,摆脱了对专家经验的依赖,有效降低技术落地门槛与人工运维成本。同时,由于仿真数据基于真实业务场景驱动生成,其物体姿态、相机参数等场景特征紧密贴合真实产线分布,有效减少了仿真数据与真实场景之间的域差异,从而可以在仿真数据用于模型训练时,提高通过仿真数据训练的模型在真实工业场景中分割性能的可靠性。此外,本方案基于真实的业务分割数据集对三类模型进行验证筛选,得到最优实例分割模型,能够适配各类场景变动引发的模型性能衰减,持续保障工业场景下工件分拣、缺陷检测等视觉任务中实例分割模型的高精度、高稳定性运行,实现了工业实例分割模型高效、自动化、高质量的闭环迭代优化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122618246B_ABST
    Figure CN122618246B_ABST
Patent Text Reader

Abstract

Embodiments of the application disclose a data processing method, device and equipment, a storage medium and a program product. The method comprises: obtaining business segmentation data sets generated by a scene instance segmentation model; performing zero-shot pose estimation to obtain pose parameters, screening candidate instances to obtain target instances in each business segmentation data; performing coordinate system conversion on the pose parameters of each target instance to obtain target pose parameters of each target instance; performing virtual rendering on the target instances in each business segmentation data to obtain H simulation data; performing model training on the scene instance segmentation model and an initial instance segmentation model according to the H simulation data to obtain a first instance segmentation model and a second instance segmentation model; and performing model verification on the scene instance segmentation model, the first instance segmentation model and the second instance segmentation model to obtain an optimal instance segmentation model. The model optimization efficiency can be improved by using the application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a data processing method, apparatus, device, storage medium, and program product. Background Technology

[0002] In the field of industrial automation, instance segmentation models are widely used in visual tasks such as workpiece sorting and defect detection. They need to be deployed and continuously iterated and optimized under specific operational scenarios to cope with performance degradation caused by factors such as production line adjustments, changes in incoming materials, and fluctuations in lighting. Currently, iterative optimization schemes for instance segmentation models in industrial scenarios heavily rely on manual intervention. When false detections or missed detections occur on the production line, engineers need to collect abnormal images, draw segmentation masks pixel by pixel to build a training dataset, then configure parameters to start model retraining or fine-tuning, and manually evaluate and verify the effect after training. However, in this manual annotation scheme, drawing a single frame mask takes several minutes to tens of minutes. Combined with the training, evaluation, and deployment process, a single iteration cycle often takes several days, severely impacting model iteration efficiency and constraining production line efficiency. Furthermore, differences in annotation habits and accuracy standards among different engineers lead to unstable training data quality, resulting in poor model optimization performance. Summary of the Invention

[0003] This application provides a data processing method, apparatus, device, storage medium, and program product that can improve the optimization efficiency and quality of instance segmentation models, and achieve efficient, automated, and high-quality closed-loop iterative optimization of instance segmentation models in industrial scenarios.

[0004] One embodiment of this application provides a data processing method, the method comprising: Obtain the business segmentation dataset generated by the scenario instance segmentation model during runtime in the business scenario; If there are abnormal segmentation data in the business segmentation dataset that can trigger the upgrade of the scenario instance segmentation model, then zero-shot pose estimation is performed on the candidate instances in each business segmentation dataset to obtain the pose parameters of the candidate instances in each business segmentation dataset. The candidate instances in each business segmentation dataset are then filtered according to the preset filtering strategy to obtain the target instance in each business segmentation dataset. Based on the 3D model corresponding to each target instance, the pose parameters of each target instance are transformed to obtain the target pose parameters of each target instance in the world coordinate system. For each target instance in the business segmentation data, the target pose parameters and 3D model are virtually rendered to obtain H simulation data. The simulation instance in a simulation data is generated based on the target instance corresponding to the business segmentation data of that simulation data. H is a positive integer. The scene instance segmentation model is trained using H simulation data to obtain the first instance segmentation model. The initial instance segmentation model is then trained using H simulation data to obtain the second instance segmentation model. The initial instance segmentation model refers to the instance segmentation model loaded with general pre-trained weights. The scenario instance segmentation model, the first instance segmentation model, and the second instance segmentation model are validated based on the business segmentation dataset to obtain the optimal instance segmentation model. The optimal instance segmentation model is determined based on either the first instance segmentation model or the second instance segmentation model. The optimal instance segmentation model is used to replace the scenario instance segmentation model deployed in the business scenario.

[0005] One embodiment of this application provides a data processing apparatus, the apparatus comprising: The data acquisition module is used to acquire the business segmentation dataset generated by the scenario instance segmentation model during runtime in the business scenario; The pose estimation module is used to perform zero-shot pose estimation on each candidate instance in the business segmentation dataset if there are abnormal segmentation data in the business segmentation dataset that could trigger the upgrade of the scenario instance segmentation model. This obtains the pose parameters of each candidate instance in the business segmentation dataset. The module then filters the candidate instances in each business segmentation dataset according to a preset filtering strategy to obtain the target instance in each business segmentation dataset. The coordinate system transformation module is used to transform the pose parameters of each target instance according to the 3D model corresponding to each target instance, so as to obtain the target pose parameters of each target instance in the world coordinate system. The virtual rendering module is used to virtually render the target pose parameters and 3D model corresponding to the target instance in each business segmentation data, resulting in H simulation data. The simulation instance in a simulation data is generated based on the target instance corresponding to the business segmentation data of that simulation data. H is a positive integer. The model training module is used to train the scene instance segmentation model based on H simulation data to obtain the first instance segmentation model, and to train the initial instance segmentation model based on H simulation data to obtain the second instance segmentation model; the initial instance segmentation model refers to the instance segmentation model loaded with general pre-trained weights. The model validation module is used to validate the scenario instance segmentation model, the first instance segmentation model, and the second instance segmentation model based on the business segmentation dataset to obtain the optimal instance segmentation model. The optimal instance segmentation model is determined based on the first instance segmentation model or the second instance segmentation model. The optimal instance segmentation model is used to replace the scenario instance segmentation model deployed in the business scenario.

[0006] In one optional implementation, the business segmentation dataset includes M business segmentation data points, and the M business segmentation data points include business segmentation data A. i Business data segmentation A i Includes color images and depth images; M is a positive integer, and i is a positive integer less than or equal to M; The pose estimation module is used to perform zero-shot pose estimation for each candidate instance in the service segmentation dataset. When obtaining the pose parameters of each candidate instance in the service segmentation dataset, the pose estimation module specifically performs the following operations: Obtain a generic instance segmentation model, and then segment the business data A using the generic instance segmentation model. i The color image in the image is segmented to obtain business segmentation data A. i The candidate mask for the corresponding N candidate instances; N is a positive integer; Data A is split into business segments. i The depth image in the image is back-projected to obtain the business segmentation data A. i The corresponding scene point cloud data is split according to N candidate masks to obtain the instance point cloud data corresponding to each of the N candidate masks. Based on business segmentation data A i The corresponding processing type performs pose estimation processing on the instance point cloud data corresponding to N candidate masks to obtain the pose parameters corresponding to the N candidate instances.

[0007] In one alternative implementation, business data A is split. i The corresponding processing type is regular geometry workpiece processing type; N candidate masks include candidate instance B. j The corresponding candidate mask; j is a positive integer less than or equal to N; The attitude estimation module is used to segment data A based on business needs. i The corresponding processing type involves performing pose estimation on the point cloud data of the instances corresponding to N candidate masks to obtain the pose parameters corresponding to the N candidate instances. Specifically, the pose estimation module performs the following operations: Obtain the set of candidate geometry types indicated by the regular geometry workpiece processing type; Based on candidate instance B j The instance point cloud data corresponding to the candidate mask are used to perform random sampling consistency fitting for each candidate geometry type in the candidate geometry type set to obtain candidate instance B. j The target geometry corresponding to a candidate instance; the target geometry corresponding to a candidate instance includes the candidate geometry type determined from the set of candidate geometry types, and the geometric parameters corresponding to the candidate instance; Based on candidate instance B j Based on the corresponding candidate geometry type and geometric parameters, construct candidate instance B. j The corresponding pose parameters.

[0008] In one alternative implementation, business data A is split. i The corresponding processing type is arbitrary shape workpiece processing type; N candidate masks include candidate instance B. j The corresponding candidate mask; j is a positive integer less than or equal to N; The attitude estimation module is used to segment data A based on business needs. i The corresponding processing type involves performing pose estimation on the point cloud data of the instances corresponding to N candidate masks to obtain the pose parameters corresponding to the N candidate instances. Specifically, the pose estimation module performs the following operations: Obtain the 3D models of T candidate objects from the candidate object set indicated by the workpiece processing type of arbitrary shape; T is a positive integer. Based on candidate instance B j The instance point cloud data of the corresponding candidate mask are used to perform point cloud registration on the 3D models corresponding to the T candidate objects to obtain T initial pose parameters and the confidence scores corresponding to the T initial pose parameters respectively. The initial pose parameters with the highest confidence level are selected as candidate instance B. j The corresponding candidate pose parameters; For candidate instance B j The corresponding candidate pose parameters are corrected using an iterative nearest-point algorithm to obtain candidate instance B. j The corresponding pose parameters.

[0009] In one optional implementation, the business segmentation dataset includes M business segmentation data points, and the M business segmentation data points include business segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; business segmentation data A i Includes N candidate instances, where N is a positive integer; The attitude estimation module is used to filter candidate instances in each business segment data according to a preset filtering strategy. When obtaining the target instance in each business segment data, the attitude estimation module specifically performs the following operations: From the pose parameters corresponding to N candidate instances, obtain the translation parameters corresponding to N candidate instances respectively; the translation parameters corresponding to a candidate instance include coordinate values ​​in multiple coordinate dimensions; Based on the translation parameters corresponding to the N candidate instances, determine the business segmentation data A. i Outlier detection range values ​​in each coordinate dimension; Among the N candidate instances, excluding the abnormal candidate instances, the candidate instances are determined as business segmentation data A. i The corresponding target instance; an anomaly candidate instance refers to a candidate instance in which any coordinate value in the translation parameter does not belong to the outlier detection range value in the coordinate dimension corresponding to that coordinate value.

[0010] In one optional implementation, the business segmentation dataset includes M business segmentation data points, and the M business segmentation data points include business segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; business segmentation data A i Includes N candidate instances, where N is a positive integer; The attitude estimation module is used to filter candidate instances in each business segment data according to a preset filtering strategy. When obtaining the target instance in each business segment data, the attitude estimation module specifically performs the following operations: Based on the candidate masks corresponding to the N candidate instances, the business segmentation data A is processed. i Feature extraction is performed on the color image and depth image to obtain a multi-dimensional feature vector corresponding to each candidate instance; the multi-dimensional feature vector corresponding to a candidate instance is used to indicate the multimodal attribute features that the candidate instance is associated with the color image and depth image. Obtain the feature detection model, and perform feature analysis on the multi-dimensional feature vector corresponding to each candidate instance through the feature detection model to obtain the feature analysis results corresponding to each candidate instance. Candidate instances whose feature analysis results indicate anomalies are identified as anomalous candidate instances. The candidate instances other than the anomalous candidate instances among the N candidate instances are identified as business segmentation data A. i The corresponding target instance.

[0011] In one optional implementation, the business segmentation dataset includes M business segmentation data points, and the M business segmentation data points include business segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; business segmentation data A i Includes N candidate instances, where N is a positive integer; The attitude estimation module is used to filter candidate instances in each business segment data according to a preset filtering strategy. When obtaining the target instance in each business segment data, the attitude estimation module specifically performs the following operations: Obtain the mask area of ​​the candidate mask corresponding to each of the N candidate instances, perform statistical processing on the N mask areas, and generate the lower quartile and upper quartile for the N mask areas; Using the lower quartile and upper quartile of N mask areas, a first mask area threshold for indicating the lower limit of the mask area and a second mask area threshold for indicating the upper limit of the mask area are generated. Candidate instances with a mask area smaller than the first mask area threshold or a mask area larger than the second mask area threshold are identified as abnormal candidate instances. The candidate instances other than the abnormal candidate instances among the N candidate instances are identified as business segmentation data A. i The corresponding target instance.

[0012] In one optional implementation, the coordinate system transformation module is used to perform coordinate system transformation on the pose parameters of each target instance based on the 3D model corresponding to each target instance. When obtaining the target pose parameters of each target instance in the world coordinate system, the coordinate system transformation module is specifically used to perform the following operations: Point cloud sampling is performed on the 3D model corresponding to each target instance to obtain the model point cloud data corresponding to each target instance. Based on the pose parameters of each target instance, the model point cloud data corresponding to each target instance is transformed to obtain the transformed scene point cloud data corresponding to each target instance. Based on the transformed scene point cloud data corresponding to each target instance, determine the target height value of each target instance on the target coordinate axis; the target height value of a target instance is used to indicate the bottom position of the object corresponding to that target instance; From the target height value of each target instance on the target coordinate axis, determine the camera installation height value, obtain the preset orientation adjustment parameter, and generate a coordinate system transformation matrix based on the camera installation height value and the preset orientation adjustment parameter; the camera installation height value is used to indicate the vertical distance of the virtual camera from the ground in the world coordinate system; the preset orientation adjustment parameter is used to correct the orientation of the target instance in the world coordinate system to the preset direction; The pose parameters of each target instance are transformed using the coordinate system transformation matrix to obtain the target pose parameters of each target instance in the world coordinate system.

[0013] In one optional implementation, the business segmentation dataset includes M business segmentation data points, and the M business segmentation data points include business segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; business segmentation data A i Includes R target instances; R is a positive integer; The virtual rendering module is used to virtually render the target pose parameters and 3D model corresponding to each target instance in each business segmentation data. When H simulation data are obtained, the virtual rendering module is specifically used to perform the following operations: Obtain randomized rendering parameters and segment data A according to business logic. i The target pose parameters and 3D models of R target instances are used to place R target instances in the initial virtual scene, resulting in a first virtual instance scene containing R simulation instances. Randomized rendering parameters are then used to render the first virtual instance scene, yielding business segmentation data A. i The corresponding first simulation dataset; one of the R simulation instances is determined based on one of the R target instances; Based on the target pose parameters and 3D models of D simulation instances, D simulation instances are placed in the initial virtual scene to obtain the second virtual instance scene. Randomized rendering parameters are used to render the second virtual instance scene to obtain business segmentation data A. i The corresponding second simulation dataset; the D simulation instances are obtained by randomly deleting some target instances from R target instances according to the preset instance discard probability; D is a positive integer less than R; Based on the target pose parameters and 3D models of F simulation instances, F simulation instances are placed in the initial virtual scene to obtain the third virtual instance scene. Randomized rendering parameters are used to render the third virtual instance scene to obtain the business segmentation data A. i The corresponding third simulation dataset; F simulation instances include R target instances and Q augmented instances; the Q augmented instances are generated by randomly adjusting the target pose parameters of the Q target instances among the R target instances within a preset adjustment range; Q is a positive integer less than or equal to R, and F is the sum of R and Q; When the first simulation dataset, second simulation dataset, and third simulation dataset corresponding to each business segmentation data are obtained, the first simulation dataset, second simulation dataset, and third simulation dataset corresponding to each business segmentation data are combined into H simulation datasets.

[0014] In one alternative implementation, the scene instance segmentation model includes a backbone network and a detection head network; the H simulation data include sample data and sample annotations; The model training module is used to train the scene instance segmentation model based on H simulation data. When the first instance segmentation model is obtained, the model training module is specifically used to perform the following operations: The network parameters of the backbone network in the scene instance segmentation model are frozen. H simulation data are input into the scene instance segmentation model. The backbone network with frozen network parameters is used to extract features from the sample data in the H simulation data to obtain the feature vectors corresponding to the H simulation data respectively. H feature vectors are input into the detection head network of the scene instance segmentation model to obtain the first detection result. Based on the first detection result and sample labels, the network parameters of the detection head network are adjusted until the network parameters of the detection head network converge to obtain the first instance segmentation model.

[0015] In one alternative implementation, the H simulation data include sample data and sample annotations; The model training module is used to train the initial instance segmentation model based on H simulation data. When obtaining the second instance segmentation model, the model training module is specifically used to perform the following operations: H simulation data are input into the initial instance segmentation model. The backbone network in the initial instance segmentation model is used to extract features from the sample data of the H simulation data to obtain the feature vectors corresponding to the H simulation data respectively. H feature vectors are input into the detection head network in the initial instance segmentation model to obtain the second detection result. Based on the second detection result and sample labels, the network parameters of the backbone network and the network parameters of the detection head network in the initial instance segmentation model are adjusted until the network parameters in the initial instance segmentation model converge to obtain the second instance segmentation model.

[0016] In one optional implementation, each business segmentation data in the business segmentation dataset includes a color image, a depth image, a first instance mask, a first instance category, and a first pose parameter corresponding to the business segmentation data output by the scene instance segmentation model. The model validation module is used to validate the scene instance segmentation model, the first instance segmentation model, and the second instance segmentation model based on the business segmentation dataset. When the optimal instance segmentation model is obtained, the model validation module specifically performs the following operations: Based on the target candidate mask, target instance category, and target pose parameters corresponding to the target instance in each business segmentation data, a pseudo-true annotation set corresponding to the business segmentation dataset is generated; the target candidate mask and target instance category corresponding to each target instance are determined during the zero-shot pose estimation process for each candidate instance in the business segmentation dataset. The first instance mask, first instance category, and first pose parameter corresponding to each business segmentation data output by the scene instance segmentation model are used as the first inference result. The second inference result of the first instance segmentation model for the business segmentation dataset and the third inference result of the second instance segmentation model for the business segmentation dataset are obtained. Based on the pseudo-true annotation set, the first inference result, the second inference result and the third inference result are evaluated respectively to obtain the first evaluation index, the second evaluation index, and the third evaluation index corresponding to the scene instance segmentation model. The optimal instance segmentation model is determined based on the first evaluation metric, the second evaluation metric, and the third evaluation metric.

[0017] In one optional implementation, when the model validation module determines the optimal instance segmentation model based on the first evaluation metric, the second evaluation metric, and the third evaluation metric, the model validation module specifically performs the following operations: If the second evaluation metric is better than the first evaluation metric and better than the third evaluation metric, then the first instance segmentation model is determined as the optimal instance segmentation model. If the third evaluation metric is better than both the first and second evaluation metrics, then the second instance segmentation model is determined as the optimal instance segmentation model. If the first evaluation metric is better than the second evaluation metric and the third evaluation metric, then increase the number of simulation data generated to obtain the number of updated simulation data, and generate an incremental simulation dataset based on the number of updated simulation data. The first instance segmentation model and the second instance segmentation model are trained using an incremental simulation dataset until the second evaluation index corresponding to the trained first instance segmentation model or the third evaluation index corresponding to the trained second instance segmentation model is better than the first evaluation index. The trained first instance segmentation model or the trained second instance segmentation model is then determined as the optimal instance segmentation model.

[0018] In one optional implementation, the business segmentation dataset includes normal segmentation data output by the scenario instance segmentation model deployed in the business scenario, and abnormal segmentation data indicated by the scenario instance segmentation model reaching an abnormal triggering condition; the abnormal triggering condition includes at least one of the following: the number of instances detected by the scenario instance segmentation model in the business scenario is less than a preset number threshold, the area deviation between the instance mask area detected by the scenario instance segmentation model and the historical average mask area is greater than a preset deviation threshold, and the instance pose parameters detected by the scenario instance segmentation model indicate that the instance is in an abnormal capture state.

[0019] One embodiment of this application provides a computer device, including a processor, a memory, and an input / output interface; The processor is connected to a memory and an input / output interface, respectively. The input / output interface is used to receive and output data, the memory is used to store computer programs, and the processor is used to call the computer programs so that the computer device containing the processor executes the method in one aspect of the embodiments of this application.

[0020] One aspect of this application provides a computer-readable storage medium storing a computer program adapted to be loaded and executed by a processor, so that a computer device having the processor performs the method of one aspect of this application.

[0021] One aspect of this application provides a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the methods provided in various optional embodiments of this application. In other words, when the computer program is executed by the processor, it implements the methods provided in various optional embodiments of this application.

[0022] Implementing the embodiments of this application will have the following beneficial effects: In this embodiment, a business segmentation dataset generated by the scene instance segmentation model during runtime is obtained. If abnormal segmentation data exists in the business segmentation dataset that could trigger an upgrade of the scene instance segmentation model, zero-shot pose estimation is performed on each candidate instance in the business segmentation dataset to obtain the pose parameters of each candidate instance. The candidate instances are then filtered according to a preset filtering strategy to obtain the target instance in each business segmentation dataset. Based on the 3D model corresponding to each target instance, coordinate system transformation is performed on the pose parameters of each target instance to obtain the target pose parameters of each target instance in the world coordinate system. Virtual rendering is then performed on the target pose parameters and 3D model corresponding to each target instance in the business segmentation dataset. H simulation data points are obtained. A simulation instance in a simulation data point is generated based on the target instance corresponding to the business segmentation data corresponding to that simulation data. H is a positive integer. The scenario instance segmentation model is trained based on the H simulation data points to obtain the first instance segmentation model. The initial instance segmentation model is trained based on the H simulation data points to obtain the second instance segmentation model. The initial instance segmentation model refers to the instance segmentation model loaded with general pre-trained weights. The scenario instance segmentation model, the first instance segmentation model, and the second instance segmentation model are validated based on the business segmentation dataset to obtain the optimal instance segmentation model. The optimal instance segmentation model is determined based on the first instance segmentation model or the second instance segmentation model. The optimal instance segmentation model is used to replace the scenario instance segmentation model deployed in the business scenario. Through the above process, business segmentation datasets generated by the scene instance segmentation model in business scenarios were collected. Abnormal segmentation data that leads to model performance degradation was accurately captured and used as the core data foundation for model iteration and optimization. Relying on automated processes such as zero-shot pose estimation, instance selection, coordinate system transformation, and virtual rendering, fully automated generation of simulation annotation data was achieved, eliminating the need for manual pixel-by-pixel annotation. This significantly shortened the model iteration cycle, improved production line efficiency, avoided subjective errors from manual annotation, unified data generation standards, and stably ensured the quality of training data, effectively improving model optimization results. Furthermore, it can automatically complete candidate instance pose calculation, accurate selection, coordinate transformation, and large-scale simulation data rendering and generation, achieving a fully automated closed loop of simulation data generation, parameter adaptation, and quality verification. No manual debugging intervention is required, eliminating reliance on expert experience and effectively reducing the technical implementation threshold and manual maintenance costs. Simultaneously, because the simulation data is generated based on real business scenarios, its object poses, camera parameters, and other scene features closely match the distribution of real production lines, effectively reducing the domain differences between simulation data and real scenes. Therefore, when using simulation data for model training, the reliability of the segmentation performance of models trained on simulation data in real industrial scenarios can be improved.Furthermore, this solution validates and filters three types of models based on real business segmentation datasets to obtain the optimal instance segmentation model. It can adapt to the performance degradation of the model caused by various scenario changes, and continuously ensure the high-precision and high-stability operation of the instance segmentation model in visual tasks such as workpiece sorting and defect detection in industrial scenarios. It realizes efficient, automated and high-quality closed-loop iterative optimization of industrial instance segmentation models. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a network interaction architecture diagram provided in an embodiment of this application; Figure 2 This is a schematic diagram of a data processing method provided in an embodiment of this application; Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 1 ; Figure 4 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 2 ; Figure 5 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 3 ; Figure 6 This is a schematic diagram of a data processing device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0026] If this application requires the collection of object data (such as user data), a prompt interface or pop-up window will be displayed before and during the collection process. This prompt interface or pop-up window is used to inform the user that certain data is being collected. The data acquisition steps will only begin after the user confirms the prompt interface or pop-up window; otherwise, the process will end. Furthermore, the acquired user data will be used in reasonable and legal scenarios or for legitimate purposes. Optionally, in scenarios where user data needs to be used but user authorization has not been obtained, authorization can be requested from the user, and the user data can only be used after authorization is granted.

[0027] It is understood that, in the specific embodiments of this application, the user data involved requires user permission or consent when the following embodiments of this application are applied to specific products or technologies, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant regions.

[0028] In the embodiments of this application, please refer to Figure 1 , Figure 1 This is a network interaction architecture diagram provided in an embodiment of this application, such as... Figure 1 As shown, the system includes a server 101 and a cluster of service devices. The cluster of service devices may include service devices 102a, 102b, 102c, ..., 102n. Communication connections can exist between the service devices in the cluster; for example, there is a communication connection between service devices 102a and 102b, and between service devices 102a and 102c. Simultaneously, any service device in the cluster can have a communication connection with the server 101; for example, there is a communication connection between service device 102a and server 101. Optionally, the communication connection between service device 102a and any other service device (e.g., service device 102b) can be achieved through a communication connection between service device 102a and server 101, and a communication connection between server 101 and service device 102b. The communication connection method is not limited; it can be established directly or indirectly through wired communication, wireless communication, or other methods. This application does not impose any restrictions on this method.

[0029] It should be understood that this data processing method can be executed jointly by any business device and server 101. Any business device can be deployed with a scene instance segmentation model, which can then be used to perform instance segmentation and other processing on objects such as workpieces in the business scene. The following explanation uses the example of this data processing method being executed jointly by business device 102n and server 101. Business device 102n can use the scene instance segmentation model to perform instance segmentation processing on the scene image obtained from the business scene, obtaining data such as instance masks and pose parameters. This allows grasping devices (such as industrial robots or robotic arms) connected to business device 102n to grasp objects in the business scene using the pose parameters.

[0030] When the business device 102n encounters anomalies such as missed detection (no target detected) or false detection (incorrect detection of a non-target) during instance segmentation using the scene instance segmentation model, the business device 102n can obtain the abnormal segmentation data generated under abnormal conditions and obtain a portion of the normal segmentation data generated during the runtime of the scene instance segmentation model. The normal segmentation data and the abnormal segmentation data are combined to form the business segmentation dataset generated during the runtime of the scene instance segmentation model in the business scenario, and the business segmentation dataset is sent to the server 101.

[0031] After receiving the business segmentation dataset, server 101 can perform zero-shot pose estimation on each candidate instance in the business segmentation dataset to obtain the pose parameters of each candidate instance. It then filters the candidate instances according to a preset filtering strategy to obtain the target instance in each business segmentation dataset. Based on the pose parameters and 3D model corresponding to each target instance in the business segmentation dataset, it performs virtual rendering to obtain H simulation data sets, where H is a positive integer. Server 101 can train a scene instance segmentation model using the H simulation data sets to obtain a first instance segmentation model. It can also train an initial instance segmentation model using the H simulation data sets to obtain a second instance segmentation model. The initial instance segmentation model refers to an instance segmentation model loaded with general pre-trained weights. Server 101 can perform model validation on the scene instance segmentation model, the first instance segmentation model, and the second instance segmentation model using the business segmentation dataset to obtain an optimal instance segmentation model. The optimal instance segmentation model is determined based on either the first or second instance segmentation model; the optimal instance segmentation model is used to replace the scene instance segmentation model deployed in the business scenario.

[0032] Server 101 can store the optimal instance partitioning model in a specified storage location. Business device 102n can periodically poll the specified storage location of server 101. When the optimal instance partitioning model is detected in the specified storage location, it will automatically download and replace the scene instance partitioning model deployed in the business scenario, thus completing hot deployment.

[0033] It should be noted that the scene instance segmentation model refers to the instance segmentation model currently deployed in actual business scenarios (such as industrial scenarios) and running online. It is a practical model providing inference services for actual visual tasks such as workpiece sorting and defect detection. For example, it can identify and segment specific workpieces in the business scenario, possessing targeted scene segmentation capabilities. The initial instance segmentation model, on the other hand, is a basic model that has not been adapted to the current business scenario. It can be a native instance segmentation model that only loads general pre-trained weights, lacking iterative training experience for this industrial business scenario and lacking targeted scene segmentation capabilities. For example, an instance segmentation model trained only on a general training set can identify and segment objects in the general training set, but lacks experience in segmenting workpieces in the current business scenario. An instance segmentation model refers to a deep learning model capable of pixel-level classification and segmentation of each target object in an image, such as Mask R-CNN, YOLACT, and SAM1-SAM3.

[0034] It should be noted that pose parameters are used to indicate the 3D position and spatial orientation information of candidate instances in the business scenario within the business segmentation data. In other words, they characterize the complete spatial orientation information of the candidate instance in the business scenario coordinate system. Specifically, the spatial orientation information can include 3D translation components and 3D rotation components, which together constitute the six-degree-of-freedom pose description of the target instance. The 3D translation components correspond to the coordinate offsets in the three orthogonal directions of the X, Y, and Z axes, used to determine the specific 3D position of the candidate instance's reference point within the business scenario space. The 3D rotation components characterize the deflection angles of the target instance around the X, Y, and Z axes, reflecting the candidate instance's own orientation posture (i.e., spatial orientation information) within the business scenario space.

[0035] It should be noted that a 3D model can refer to a 3D digital entity model of the object corresponding to the target instance. It contains accurate 3D information such as the object's actual size, shape, surface texture, and spatial topology. It is the core basic carrier for realizing zero-sample pose calculation of instances, cross-coordinate system pose transformation, and realistic scene simulation rendering.

[0036] It should be noted that any business device in the business device cluster can be an onboard device installed on an industrial robot in a business scenario (such as an industrial scenario), or a terminal device that is connected to the industrial robot. That is, it can perform instance segmentation on scene images in the business scenario through a scene instance segmentation model to obtain data such as instance mask, instance category, confidence level and instance pose parameters corresponding to candidate instances, and can control the industrial robot to perform grasping operations through instance pose parameters.

[0037] It is understood that the business equipment and server mentioned in the embodiments of this application can also be a type of computer equipment. Specifically, the business equipment mentioned above can be an electronic device, including but not limited to mobile phones, tablets, desktop computers, laptops, handheld computers, augmented reality / virtual reality (AR / VR) devices, and other mobile internet devices (MIDs) with network access capabilities. Figure 1 As shown, the business device can be a mobile phone (as shown in business device 102a), a desktop computer (as shown in business device 102b), a tablet computer (as shown in business device 102c), or an onboard device for an industrial robot (as shown in business device 102n), etc. Figure 1 Only a portion of the devices are listed. The servers mentioned above can be standalone physical servers, server clusters consisting of multiple physical servers, or distributed systems.

[0038] For details, please see Figure 2 , Figure 2 This is a schematic diagram illustrating a data processing method provided in an embodiment of this application. For example... Figure 2 As shown, service device 2A can segment the scene images (including scene images 201a, ..., scene images 201b, etc.) corresponding to service scene 201 using the deployed scene instance segmentation model 202, obtaining service segmentation data corresponding to scene image 201a and scene image 201b. Now, assuming that service device 2A detects a missed detection or false detection when scene instance segmentation model 202 segments scene image 201b, service device 2A can identify the service segmentation data corresponding to scene image 201b as abnormal segmentation data. Service device 2A can connect the abnormal segmentation data with some normal segmentation data (such as the service segmentation data corresponding to scene image 201a) to form a service segmentation dataset 203. Service device 2A can then upload this service segmentation dataset 203 to server 2B.

[0039] It should be noted that scene images may include, but are not limited to, the color images and depth images (such as RGB-D images) corresponding to the scene; a business segmentation dataset may include, but is not limited to, the color images and depth images corresponding to a scene, and the instance masks, instance categories, confidence scores, and instance pose parameters output by the scene instance segmentation model 202 for the color and depth images. The number of abnormal segmentation data in the business segmentation dataset 203 may be one or more, and the ratio of abnormal segmentation data to normal segmentation data may be 1:1.

[0040] Server 2B can obtain the business segmentation dataset 203 generated by the scene instance segmentation model 202 during runtime. If there is abnormal segmentation data in the business segmentation dataset 203 that could trigger an upgrade of the scene instance segmentation model 202, then Server 2B can perform zero-shot pose estimation on each candidate instance in the business segmentation dataset 203 to obtain the pose parameters of each candidate instance. Server 2B can filter the candidate instances in each business segmentation dataset according to a preset filtering strategy to obtain the target instance in each business segmentation dataset. Server 2B can perform coordinate system transformation on the pose parameters of each target instance according to the 3D model corresponding to each target instance to obtain the target pose parameters of each target instance located in the world coordinate system.

[0041] Furthermore, server 2B can perform virtual rendering of the target pose parameters and 3D model corresponding to the target instance in each business segmentation data, obtaining H simulation data. Here, a simulation instance in one simulation data is generated based on the target instance corresponding to the business segmentation data corresponding to that simulation data; H is a positive integer. Server 2B can train the scene instance segmentation model 202 based on the H simulation data to obtain the first instance segmentation model 204; server 2B can also train the initial instance segmentation model 205 based on the H simulation data to obtain the second instance segmentation model 206.

[0042] It should be noted that the initial instance segmentation model 205 can refer to an instance segmentation model loaded with general pre-trained weights, or the initial instance segmentation model 205 can be the predecessor of the scene instance segmentation model 202. That is, the scene instance segmentation model 202 can be obtained by training the initial instance segmentation model 205 with training samples generated from data in the business scenario.

[0043] Server 2B can perform model validation on scene instance segmentation model 202, first instance segmentation model 204 and second instance segmentation model 206 respectively based on business segmentation dataset 203 to obtain the optimal instance segmentation model 207.

[0044] The optimal instance segmentation model 207 is determined based on the first instance segmentation model 204 or the second instance segmentation model 206; the optimal instance segmentation model 207 is used to replace the scenario instance segmentation model deployed in the business scenario.

[0045] It should be noted that during the model validation process, server 2B can evaluate the model performance of the three instance segmentation models, namely scene instance segmentation model 202, first instance segmentation model 204, and second instance segmentation model 206, using business segmentation dataset 203 and pseudo-real annotation set. If the model performance of first instance segmentation model 204 or second instance segmentation model 206 is better than the three instance segmentation models, then server 2B can determine first instance segmentation model 204 or second instance segmentation model 206 as the optimal instance segmentation model 207.

[0046] If the performance of scene instance segmentation model 202 is better than that of the three instance segmentation models, it can be considered that the current model training cannot improve the model quality of the instance segmentation model. Server 2B can return to the step of generating simulation data, increase the amount of simulation data generated, generate more simulation data, and obtain incremental simulation data. The first instance segmentation model 204 and the second instance segmentation model 206 are trained again using the incremental simulation data until the performance of the first instance segmentation model 204 or the second instance segmentation model 206 is better than that of scene instance segmentation model 202. At this point, server 2B can determine the first instance segmentation model 204 or the second instance segmentation model 206 as the optimal instance segmentation model 207.

[0047] Furthermore, server 2B can store the optimal instance segmentation model 207 in a designated storage location. Business device 2A can periodically poll the designated storage location on server 2B. Upon detecting the optimal instance segmentation model 207 in the designated storage location, it automatically downloads and replaces the scene instance segmentation model 202 deployed in the business scenario, completing the model deployment. In other words, business device 2A can subsequently use the deployed optimal instance segmentation model 207 to segment newly generated scene images in business scenario 201, obtaining new business segmentation data.

[0048] It should be understood that in the subsequent instance segmentation process, when the business device 2A encounters anomalies such as missed detection or false detection during instance segmentation processing using the optimal instance segmentation model 207, the business device 2A and the server 2B can adopt the same execution process described above to continue iteratively optimizing the optimal instance segmentation model 207.

[0049] Through the above process, business segmentation datasets generated by the scene instance segmentation model in business scenarios were collected. Abnormal segmentation data that leads to model performance degradation was accurately captured and used as the core data foundation for model iteration and optimization. Relying on automated processes such as zero-shot pose estimation, instance selection, coordinate system transformation, and virtual rendering, fully automated generation of simulation annotation data was achieved, eliminating the need for manual pixel-by-pixel annotation. This significantly shortened the model iteration cycle, improved production line efficiency, avoided subjective errors from manual annotation, unified data generation standards, and stably ensured the quality of training data, effectively improving model optimization results. Furthermore, it can automatically complete candidate instance pose calculation, accurate selection, coordinate transformation, and large-scale simulation data rendering and generation, achieving a fully automated closed loop of simulation data generation, parameter adaptation, and quality verification. No manual debugging intervention is required, eliminating reliance on expert experience and effectively reducing the technical implementation threshold and manual maintenance costs. Simultaneously, because the simulation data is generated based on real business scenarios, its object poses, camera parameters, and other scene features closely match the distribution of real production lines, effectively reducing the domain differences between simulation data and real scenes. Therefore, when using simulation data for model training, the reliability of the segmentation performance of models trained on simulation data in real industrial scenarios can be improved. Furthermore, this solution validates and filters three types of models based on real business segmentation datasets to obtain the optimal instance segmentation model. It can adapt to the performance degradation of the model caused by various scenario changes, and continuously ensure the high-precision and high-stability operation of the instance segmentation model in visual tasks such as workpiece sorting and defect detection in industrial scenarios. It realizes efficient, automated and high-quality closed-loop iterative optimization of industrial instance segmentation models.

[0050] Further, please see Figure 3 , Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 1 This data processing method can be executed by a server. The data processing method may include at least the following steps S301-S306: Step S301: Obtain the business segmentation dataset generated by the scenario instance segmentation model during runtime in the business scenario.

[0051] In this embodiment, the server can obtain the business segmentation dataset generated by the scene instance segmentation model during runtime in the business scenario. The business segmentation dataset includes normal segmentation data output by the scene instance segmentation model deployed in the business scenario, and abnormal segmentation data indicated by the scene instance segmentation model reaching an abnormal triggering condition. The abnormal triggering condition includes at least one of the following: the number of instances detected by the scene instance segmentation model in the business scenario is less than a preset number threshold; the area deviation between the instance mask area detected by the scene instance segmentation model and the historical average mask area is greater than a preset deviation threshold; or the instance pose parameters detected by the scene instance segmentation model indicate that the instance is in an abnormal capture state.

[0052] Specifically, the business segmentation dataset can be generated by business devices deployed with the scene instance segmentation model during instance segmentation in a business scenario, and these devices can upload the dataset to a server. During instance segmentation, the business device can obtain a preset threshold. When the business device segments a scene image in the business scenario using the scene instance segmentation model, it obtains the instance mask, instance category, confidence score, and instance pose parameters corresponding to the instances in the scene image. If the business device detects that the number of instances corresponding to the instances in the scene image is less than the preset threshold (i.e., a missed detection), the business device can determine that the scene instance segmentation model has reached an abnormal trigger condition. Alternatively, the business device can obtain the instance mask area corresponding to each instance, a preset deviation threshold, and the historical average mask area corresponding to the business scenario. The business device can calculate the deviation between the instance mask area corresponding to each instance mask and the historical average mask area. If the deviation between the instance mask area corresponding to one or more instance masks in the scene image and the historical average mask area is greater than the preset deviation threshold (i.e., a false detection), the business device can determine that the scene instance segmentation model has reached an abnormal trigger condition.

[0053] The preset quantity threshold is used to characterize the minimum number of instances that the business device expects to detect when performing instance segmentation on a scene image in a business scenario using the scene instance segmentation model.

[0054] Taking an instance mask as an example, the service device can determine the absolute value of the difference between the instance mask area corresponding to the instance mask and the historical average mask area as the mask area difference, and determine the ratio between the mask area difference and the historical average mask area as the area deviation corresponding to the instance mask.

[0055] Optionally, the business device can determine whether the current instance can be grasped normally based on the instance pose parameters of each instance. Taking an instance as an example, the business device can perform collision detection and reachability analysis on the instance pose parameters of the instance, and then determine whether the current posture meets the preset grasping constraints. If the preset grasping constraints are not met, it can be determined that the instance pose parameters of the instance indicate that the instance is in an abnormal grasping state, and thus it can be determined that the scene instance segmentation model has reached the abnormal triggering condition.

[0056] Optionally, to determine whether the scene instance segmentation model has reached the abnormal triggering condition, it is also possible to dynamically determine whether the segmentation result of the current scene image processed by the scene instance segmentation model is abnormal based on the feature distribution of historical normal data, using an online learning anomaly detector (such as isolated forest, autoencoder reconstruction error, etc.).

[0057] It should be noted that when the scene instance segmentation model reaches an abnormal trigger condition, the business segmentation data generated by the scene instance segmentation model can be considered abnormal segmentation data. In this case, the scene instance segmentation model can be considered to need iterative upgrades to obtain an instance segmentation model with better performance. At this time, the business device can combine the above-mentioned abnormal segmentation data with the normal segmentation data output by the scene instance segmentation model to form a business segmentation dataset.

[0058] It should be noted that a business segmentation dataset may include, but is not limited to, a color image, a depth image, and instance masks, instance categories, confidence scores, and instance pose parameters output by the scene instance segmentation model for a given color image and depth image within a business scenario. The business segmentation dataset may contain one or more anomalous segmentation data points, and the ratio of anomalous to normal segmentation data points may be 1:1.

[0059] The business device can send the business segmentation dataset to the server, such as uploading it to a fixed directory on the server. At this time, the server can obtain the business segmentation dataset generated by the scenario instance segmentation model in the business scenario when it runs from the fixed directory.

[0060] Step S302: If there is abnormal segmentation data in the business segmentation dataset that can trigger the upgrade of the scenario instance segmentation model, then zero-shot pose estimation is performed on the candidate instances in each business segmentation dataset to obtain the pose parameters of the candidate instances in each business segmentation dataset. The candidate instances in each business segmentation dataset are then filtered according to the preset filtering strategy to obtain the target instance in each business segmentation dataset.

[0061] In this embodiment, if there is abnormal segmentation data in the business segmentation dataset that could trigger an upgrade of the scene instance segmentation model, the server can perform instance segmentation processing on the color image in each business segmentation data in the business segmentation dataset using a general instance segmentation model to obtain a candidate mask for each candidate instance in the business segmentation data. Furthermore, the server can perform back-projection processing on the depth image in each business segmentation data to obtain the scene point cloud data corresponding to each business segmentation data. For each candidate instance, the server can extract the instance point cloud data corresponding to that candidate instance from the scene point cloud data corresponding to that candidate instance based on the candidate mask of that candidate instance.

[0062] Furthermore, the server can perform pose estimation processing on the instance point cloud data corresponding to the candidate mask of each candidate instance to obtain the pose parameters corresponding to each candidate instance.

[0063] The server can filter candidate instances in each business segment data according to a preset filtering strategy to obtain the target instance in each business segment data.

[0064] It should be noted that zero-shot pose estimation refers to the process of directly calculating the rotation angle and translation vector (i.e., pose parameters) of any newly emerging instance (candidate instance) of an unknown category in 3D space in real time by analyzing its corresponding instance point cloud data, without prior training for a specific object category. No model training or fine-tuning is required; the general instance segmentation model in this scheme does not need to be retrained for new objects.

[0065] The preset filtering strategies can include, but are not limited to, spatial distribution filtering strategies, multimodal feature distribution filtering strategies, and mask area distribution filtering strategies. Spatial distribution filtering strategies refer to filtering strategies that, based on the three-dimensional spatial position distribution of candidate instances in the camera coordinate system (such as center point coordinates, or translation parameters in pose parameters), use statistical mean and standard deviation to set a reasonable range and eliminate outlier instances that deviate from the overall distribution. Multimodal feature distribution filtering strategies refer to filtering strategies that fuse multimodal information such as depth images, color images (RGB images), and candidate masks to extract features such as depth continuity, edge alignment, and texture consistency of the instance region of candidate instances, construct multidimensional feature vectors, and then use feature detection models (such as isolated forests, One-Class SVMs, etc.) to identify and eliminate candidate instances with abnormal feature combinations. Mask area distribution filtering strategies refer to statistically calculating the mask area of ​​candidate instances, calculating the median and interquartile range (IQR) of the mask area, and filtering out outlier instances with a mask area less than (Q1...). A screening strategy is used to eliminate candidate instances with a threshold of 1.5 × IQR or greater than (Q3 + 1.5 × IQR). Here, Q1 is the lower quartile, Q3 is the upper quartile, and IQR (Interquartile Range) indicates the distribution range of the middle 50% of the data, reflecting the dispersion of the data. IQR is calculated as IQR = Q3 - Q1. The specific screening process for the three preset screening strategies described above can be found in the relevant descriptions below.

[0066] Step S303: Based on the 3D model corresponding to each target instance, perform coordinate system transformation on the pose parameters of each target instance to obtain the target pose parameters of each target instance in the world coordinate system.

[0067] In this embodiment of the application, the server can convert the pose parameters of each target instance to the world coordinate system to obtain the target pose parameters of each target instance located in the world coordinate system.

[0068] Specifically, the server can perform point cloud sampling on the 3D model corresponding to each target instance to obtain the model point cloud data for each target instance. Alternatively, the server can pre-store the model point cloud data of the 3D model corresponding to each target instance, and the server can directly obtain the model point cloud data corresponding to each target instance.

[0069] Among them, the 3D model is a digital stereoscopic model that fits the structure of the real target object, and the pose parameters of the target instance are the original position and attitude data of the target instance without ground calibration and orientation correction.

[0070] The server can reconstruct the point cloud from the 3D models of each target instance. Using the pose parameters of each target instance and its corresponding model point cloud data, it can perform coordinate transformation and restoration of the virtual scene point cloud, obtaining transformed scene point cloud data for each target instance. Furthermore, the server can accurately calculate the bottom position of each target instance, thereby determining the scene ground height calibration and the camera installation height of the virtual camera. Simultaneously, based on preset orientation adjustment parameters, it performs orientation deviation correction of the global coordinate system, generating a precise coordinate system transformation matrix adapted to the current business scenario, possessing both height calibration and orientation correction functions. Based on this transformation matrix, the server can perform coordinate system transformation on the pose parameters of all target instances, obtaining high-precision target pose parameters for each target instance in the world coordinate system.

[0071] The server can store the camera mounting height value corresponding to each business segment and the target pose parameters of the target instance in the world coordinate system in each business segment as a scene description file for subsequent virtual rendering. Among them, the camera mounting height value is used to indicate the camera position setting of the virtual camera during subsequent virtual rendering.

[0072] Step S304: Perform virtual rendering on the target pose parameters and 3D model corresponding to the target instance in each business segmentation data to obtain H simulation data; the simulation instance in a simulation data is generated based on the target instance corresponding to the business segmentation data corresponding to that simulation data; H is a positive integer.

[0073] In this embodiment of the application, the server can load the above-mentioned scene description file through a virtual simulation engine, and perform virtual rendering on the target pose parameters and 3D model corresponding to the target instance in each business segmentation data according to the following three types of rendering methods, to obtain H simulation data.

[0074] Virtual simulation engines refer to software systems used to generate realistic rendered images and labeled data, such as Dexverse, NVIDIA Isaac Sim, and Blender, which support pose reproduction and randomization augmentation based on real-world scenes.

[0075] Taking a business segmentation dataset as an example, the server can perform a first-type rendering method (i.e., reproducing the business scenario) on each target instance in the business segmentation dataset to obtain the first simulation dataset corresponding to the business segmentation dataset. That is, the server can place all target instances in the business segmentation dataset in the initial virtual scene according to the target pose parameters and 3D model of each target instance in the business segmentation dataset to obtain the first virtual instance scene of all target instances. The server then uses randomized rendering parameters to render the first virtual instance scene to obtain the first simulation dataset corresponding to the business segmentation dataset.

[0076] Randomized rendering parameters refer to different parameters such as illumination intensity / color, texture mapping, camera noise, and depth sensor error. The first simulation dataset can include multiple first simulation datasets. The simulation instances in each first simulation dataset are identical, and the only difference between each first simulation dataset is at least one of the following parameters: illumination intensity / color, texture mapping, camera noise, or depth sensor error. A simulation instance in a first simulation dataset can be a target instance in the business segmentation data. For example, if the business segmentation data includes target instance 1 and target instance 2, any first simulation dataset corresponding to the business segmentation data includes simulation instance 1 and simulation instance 2, where simulation instance 1 is identical to target instance 1, and simulation instance 2 is identical to target instance 2.

[0077] Furthermore, the server can perform a second type of rendering on each target instance in the segmented data of the service (i.e., randomly discard some target instances to simulate occlusion or missed detection in the actual scene) to obtain the second simulation dataset corresponding to the segmented data of the service.

[0078] The second simulation dataset may include multiple second simulation datasets. The simulation instances in each second simulation dataset may be the same or different, but each simulation instance in the second simulation dataset originates from the target instance in the business segmentation data. For example, if the business segmentation data includes target instance 1, target instance 2, and target instance 3, the corresponding second simulation dataset includes second simulation data a and second simulation data b. Second simulation data a may include simulation instance 1 and simulation instance 2, and second simulation data b may include simulation instance 1 and simulation instance 3. Alternatively, second simulation data a may include simulation instance 1 and simulation instance 2, and second simulation data b may include simulation instance 1 and simulation instance 2, in which case at least one parameter among illumination intensity / color, texture mapping, camera noise, or depth sensor error differs between second simulation data a and second simulation data b. Simulation instance 1 is the same as target instance 1, simulation instance 2 is the same as target instance 2, and simulation instance 3 is the same as target instance 3.

[0079] Furthermore, the server can perform a third type of rendering on each target instance in the business segmentation data (i.e., randomly expand the target instance to simulate a scenario where similar workpieces are densely stacked) to obtain the third simulation dataset corresponding to the business segmentation data.

[0080] The third simulation dataset can include multiple third simulation datasets. The simulation instances in each third simulation dataset can be the same or different, but each simulation instance in the third simulation dataset originates from the target instance in the service segmentation data, or is obtained by fine-tuning the target instance in the service segmentation data. For example, if the service segmentation data includes target instance 1 and target instance 2, one third simulation dataset in the corresponding third simulation dataset can include simulation instance 1, simulation instance 2, and simulation instance 3. Simulation instance 1 is the same as target instance 1, simulation instance 2 is the same as target instance 2, and simulation instance 3 can be obtained by fine-tuning either target instance 1 or target instance 2. For example, by subjecting the target pose parameters of target instance 1 to a small-range random perturbation (e.g., translation ±5cm, rotation ±15°), simulation instance 3 is obtained.

[0081] When the server obtains the first simulation dataset, the second simulation dataset, and the third simulation dataset corresponding to each business segmentation data, it can combine the first simulation dataset, the second simulation dataset, and the third simulation dataset corresponding to each business segmentation data into H simulation datasets.

[0082] Step S305: Train the scene instance segmentation model based on H simulation data to obtain the first instance segmentation model; train the initial instance segmentation model based on H simulation data to obtain the second instance segmentation model; the initial instance segmentation model refers to the instance segmentation model loaded with general pre-trained weights.

[0083] In this embodiment, the scene instance segmentation model includes a backbone network and a detection head network. The server can adjust the network parameters of the detection head network in the scene instance segmentation model based on H simulation data until the network parameters of the detection head network converge to obtain the first instance segmentation model.

[0084] The server can adjust the network parameters of the backbone network and the detection head network in the initial instance segmentation model based on H simulation data until the network parameters in the initial instance segmentation model converge to obtain the second instance segmentation model.

[0085] The initial instance segmentation model refers to the instance segmentation model loaded with general pre-trained weights.

[0086] Step S306: Based on the business segmentation dataset, perform model validation on the scenario instance segmentation model, the first instance segmentation model, and the second instance segmentation model to obtain the optimal instance segmentation model; the optimal instance segmentation model is determined based on the first instance segmentation model or the second instance segmentation model; the optimal instance segmentation model is used to replace the scenario instance segmentation model deployed in the business scenario.

[0087] In this embodiment of the application, the server can perform model validation on the scene instance segmentation model, the first instance segmentation model and the second instance segmentation model based on the business segmentation dataset to obtain the optimal instance segmentation model.

[0088] The optimal instance segmentation model is determined based on either the first instance segmentation model or the second instance segmentation model; the optimal instance segmentation model is used to replace the scenario instance segmentation model deployed in the business scenario.

[0089] Specifically, during the model validation process, the server can determine the first evaluation metric corresponding to the scene instance segmentation model, the second evaluation metric corresponding to the first instance segmentation model, and the third evaluation metric corresponding to the second instance segmentation model through the business segmentation dataset and the pseudo-real annotation set.

[0090] If the first evaluation metric, the second evaluation metric, and the third evaluation metric indicate that the first instance segmentation model or the second instance segmentation model is the best performing instance segmentation model among the three instance segmentation models, then the server can determine the first instance segmentation model or the second instance segmentation model as the optimal instance segmentation model.

[0091] Optionally, if the first evaluation metric, the second evaluation metric, and the third evaluation metric indicate that the scene instance segmentation model is the instance segmentation model with the best performance among the three instance segmentation models, the server can increase the number of simulation data generated above, obtain an incremental simulation dataset, and use the incremental simulation dataset to train the first instance segmentation model and the second instance segmentation model respectively, until the first instance segmentation model or the second instance segmentation model trained is the instance segmentation model with the best performance among the three instance segmentation models. The first instance segmentation model or the second instance segmentation model trained is then determined as the optimal instance segmentation model.

[0092] Furthermore, the server can store the optimal instance segmentation model in a designated storage location. Business devices can periodically poll the server's designated storage location. When the optimal instance segmentation model is detected in the designated storage location, it will automatically download and replace the scene instance segmentation model deployed in the business scenario, thus completing the model deployment. In other words, subsequent business devices can use the deployed optimal instance segmentation model to segment newly generated scene images in the business scenario to obtain new business segmentation data.

[0093] Through the above process, business segmentation datasets generated by the scene instance segmentation model in business scenarios were collected. Abnormal segmentation data that leads to model performance degradation was accurately captured and used as the core data foundation for model iteration and optimization. Relying on automated processes such as zero-shot pose estimation, instance selection, coordinate system transformation, and virtual rendering, fully automated generation of simulation annotation data was achieved, eliminating the need for manual pixel-by-pixel annotation. This significantly shortened the model iteration cycle, improved production line efficiency, avoided subjective errors from manual annotation, unified data generation standards, and stably ensured the quality of training data, effectively improving model optimization results. Furthermore, it can automatically complete candidate instance pose calculation, accurate selection, coordinate transformation, and large-scale simulation data rendering and generation, achieving a fully automated closed loop of simulation data generation, parameter adaptation, and quality verification. No manual debugging intervention is required, eliminating reliance on expert experience and effectively reducing the technical implementation threshold and manual maintenance costs. Simultaneously, because the simulation data is generated based on real business scenarios, its object poses, camera parameters, and other scene features closely match the distribution of real production lines, effectively reducing the domain differences between simulation data and real scenes. Therefore, when using simulation data for model training, the reliability of the segmentation performance of models trained on simulation data in real industrial scenarios can be improved. Furthermore, this solution validates and filters three types of models based on real business segmentation datasets to obtain the optimal instance segmentation model. It can adapt to the performance degradation of the model caused by various scenario changes, and continuously ensure the high-precision and high-stability operation of the instance segmentation model in visual tasks such as workpiece sorting and defect detection in industrial scenarios. It realizes efficient, automated and high-quality closed-loop iterative optimization of industrial instance segmentation models.

[0094] Further, please see Figure 4 , Figure 4 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 2 This data processing method can be executed by a server. The data processing method may include at least the following steps S401-S409: Step S401: Obtain the business segmentation dataset generated by the scenario instance segmentation model during runtime in the business scenario.

[0095] In this embodiment, the specific implementation steps of step S401 can be found above. Figure 3 The specific details of step S01 in the embodiments will not be repeated here.

[0096] It should be noted that after the server obtains the business segmentation dataset generated by the scenario instance segmentation model during runtime in the business scenario, it can generate a trigger file based on the business segmentation dataset (including abnormal segmentation data and normal segmentation data), the scenario instance segmentation model, and the configuration file for subsequent model training of the scenario instance segmentation model. This trigger file is used to trigger the model iteration optimization process of subsequent steps S402-S409. In other words, the subsequent steps are carried out based on the trigger file.

[0097] Step S402: If there is abnormal segmentation data in the business segmentation dataset that can trigger the upgrade of the scenario instance segmentation model, then zero-sample pose estimation is performed on the candidate instances in each business segmentation dataset to obtain the pose parameters of the candidate instances in each business segmentation dataset.

[0098] In this embodiment, each segmentation data in the service segmentation dataset can include a color image and a depth image, and the color image and depth image of each service segmentation data can be different from each other. If there is abnormal segmentation data in the service segmentation dataset that can trigger the upgrade of the scene instance segmentation model, the server can obtain a general instance segmentation model and perform instance segmentation processing on the color image in each service segmentation data using the general instance segmentation model to obtain the candidate mask of each candidate instance corresponding to each service segmentation data. The server can perform back projection processing on the depth image in each service segmentation data to obtain the scene point cloud data corresponding to each service segmentation data. For each service segmentation data, the server can split the corresponding scene point cloud data using the candidate mask of each candidate instance in each service segmentation data to obtain the instance point cloud data corresponding to each candidate instance in each service segmentation data. Furthermore, the server can perform pose estimation processing on the instance point cloud data corresponding to each candidate instance in each service segmentation data using the processing type corresponding to each service segmentation data to obtain the pose parameters corresponding to each candidate instance in each service segmentation data.

[0099] The business segmentation dataset includes M business segmentation data points, and the M business segmentation data points include business segmentation data A. i Business data segmentation A i This includes color images and depth images; M is a positive integer, and i is a positive integer less than or equal to M. The server can obtain a general instance segmentation model and segment the business data A using this model. i The color image in the image is segmented to obtain business segmentation data A. i The candidate mask for the corresponding N candidate instances; N is a positive integer. The server can segment data A for business purposes. i The depth image in the image is back-projected to obtain the business segmentation data A. i The corresponding scene point cloud data is split according to N candidate masks to obtain the instance point cloud data corresponding to each of the N candidate masks. That is, the scene point cloud data is split from the scene point cloud data according to each candidate mask to obtain the instance point cloud data corresponding to each candidate mask. In other words, the scene point cloud data is the point cloud data composed of N candidate instances, and the instance point cloud data is the point cloud data corresponding to a single candidate instance.

[0100] Furthermore, the server can segment data A according to business needs. i The corresponding processing type involves performing pose estimation on the instance point cloud data corresponding to each of the N candidate masks to obtain the pose parameters corresponding to each of the N candidate instances. Among these, the service segmentation data A... iThe corresponding processing type refers to the data segmentation A for this business in the aforementioned trigger file. i The associated processing type indicates the processing type used when the scene instance segmentation model performs pose estimation on the color and depth images in the segmented data during the generation of the segmented data. In other words, the same processing type used by the scene instance segmentation model to generate the segmented data can be used in the pose estimation process.

[0101] Specifically, the N candidate masks include candidate instance B. j The corresponding candidate mask; j is a positive integer less than or equal to N; if the business segmentation data A i The corresponding processing type is regular geometry workpiece processing type, then the server segments the data A according to the business logic. i The corresponding processing type, performing pose estimation processing on the instance point cloud data corresponding to N candidate masks to obtain the pose parameters corresponding to N candidate instances, can be implemented as follows: The server can obtain the set of candidate geometric types indicated by the regular geometric workpiece processing type; based on candidate instance B j The point cloud data of the corresponding candidate mask are used to perform random sample consensus fitting (such as Random Sample Consensus, RANSAC algorithm) on each candidate geometry type in the candidate geometry type set to obtain candidate instance B. j The corresponding target geometry; the target geometry corresponding to a candidate instance includes the candidate geometry type determined from the candidate geometry type set, and the geometric parameters corresponding to the candidate instance; based on candidate instance B j Based on the corresponding candidate geometry type and geometric parameters, construct candidate instance B. j The corresponding pose parameters.

[0102] The candidate geometry type set may include, but is not limited to, regular geometry types such as cylinders, circles, rectangles, and spheres.

[0103] More specifically, taking the random sampling consistency fitting process for a candidate geometry type (such as a cylinder) in the candidate geometry type set as an example, the server can perform random sampling consistency fitting on candidate instance B. j The corresponding instance point cloud data is randomly sampled to obtain a minimum sample set. This minimum sample set includes the minimum number of sampling points required to fit a candidate geometry type (e.g., the minimum number of data points needed to fit a cylinder). Furthermore, the server can make geometric assumptions using the minimum sample set, such as calculating a hypothetical cylinder (including its radius, axial direction, etc.) based on the sampling points. The server can then perform a consistency check on this cylinder, i.e., statistically analyze candidate instances B. jThe server determines the number of data points in the corresponding instance point cloud data whose coordinates lie within the cylinder's area, i.e., the number of inliers. The server can continue repeating the above process of random sampling, geometric assumptions, and consistency checks until a preset number of iterations is reached. The cylinder with the most inliers is then selected as the candidate instance B. j The corresponding cylinder.

[0104] When candidate instance B is obtained j When fitting each candidate geometry (i.e., the candidate geometry corresponding to each candidate geometry type in the candidate geometry type set), we can calculate the fitting inlier rate (the ratio between the number of inlier points and the total number of points), fitting residuals, and other data for each candidate geometry. Based on the fitting inlier rate, fitting residuals, and other data, we can determine candidate instance B from each candidate geometry. j The corresponding target geometry.

[0105] Furthermore, the server determines the candidate instance B based on... j Based on the corresponding candidate geometry type and geometric parameters, construct candidate instance B. j The specific implementation process of the corresponding pose parameters can be: the server can determine the pose parameters based on candidate instance B. j The geometric parameters are used to determine the geometric center point of the target geometry, and the coordinates of this geometric center point are used to determine candidate instance B. j At the origin of the object's coordinate system (i.e., the translation parameter), the server can construct rotation parameters (or rotation matrices) based on the candidate geometry type of the target geometry. For example, if the candidate geometry type is planar, the server can use the normal vector in the geometric parameters as the Z-axis, obtain the X-axis by cross product of the normal vector and any non-parallel vector, and then obtain the Y-axis by cross product of the Z-axis and X-axis. The X, Y, and Z axes then constitute the rotation parameters of the target geometry. If the candidate geometry type is cylindrical, the server can use the axis direction vector in the geometric parameters as the Z-axis, and similarly obtain the X and Y axes. The X, Y, and Z axes then constitute the rotation parameters of the target geometry. Furthermore, the server can use the above translation and rotation parameters to form candidate instance B. j The corresponding pose parameters.

[0106] Optionally, if the business splits data A i The corresponding processing type is arbitrary shape workpiece processing type, then the server segments data A according to the business. iThe specific implementation process for the corresponding processing type, which performs pose estimation processing on the instance point cloud data corresponding to N candidate masks to obtain the pose parameters corresponding to N candidate instances, can be as follows: The server can obtain the 3D models corresponding to T candidate objects in the candidate object set indicated by the processing type of the workpiece of arbitrary shape; T is a positive integer. The server can then perform pose estimation processing on the instance point cloud data corresponding to N candidate masks to obtain the pose parameters corresponding to N candidate instances. j The point cloud data of the corresponding candidate mask is used to perform point cloud registration on the 3D models corresponding to the T candidate objects, resulting in T initial pose parameters and their respective confidence scores. The server can then determine the initial pose parameter with the highest confidence score as candidate instance B. j The corresponding candidate pose parameters; for candidate instance B j The corresponding candidate pose parameters are corrected using an iterative nearest-point algorithm to obtain candidate instance B. j The corresponding pose parameters.

[0107] The server can perform point cloud sampling on the 3D models corresponding to T candidate objects to obtain the model point cloud data corresponding to each candidate object; or, the server can pre-sample the 3D models of the candidate objects to obtain the model point cloud data corresponding to the 3D models of the candidate objects, and associate and store the model point cloud data corresponding to each candidate object with the 3D models of the candidate objects. Thus, the server can obtain the model point cloud data corresponding to the 3D models of the candidate objects simultaneously.

[0108] Furthermore, the server can evaluate candidate instance B. j The first key point and the first local feature descriptor are extracted from the instance point cloud data of the corresponding candidate mask; T second key points and T second local feature descriptors can be extracted from T model point cloud data respectively.

[0109] For a model point cloud dataset, the server can establish an initial candidate point pair matching relationship between the first keypoint of the instance point cloud dataset and the second keypoint of the model point cloud dataset using the first local feature descriptor and the corresponding second local feature descriptor. This yields matched point pairs. A random sampling consensus algorithm is then introduced to randomly sample the minimum point set (usually 3 pairs) from the matched point pairs. The candidate pose parameters corresponding to the model point cloud dataset are calculated, and the number of interior points satisfying the candidate pose parameters is counted. This process is iterated to obtain the candidate pose parameters for each iteration. The candidate pose parameter with the largest number of interior points is selected as the candidate instance B. j The initial pose parameters corresponding to the point cloud data of the model.

[0110] When candidate instance B is obtained jWhen given the initial pose parameters corresponding to T model point cloud data, the server can calculate the confidence level for each of the T initial pose parameters. For example, the inlier ratio (the proportion of inliers to the total number of points in the instance point cloud data) corresponding to each initial pose parameter can be used to determine the confidence level for each initial pose parameter. Furthermore, the server can determine the initial pose parameter with the highest confidence level as candidate instance B. j The corresponding candidate pose parameters; can be used for candidate instance B j The corresponding candidate pose parameters are corrected using an iterative nearest point algorithm (such as the Iterative Closest Point, ICP algorithm) to obtain candidate instance B. j The corresponding pose parameters.

[0111] Optionally, the server performs pose estimation processing on the point cloud data corresponding to the N candidate masks to obtain the pose parameters corresponding to the N candidate instances. The specific implementation process can also be as follows: The server can store a template library for pose estimation. This template library stores multimodal template images, virtual pose parameters, and point cloud data of one or more objects from multiple viewpoints, as well as image features generated based on the multimodal template images of each object from each viewpoint. The multimodal template images include, but are not limited to, color images and depth images. The server can obtain candidate instance B. j Based on the corresponding instance features, semantic visual matching is performed on the image features corresponding to each multimodal template image in the template library to obtain candidate instance B. j The corresponding target multimodal template image and visual fusion score. The target multimodal template image is a template from the template library that matches candidate instance B. j Multimodal template images that match in the semantic visual dimension; visual fusion scores are used to characterize candidate instance B. j The degree of matching between the image features and the target multimodal template image. The server can determine the match between candidate instance B and the target multimodal template image based on the point cloud data indicated by the target multimodal template image in the template library and the virtual pose parameters of the target multimodal template image. j Spatial matching is performed to obtain candidate instance B. j The corresponding spatial matching score; based on candidate instance B j The corresponding visual fusion score and spatial matching score determine the spatial verification result of the target multimodal template image; if the spatial verification result indicates that the spatial verification is passed, then candidate instance B is selected. j The corresponding target multimodal template image, for candidate instance B j Perform pose calculation to obtain candidate instance B. j The corresponding pose parameters.

[0112] It should be noted that in the zero-shot pose estimation process based on the template library described above, there is no need to retrain / fine-tune the weights of the segmentation network when changing to a new target object. Only offline template resources (such as multimodal template images from the template library) are required to complete the migration of the pose estimation system without changing any parameters of the deep learning model.

[0113] It should be noted that during the zero-shot pose estimation process on the server, in addition to outputting the pose parameters corresponding to each candidate instance, it can also output the candidate mask and instance category corresponding to each candidate instance. This can be used as the target candidate mask and target instance category corresponding to the target instance after the target instance is selected in the subsequent process, to generate the pseudo-true annotation set corresponding to the business segmentation dataset.

[0114] This step enables zero-shot pose estimation using a pre-trained general instance segmentation model, which automatically outputs instance segmentation masks and pose parameters as pseudo-real values ​​in a pseudo-real annotation set without any manual annotation. This fills the technical gap in automatic annotation generation in the field of automatic iteration of visual models and reduces annotation costs.

[0115] Furthermore, the server can filter candidate instances in each business segment data according to a preset filtering strategy to obtain the target instance in each business segment data. The filtering process can refer to at least one of the following steps: S403 (corresponding to the specific filtering process of the spatial distribution filtering strategy in the preset filtering strategy), S404 (corresponding to the specific filtering process of the multimodal feature distribution filtering strategy in the preset filtering strategy), and S405 (corresponding to the specific filtering process of the mask area distribution filtering strategy in the preset filtering strategy). That is, the server can filter candidate instances in each business segment data according to any one of steps S403, S404, and S405 to obtain the target instance in each business segment data. Alternatively, the server can also filter candidate instances in each business segment data according to any two of steps S403, S404, and S405; or, the server can execute the three filtering processes of steps S403, S404, and S405 sequentially to filter candidate instances in each business segment data, without any limitation.

[0116] Step S403: Obtain the translation parameters corresponding to the N candidate instances, and calculate the business segmentation data A. i Outlier detection thresholds are set for each coordinate dimension, and outlier candidate instances whose coordinate values ​​in the translation parameters exceed the outlier detection thresholds for the corresponding coordinate dimensions are removed.

[0117] In this embodiment of the application, the service segmentation dataset includes M service segmentation data, and the M service segmentation data includes service segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; business segmentation data A i There are N candidate instances, where N is a positive integer. The server can obtain the translation parameters corresponding to each of the N candidate instances from their respective pose parameters. The translation parameter for a candidate instance includes coordinate values ​​across multiple dimensions. For example, an optional translation parameter can be represented as (1,2,3), meaning that the translation parameter indicates a coordinate value of 1 on the X-axis, 2 on the Y-axis, and 3 on the Z-axis.

[0118] Furthermore, the server can determine the business segmentation data A based on the translation parameters corresponding to the N candidate instances. i Outlier detection range value in each coordinate dimension.

[0119] Specifically, each coordinate dimension can include, but is not limited to, the X-axis, Y-axis, and Z-axis dimensions. The server can calculate the mean of the N translation parameters along the X-axis dimension for each of the N candidate instances. ) and standard deviation ( The mean ± 2 standard deviations along the X-axis coordinate dimension are defined as the business segmentation data A. i The outlier detection range value along the X-axis coordinate dimension, i.e., the outlier detection range value along the X-axis coordinate dimension is [ , Similarly, the server can calculate the business segmentation data A. i Outlier detection range and business segmentation data A along the Y-axis coordinate dimension i The outlier detection range value along the Z-axis coordinate dimension.

[0120] The server can identify the candidate instances (excluding abnormal candidate instances) from among N candidate instances as the business segmentation data A. i The corresponding target instance. Here, an anomaly candidate instance refers to a candidate instance where any coordinate value in the translation parameters does not belong to the outlier detection range of the corresponding coordinate dimension. In other words, the server can remove anomaly candidate instances from N candidate instances where any coordinate value in the translation parameters exceeds the outlier detection threshold of the corresponding coordinate dimension, thereby obtaining the business segmentation data A. i The target instance in.

[0121] Step S404: Segment the business data A according to the candidate masks of N candidate instances. iMultidimensional feature vectors are extracted from color and depth images. Based on the feature detection model, candidate instances of feature anomalies are detected and then eliminated.

[0122] In this embodiment of the application, the service segmentation dataset includes M service segmentation data, and the M service segmentation data includes service segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; business segmentation data A i There are N candidate instances, where N is a positive integer. The server can segment the business data A based on the candidate masks corresponding to the N candidate instances. i Feature extraction is performed on the color and depth images to obtain a multidimensional feature vector corresponding to each candidate instance. The multidimensional feature vector corresponding to a candidate instance indicates the multimodal attribute features associated with that candidate instance and the color and depth images. The server can obtain a feature detection model and perform feature analysis on the multidimensional feature vector corresponding to each candidate instance to obtain the feature analysis result for each candidate instance. Candidate instances whose feature analysis results indicate abnormal results are identified as abnormal candidate instances. The candidate instances other than the abnormal candidate instances among the N candidate instances are identified as the business segmentation data A. i The corresponding target instance.

[0123] Specifically, the server can combine the candidate masks corresponding to N candidate instances and the business segmentation data A. i The modal data, such as color images and depth images, are used to extract features such as depth continuous variance, edge alignment error with depth map, and RGB texture consistency for each candidate instance. These features are then combined to form a multidimensional feature vector for each candidate instance.

[0124] Among them, depth continuity variance is used to characterize the spatial smoothness of depth values ​​within the instance mask region of the candidate instance. Edge-depth map alignment error is used to characterize the spatial alignment between detected edges (from the color image) and the corresponding locations in the depth image where depth jumps occur. RGB texture consistency is used to characterize the uniformity of RGB color or texture within the instance mask region of the candidate instance.

[0125] Furthermore, the server can acquire a feature detection model, which may include, but is not limited to, Isolation Forest and One-Class SVM models. The server can use the feature detection model to perform feature analysis on the multi-dimensional feature vectors corresponding to each candidate instance, outputting feature anomaly results as the feature analysis results of the multi-dimensional feature vectors detected as feature anomalies. Candidate instances whose feature analysis results indicate feature anomalies are identified as anomalous candidate instances. From N candidate instances, these anomalous candidate instances are removed to obtain the business segmentation data A. i The corresponding target instance.

[0126] Step S405: Calculate the quartiles of the mask area of ​​N candidate instances, construct a first mask area threshold and a second mask area threshold using the quartiles, and remove candidate instances whose mask area is less than the first mask area threshold or greater than the second mask area threshold.

[0127] In this embodiment of the application, the service segmentation dataset includes M service segmentation data, and the M service segmentation data includes service segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; business segmentation data A i There are N candidate instances, where N is a positive integer. The server can obtain the mask area of ​​the candidate mask corresponding to each of the N candidate instances, and perform statistical processing on the N mask areas to obtain the quartiles of the mask areas of the N candidate instances. The quartiles include the median, lower quartile (Q1), and upper quartile (Q3). Specifically, the server can sort the N mask areas to obtain sorted N mask areas. From the sorted N mask areas, the server determines the median of the N mask areas. The sorted N mask areas are then divided into a lower half and an upper half based on the median. The median corresponding to the lower half is determined as the lower quartile (Q1) of the N mask areas; the median corresponding to the upper half is determined as the upper quartile (Q3) of the N mask areas.

[0128] Furthermore, the server can use the lower and upper quartiles of N mask areas to generate a first mask area threshold indicating the lower limit of the mask area and a second mask area threshold indicating the upper limit of the mask area. Specifically, the server can calculate the interquartile range (IQR) using the lower and upper quartiles, where IQR = Q3 - Q1. The server can use the lower quartiles and the interquartile range to generate the first mask area threshold indicating the lower limit of the mask area. An optional first mask area threshold can be expressed as "Q1 - 1.5 × IQR". The server can use the upper quartiles and the interquartile range to generate the second mask area threshold indicating the upper limit of the mask area. An optional second mask area threshold can be expressed as "Q3 + 1.5 × IQR".

[0129] Furthermore, the server can remove candidate instances from N candidate instances whose mask area is less than the first mask area threshold or whose mask area is greater than the second mask area threshold to obtain business segmentation data A. i The corresponding target instance.

[0130] Step S406: Based on the 3D model corresponding to each target instance, perform coordinate system transformation on the pose parameters of each target instance to obtain the target pose parameters of each target instance in the world coordinate system.

[0131] In this embodiment, the server can perform point cloud sampling on the 3D model corresponding to each target instance to obtain the model point cloud data corresponding to each target instance. It should be noted that since the aforementioned steps have already obtained the model point cloud data of the candidate objects corresponding to the target instance, point cloud sampling is not required here; the model point cloud data corresponding to the target instances obtained in the aforementioned steps can be used directly. Simultaneously, for target instances in business segmentation data with a processing type of regular geometry workpiece processing, the server can cache the 3D models corresponding to the candidate geometries for each candidate geometry type in the candidate geometry type set, or the model point cloud data of the 3D models corresponding to the candidate geometries. If the server only caches the 3D models corresponding to the candidate geometries, then for target instances in business segmentation data with a processing type of regular geometry workpiece processing, the server can perform point cloud sampling on the 3D model corresponding to each target instance to obtain the model point cloud data corresponding to each target instance.

[0132] Furthermore, the server can perform transformation processing on the model point cloud data corresponding to each target instance based on the pose parameters of each target instance, to obtain the transformed scene point cloud data corresponding to each target instance. An optional transformation processing implementation method can be found in formula (1): (1); As shown in formula (1), Used to indicate point cloud data in a changing scene Used to indicate the pose parameters of a target instance. Used to indicate the model point cloud data corresponding to the target instance.

[0133] The server can determine the target height value of each target instance on the target coordinate axis based on the transformed scene point cloud data corresponding to each target instance. For example, the server can take the minimum value of each data point in the transformed scene point cloud data of a target instance on the target coordinate axis as the target height value of the target instance on the target coordinate axis. Among them, the target height value of a target instance is used to indicate the bottom position of the object corresponding to the target instance; since in industrial scenes, the engineering general coordinate system stipulates that the Z-axis represents the vertical height, the target coordinate axis can be the Z-axis. An optional implementation method for determining the target height value can be found in formula (2): (2); As shown in formula (2), Used to indicate the target height value of the target instance on the Z-axis, i.e., from the transformed scene point cloud data. The target height is determined by taking all elements from the third column of the entire point cloud matrix, which represents the world coordinate system Z-values ​​of each data point in the transformed scene point cloud data. Then, the minimum Z-value is obtained using the "min()" function and taken as the target height. The colon ":" in the text indicates that all rows should be selected, meaning from... Extract all data from the 3rd column of all rows (in programming, the index starts counting from 0, and the value 2 corresponds to the 3rd column).

[0134] The server can determine the camera installation height value from the target height value of each target instance on the target coordinate axis. For example, the server can determine the camera installation height value as the median of the target height values ​​corresponding to each target instance. The server can obtain the preset orientation adjustment parameters and generate a coordinate system transformation matrix based on the camera installation height value and the preset orientation adjustment parameters. Among them, the camera installation height value is used to indicate the vertical distance of the virtual camera from the ground in the world coordinate system; the preset orientation adjustment parameters are used to correct the orientation of the target instance in the world coordinate system to the preset direction, that is, to rotate and align the orientation of all objects, so that the front of the objects is uniform, and to achieve the regular alignment of the object posture in the scene. The server can perform coordinate system transformation on the pose parameters of each target instance according to the coordinate system transformation matrix to obtain the target pose parameters of each target instance in the world coordinate system. An optional method for determining the target pose parameters can be found in formula (3): (3); As shown in formula (3), Used to indicate the target pose parameters located in the world coordinate system; Used to indicate the coordinate system transformation matrix, that is, the transformation matrix that transforms the pose parameters from the camera coordinate system to the world coordinate system; Let be the pose parameters of a target instance, i.e., the pose matrix of the target instance in the camera coordinate system. The coordinate transformation matrix can be represented as " ”, t=[0,0,-z cam ]. To preset the orientation adjustment parameters, z cam Set the camera's height value.

[0135] Optionally, in addition to the above-mentioned process of calculating the target height value of the target instance on the Z-axis and generating a coordinate transformation matrix for coordinate system transformation, the coordinate system transformation process can also be carried out by directly determining the ground position by detecting the ground plane (such as RANSAC plane fitting), and then transforming the pose parameters in the camera coordinate system to the world coordinate system with the ground as Z=0 to obtain the target pose parameters.

[0136] It should be noted that when the target pose parameters of each target instance in each business segmentation data are obtained in the world coordinate system, the server can record the camera installation height value (used to set the position of the virtual camera during subsequent virtual rendering) and store the target pose parameters of each target instance in each business segmentation data in the world coordinate system as a scene description file (such as a description file in JSON format).

[0137] Step S407: Perform virtual rendering on the target pose parameters and 3D model corresponding to the target instance in each business segmentation data to obtain H simulation data; the simulation instance in a simulation data is generated based on the target instance corresponding to the business segmentation data corresponding to that simulation data; H is a positive integer.

[0138] In this embodiment, the server can call a virtual simulation engine (such as Dexverse, Isaac Sim, or Blender), load the scene description file mentioned above through the virtual simulation engine, and perform virtual rendering on the target pose parameters and 3D model corresponding to the target instance in each business segmentation data to obtain H simulation data.

[0139] Specifically, the business segmentation dataset includes M business segmentation data points, and the M business segmentation data points include business segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; business segmentation data A iThis includes R target instances; R is a positive integer. The server can obtain randomized rendering parameters, which include, but are not limited to, light intensity / color, texture mapping, camera noise, depth sensor error, etc. In other words, during this type of virtual rendering (hereinafter referred to as the first virtual rendering), the virtual simulation engine can assign different light intensities / colors, texture mappings, camera noise, depth sensor errors, etc., to achieve simulation data with different rendering effects.

[0140] Furthermore, the server can use a virtual simulation engine to segment data A according to business needs. i The target pose parameters and 3D models of R target instances are used to place R target instances in the initial virtual scene, resulting in a first virtual instance scene containing R simulation instances. Randomized rendering parameters are then used to render the first virtual instance scene, yielding business segmentation data A. i The corresponding first simulation dataset. One of the R simulation instances is determined based on one of the R target instances. In other words, the server simulates business segmentation data A in this type of virtual rendering. i The target pose parameters of each target instance are used to accurately place the CAD model of each target instance, generating business segmentation data A that can reproduce the business scenario. i Multiple first simulation data of the target instance are used to form business segmentation data A. i The corresponding first simulation dataset.

[0141] The server can drop instances based on a preset instance drop probability p. drop (0) <p drop <1) and R, determine the number of instances to be dropped, and then randomly drop an equal number of target instances from the R target instances to obtain D simulation instances, where D is a positive integer less than R, and D = R - the number of instances dropped. Then, the server can place D simulation instances in the initial virtual scene based on the target pose parameters and 3D model of the D simulation instances to obtain a second virtual instance scene. Randomized rendering parameters are used to render the second virtual instance scene to obtain multiple second simulation data sets. These multiple second simulation data sets are then combined to form business segmentation data A. i The corresponding second simulation dataset. In other words, during this type of virtual rendering (hereinafter referred to as the second virtual rendering), the server randomly discards some target instances to generate second simulation data to simulate occlusion or missed detection in the actual scene.

[0142] The server can randomly add 1 to R augmented instances based on R target instances, such as Q augmented instances, where Q is a positive integer less than or equal to R. The Q augmented instances are generated by randomly adjusting the target pose parameters of Q target instances out of the R target instances within a preset adjustment range. For example, shifting the target pose parameters of a target instance by ±5cm or rotating it by ±15° yields the corresponding augmented instance. The server can combine the R target instances and Q augmented instances into F simulation instances, where F is the sum of R and Q. Furthermore, based on the target pose parameters and 3D model of the F simulation instances, the server can place the F simulation instances in the initial virtual scene to obtain a third virtual instance scene. Randomized rendering parameters are then used to render the third virtual instance scene, resulting in multiple third simulation data sets. These multiple third simulation data sets are then combined to form business segmentation data A. i The corresponding third simulation dataset. In other words, during this type of virtual rendering (hereinafter referred to as the third virtual rendering), the server randomly adjusts the target pose parameters of some target instances to obtain augmented instances with slightly different target pose parameters. Based on the original R target instances and D augmented instances, the third simulation data is generated to simulate a scenario of densely stacked similar workpieces.

[0143] This step, based on the reproduction of poses in real-world scenarios (the first virtual rendering mentioned above), introduces two strategies: randomly losing some instances (the second virtual rendering mentioned above) and randomly expanding instances (the third virtual rendering mentioned above), generating three types of simulation data (complete reproduction, random loss, and random expansion). The model is trained using these three types of simulation data, which can significantly improve the generalization ability of the trained model in similar scenarios, effectively avoid overfitting, and enable the trained model to adapt to dynamic situations such as changes in the number of workpieces and changes in the degree of occlusion in actual production lines.

[0144] When the server obtains the first simulation dataset, the second simulation dataset, and the third simulation dataset corresponding to each business segmentation data, it can combine the first simulation dataset, the second simulation dataset, and the third simulation dataset corresponding to each business segmentation data into H simulation datasets.

[0145] It should be noted that, assuming the number of business segmentation data in the business segmentation dataset is 500, the server can generate 1000 simulation data for each type of virtual rendering during the above three types of virtual rendering. For example, if one business segmentation data generates two simulation data in one type of virtual rendering, then H can be equal to 3000.

[0146] Optionally, during the virtual rendering process that generates H simulation data points, the random loss / amplification ratio of the target instance can be set to adaptive, such as dynamically adjusting it based on the evaluation results (increasing the amplification ratio if the generalization is poor). Alternatively, a domain randomization parameter can be introduced to automatically search (e.g., using Bayesian optimization) to find the optimal range of randomization loss / amplification ratios.

[0147] This step accurately reproduces the scene layout in the virtual simulation engine based on the automatically labeled real scene poses and automatically completes the coordinate system transformation (from camera coordinate system to world coordinate system) without manual intervention for alignment. This ensures that the simulation data and the poses of the real scene are completely consistent, eliminating the need for secondary verification and solving the domain alignment problem between simulation data and real data.

[0148] To address the shortcomings of existing technologies in the third category, such as the lack of visual annotation capabilities and the inability to process instance segmentation data, this invention integrates a pre-trained large segmentation model (SAM) and a general zero-shot pose estimation module, which can automatically output instance segmentation masks and 6D poses as pseudo-ground values. Beneficial effects: It fills the technical gap in the field of automatic annotation generation in the automatic iteration of visual models, enabling task orchestration systems to truly possess visual data processing capabilities.

[0149] Step S408: Train the scene instance segmentation model based on H simulation data to obtain the first instance segmentation model; train the initial instance segmentation model based on H simulation data to obtain the second instance segmentation model; the initial instance segmentation model refers to the instance segmentation model loaded with general pre-trained weights.

[0150] In this embodiment, the scene instance segmentation model includes a backbone network and a detection head network; the H simulation data include sample data and sample annotations; the sample corresponding to one simulation data may include the service segmentation data A obtained during virtual rendering. i The corresponding simulated color image and simulated depth image, the annotation corresponding to a simulated data can include the business segmentation data A obtained during virtual rendering. i The corresponding simulation instance mask and simulation pose parameters (determined by the target pose parameters of the target instance); the above sample data may include samples corresponding to H simulation data respectively, and the above sample annotation may include annotations corresponding to H simulation data respectively.

[0151] The specific implementation process of the server training the scene instance segmentation model based on H simulation data to obtain the first instance segmentation model can be as follows: freeze the network parameters of the backbone network in the scene instance segmentation model, input the H simulation data into the scene instance segmentation model, extract features from the sample data in the H simulation data through the backbone network with frozen network parameters, and obtain the feature vectors corresponding to the H simulation data respectively; input the H feature vectors into the detection head network in the scene instance segmentation model to obtain the first detection result, and adjust the network parameters of the detection head network according to the first detection result and sample labeling until the network parameters of the detection head network converge to obtain the first instance segmentation model.

[0152] The specific implementation process of the server training the initial instance segmentation model based on H simulation data to obtain the second instance segmentation model can be as follows: inputting H simulation data into the initial instance segmentation model, extracting features from the sample data in the H simulation data through the backbone network in the initial instance segmentation model to obtain feature vectors corresponding to the H simulation data respectively; inputting the H feature vectors into the detection head network in the initial instance segmentation model to obtain the second detection result; and adjusting the network parameters of the backbone network and the detection head network in the initial instance segmentation model according to the second detection result and sample annotation, until the network parameters in the initial instance segmentation model converge to obtain the second instance segmentation model.

[0153] It should be noted that during the model training process described above, the server can load the configuration file in the trigger file, automatically read the original training set path, model architecture, hyperparameters, and other data, supplement the training set directory indicated by the generated H simulation data sets, and pre-configure two training modes: fine-tuning mode and retraining mode. Fine-tuning mode refers to the process of using newly generated simulation data to perform limited iterative training on an existing pre-trained model (the scene instance segmentation model deployed on-site). In this mode, the original model weights of the scene instance segmentation model are loaded, the backbone network is frozen, only the detection head network is updated, and a small learning rate and fewer training epochs are set, as in the mode described above where the first instance segmentation model is trained based on the scene instance segmentation model. Retraining mode refers to training the model from scratch or completely retraining the model using newly generated simulation data. In this mode, the model does not load the original weights (i.e., it is not trained on top of the scene instance segmentation model), but only loads general pre-trained weights. All network parameters are trainable, and there are more training epochs. The learning rate is set to a standard level. This mode is suitable for situations where the model architecture has changed or where the distribution of simulation data differs significantly from the original training set. This is similar to the mode described above, where a second instance segmentation model is trained based on the initial instance segmentation model. The retrain mode requires more computational resources and training time than the fine-tuning mode, but may achieve better scene adaptability.

[0154] Step S409: Based on the business segmentation dataset, perform model validation on the scenario instance segmentation model, the first instance segmentation model, and the second instance segmentation model to obtain the optimal instance segmentation model; the optimal instance segmentation model is determined based on the first instance segmentation model or the second instance segmentation model; the optimal instance segmentation model is used to replace the scenario instance segmentation model deployed in the business scenario.

[0155] In this embodiment, each segmentation data in the segmentation dataset includes a color image, a depth image, a first instance mask, a first instance category, and a first pose parameter corresponding to the segmentation data output by the scene instance segmentation model. The server performs model validation on the scene instance segmentation model, the first instance segmentation model, and the second instance segmentation model based on the segmentation dataset to obtain the optimal instance segmentation model. Specifically, the server can generate a pseudo-annotation set corresponding to the segmentation dataset based on the target candidate mask, target instance category, and target pose parameter corresponding to the target instance in each segmentation data. The target candidate mask and target instance category corresponding to each target instance are determined during zero-shot pose estimation of the candidate instances in each segmentation data.

[0156] Furthermore, the server can use the first instance mask, first instance category, and first pose parameter corresponding to each service segmentation data output by the scene instance segmentation model as the first inference result corresponding to the scene instance segmentation model. The server can input the color image and depth image included in each service segmentation data in the service segmentation dataset into the first instance segmentation model, and perform instance segmentation processing and pose estimation processing on the color image and depth image included in each service segmentation data through the first instance segmentation model to obtain the second instance mask, second instance category, and second pose parameter corresponding to each service segmentation data output by the first instance segmentation model. The second instance mask, second instance category, and second pose parameter corresponding to each service segmentation data are used as the second inference result of the first instance segmentation model for the service segmentation dataset.

[0157] The server can input the color image and depth image included in each business segmentation data in the business segmentation dataset into the second instance segmentation model. The second instance segmentation model performs instance segmentation processing and pose estimation processing on the color image and depth image included in each business segmentation data, and obtains the third instance mask, third instance category and third pose parameters corresponding to each business segmentation data output by the second instance segmentation model. The third instance mask, third instance category and third pose parameters corresponding to each business segmentation data are used as the third inference result of the second instance segmentation model for the business segmentation dataset.

[0158] The server can evaluate the first, second, and third inference results based on the pseudo-real and true annotation sets, respectively, to obtain the first evaluation metric, the second evaluation metric, and the third evaluation metric corresponding to the scene instance segmentation model. Here, the first, second, and third evaluation metrics are all evaluation metrics; they are simply referred to as the first evaluation metric, the second evaluation metric, and the third evaluation metric to distinguish them from the different instance segmentation models they correspond to.

[0159] The specific evaluation metrics include, but are not limited to, mAP@0.75, instance recall, average 6D pose angle, and translation difference. mAP@0.75 indicates the average precision of instance segmentation. The average precision of instance segmentation corresponding to an instance segmentation model can be determined based on the instance mask output by the instance segmentation model and the target candidate mask corresponding to the target instance in each service segmentation data in the pseudo-annotation set. Instance recall indicates the average of the individual instance recall rates corresponding to M service segmentation data in the service segmentation dataset. The individual instance recall rate corresponding to a service segmentation data can be the ratio between the number of instances output by an instance segmentation model for a service segmentation data and the number of target instances of that service segmentation data in the pseudo-annotation set. Average 6D pose angle and translation difference indicate the translational and angular deviations between the pose parameters predicted by an instance segmentation model and the target pose parameters in the pseudo-annotation set.

[0160] Furthermore, the server can determine the optimal instance segmentation model based on the first evaluation metric, the second evaluation metric, and the third evaluation metric. Specifically, if the second evaluation metric is better than both the first and third evaluation metric, the server can determine the first instance segmentation model as the optimal instance segmentation model.

[0161] Optionally, if the third evaluation metric is better than both the first and second evaluation metric, the server can determine the second instance segmentation model as the optimal instance segmentation model.

[0162] Optionally, if the first evaluation metric is better than both the second and third evaluation metric, then the current simulation data quality is considered insufficient to improve the performance of the scene instance segmentation model or the initial instance segmentation model. The server can then increase the number of simulation data generated, resulting in updated simulation data. For example, for the second virtual rendering process, assuming 1000 second simulation data points were originally generated for the business segmentation dataset, the number is increased from 1000 to 3000. Similarly, for the third virtual rendering process, the number of simulation data points is also increased from 1000 to 3000. That is, the number of updated simulation data points is now 4000. The incremental simulation dataset generated based on this updated dataset can include the newly generated 2000 second simulation data points and 2000 third simulation data points. The server can use the incremental simulation dataset to train the first instance segmentation model and the second instance segmentation model respectively, until the second evaluation metric corresponding to the trained first instance segmentation model or the third evaluation metric corresponding to the trained second instance segmentation model is better than the first evaluation metric. The trained first instance segmentation model or the trained second instance segmentation model is then determined as the optimal instance segmentation model.

[0163] Alternatively, the server can return to the virtual rendering process and train the scene instance segmentation model and the initial instance segmentation model together with the originally generated H simulation data and the incremental simulation dataset. For example, after adding 2000 second simulation data and 2000 third simulation data, the scene instance segmentation model and the initial instance segmentation model can be retrained based on the originally generated 1000 first simulation data, 3000 second simulation data, and 3000 third simulation data, until the second evaluation index corresponding to the trained first instance segmentation model or the third evaluation index corresponding to the trained second instance segmentation model is better than the first evaluation index. The trained first instance segmentation model or the trained second instance segmentation model is then determined as the optimal instance segmentation model.

[0164] Optionally, in addition to determining the optimal instance segmentation model based on the first evaluation metric, the second evaluation metric, and the third evaluation metric, the process of validating the scenario instance segmentation model, the first instance segmentation model, and the second instance segmentation model can also introduce an A / B testing mechanism. This involves deploying the new model (the first instance segmentation model and the second instance segmentation model) to the business devices corresponding to the business scenario for small-scale verification, and deciding whether to fully deploy the first instance segmentation model or the second instance segmentation model based on online feedback metrics (such as capture success rate).

[0165] Furthermore, the server can store the optimal instance segmentation model in a designated storage location. Business devices can periodically poll the server's designated storage location. When the optimal instance segmentation model is detected in the designated storage location, it will automatically download and replace the scene instance segmentation model deployed in the business scenario, thus completing the model deployment. In other words, subsequent business devices can use the deployed optimal instance segmentation model to segment newly generated scene images in the business scenario to obtain new business segmentation data.

[0166] It should be understood that in the subsequent instance segmentation process, when the business device encounters anomalies such as missed detection or false detection during instance segmentation processing using the optimal instance segmentation model, the business device and the server can take the same execution process (i.e., steps S401 to S409) to continue iteratively optimizing the optimal instance segmentation model.

[0167] It should be noted that since the annotations contain a variety of visually relevant information such as masks, instance categories, and pose parameters, the model iterative optimization process mentioned in this application can be used not only for the iterative optimization of instance segmentation models, but also for the iterative optimization of models that rely on annotation sets and can be retrained, such as object detection models and pose estimation models.

[0168] Through the above process, business segmentation datasets generated by the scene instance segmentation model in business scenarios were collected. Abnormal segmentation data that leads to model performance degradation was accurately captured and used as the core data foundation for model iteration and optimization. Relying on automated processes such as zero-shot pose estimation, instance selection, coordinate system transformation, and virtual rendering, fully automated generation of simulation annotation data was achieved, eliminating the need for manual pixel-by-pixel annotation. This significantly shortened the model iteration cycle, improved production line efficiency, avoided subjective errors from manual annotation, unified data generation standards, and stably ensured the quality of training data, effectively improving model optimization results. Furthermore, it can automatically complete candidate instance pose calculation, accurate selection, coordinate transformation, and large-scale simulation data rendering and generation, achieving a fully automated closed loop of simulation data generation, parameter adaptation, and quality verification. No manual debugging intervention is required, eliminating reliance on expert experience and effectively reducing the technical implementation threshold and manual maintenance costs. Simultaneously, because the simulation data is generated based on real business scenarios, its object poses, camera parameters, and other scene features closely match the distribution of real production lines, effectively reducing the domain differences between simulation data and real scenes. Therefore, when using simulation data for model training, the reliability of the segmentation performance of models trained on simulation data in real industrial scenarios can be improved. Furthermore, this solution validates and filters three types of models based on real business segmentation datasets to obtain the optimal instance segmentation model. This model can adapt to performance degradation caused by various scene changes, continuously ensuring high-precision and high-stability operation of the instance segmentation model in visual tasks such as workpiece sorting and defect detection in industrial scenarios. Moreover, this application constructs a complete closed-loop system encompassing "on-site anomaly triggering – automatic data upload – zero-sample pose estimation pseudo-ground value annotation – scene pose reproduction and virtual rendering – automatic training and comparative evaluation – automatic deployment of the optimal model," achieving efficient, automated, and high-quality closed-loop iterative optimization of the instance segmentation model.

[0169] Please see again Figure 5 , Figure 5 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 3 This describes the complete process of iterative optimization of a scenario instance segmentation model in a business scenario. This data processing method can be executed jointly by the business device and the server. Specifically, this data processing method may include at least the following steps S501-S514: Step S501: The business device runs the scenario instance segmentation model in the business scenario.

[0170] In this embodiment of the application, all business devices can be deployed with a scene instance segmentation model, so that the scene image obtained from the business scene can be segmented by the scene instance segmentation model to obtain data such as instance mask and pose parameters.

[0171] Step S502: Perform anomaly detection on the instance mask and instance pose parameters output by the scene instance segmentation model.

[0172] In this embodiment, the service device can perform anomaly determination on the instance mask and instance pose parameters output by the scene instance segmentation model. Specifically, the service device can determine whether the number of instances detected by the scene instance segmentation model in the service scene is less than a preset number threshold; or the service device can determine whether the area deviation between the instance mask area detected by the scene instance segmentation model and the average area of ​​historical masks is greater than a preset deviation threshold; or the service device can determine whether the instance pose parameters detected by the scene instance segmentation model indicate that the instance is in an abnormal capture state.

[0173] When the business device determines that the number of instances detected by the scene instance segmentation model in the business scene is less than the preset number threshold, or the area deviation between the instance mask area detected by the scene instance segmentation model and the average area of ​​the historical mask is greater than the preset deviation threshold, or the instance pose parameters detected by the scene instance segmentation model indicate that the instance is in an abnormal capture state, the business device can consider that the scene instance segmentation model is in an abnormal segmentation process and can jump to step S503.

[0174] When the business device determines that the number of instances detected by the scene instance segmentation model in the business scene is greater than or equal to a preset number threshold, and the area deviation between the instance mask area detected by the scene instance segmentation model and the average area of ​​the historical mask is less than or equal to a preset deviation threshold, and the instance pose parameters detected by the scene instance segmentation model indicate that the instance is in a normal capture state, the business device can consider that the scene instance segmentation model is performing normally at this time, and can jump to step S501, that is, continue to perform instance segmentation processing on the scene image obtained from the business scene through the scene instance segmentation model.

[0175] Step S503: Combine the abnormal segmentation data with some normal segmentation data to form a business segmentation dataset, and upload it to the server.

[0176] In this embodiment of the application, the service device can combine the abnormal segmentation data indicated by the scene instance segmentation model reaching the abnormal trigger condition (i.e., the above-mentioned judgment that the scene instance segmentation model segmentation processing is abnormal) with some of the normal segmentation data output by the scene instance segmentation model to form a service segmentation dataset, and upload the service segmentation dataset to a specified directory of the server.

[0177] Step S504: Obtain the business segmentation dataset generated by the scenario instance segmentation model during runtime in the business scenario.

[0178] In this embodiment of the application, the server can obtain the business segmentation dataset generated by the scenario instance segmentation model during runtime in the business scenario from a specified directory.

[0179] Step S505: Perform zero-shot pose estimation on each candidate instance in the business segmentation dataset to obtain the pose parameters of each candidate instance in the business segmentation dataset.

[0180] In this embodiment, the server can perform instance segmentation processing on the color image in each service segmentation data set in the service segmentation dataset to obtain a candidate mask for each candidate instance in the service segmentation data set. Furthermore, the server can perform back-projection processing on the depth image in each service segmentation data set to obtain scene point cloud data corresponding to each service segmentation data set. For each candidate instance, the server can extract the instance point cloud data corresponding to that candidate instance from the scene point cloud data corresponding to that candidate instance based on the candidate mask of that candidate instance. Further, the server can perform pose estimation processing on the instance point cloud data corresponding to the candidate mask of each candidate instance to obtain the pose parameters corresponding to each candidate instance. The specific implementation process of step S505 can also be found in [reference needed]. Figure 4 The relevant descriptions in step S402 of the embodiment will not be repeated here.

[0181] Step S506: Filter the candidate instances in each business segmentation data according to the preset filtering strategy to obtain the target instance in each business segmentation data.

[0182] In this embodiment, the preset filtering strategy may include, but is not limited to, spatial distribution filtering strategy, multimodal feature distribution filtering strategy, and mask area distribution filtering strategy. The server can execute the spatial distribution filtering strategy, multimodal feature distribution filtering strategy, and mask area distribution filtering strategy respectively to filter candidate instances in each service segmentation data, thereby obtaining the target instance in each service segmentation data. The specific implementation process of step S506 can also be found in [reference needed]. Figure 4 The relevant descriptions of steps S403 to S405 in the embodiments will not be repeated here.

[0183] Step S507: Perform coordinate system transformation on the pose parameters of each target instance.

[0184] In this embodiment, the server can perform coordinate system transformation on the pose parameters of each target instance based on the 3D model corresponding to each target instance, thereby obtaining the target pose parameters of each target instance located in the world coordinate system. The specific implementation process of step S507 can also be found in [reference needed]. Figure 4 The relevant descriptions in step S406 of the embodiment will not be repeated here.

[0185] Step S508: Perform virtual rendering on the target instance in each business segmentation data to obtain H simulation data.

[0186] In this embodiment, the server can perform virtual rendering on the target pose parameters and 3D model corresponding to the target instance in each service segmentation data to obtain H simulation data. Specifically, for each service segmentation data, the server can execute the above-mentioned first virtual rendering, second virtual rendering, and third virtual rendering processes through the virtual simulation engine to generate H simulation data containing the three types of simulation data obtained from these three types of virtual rendering. The specific implementation process of step S508 can also be found in [reference needed]. Figure 4 The relevant descriptions in step S407 of the embodiment will not be repeated here.

[0187] Step S509: Perform model fine-tuning training to obtain the first instance segmentation model.

[0188] In this embodiment, the server can fine-tune the training of the scene instance segmentation model to obtain a first instance segmentation model. Specifically, the server can freeze the network parameters of the backbone network in the scene instance segmentation model, input H simulation data into the scene instance segmentation model, extract features from the sample data in the H simulation data through the backbone network with frozen network parameters, and obtain feature vectors corresponding to the H simulation data respectively; input the H feature vectors into the detection head network in the scene instance segmentation model to obtain a first detection result, and adjust the network parameters of the detection head network according to the first detection result and sample annotation until the network parameters of the detection head network converge to obtain the first instance segmentation model.

[0189] Step S510: Train the model from scratch to obtain the second instance segmentation model.

[0190] In this embodiment, the server can train the initial instance segmentation model from scratch to obtain a second instance segmentation model. Specifically, the server can input H simulation data into the initial instance segmentation model, extract features from the sample data in the H simulation data through the backbone network of the initial instance segmentation model, and obtain feature vectors corresponding to the H simulation data respectively; input the H feature vectors into the detection head network of the initial instance segmentation model to obtain a second detection result; and adjust the network parameters of the backbone network and the detection head network of the initial instance segmentation model according to the second detection result and the sample label, until the network parameters of the initial instance segmentation model converge to obtain the second instance segmentation model.

[0191] Step S511: Compare and evaluate the scene instance segmentation model, the first instance segmentation model, and the second instance segmentation model.

[0192] In this embodiment, the server can compare and evaluate the performance of the three instance segmentation models—scene instance segmentation model, first instance segmentation model, and second instance segmentation model—based on the business segmentation dataset and the pseudo-real annotation set, to obtain the first evaluation index corresponding to the scene instance segmentation model, the second evaluation index corresponding to the first instance segmentation model, and the third evaluation index corresponding to the second instance segmentation model.

[0193] Step S512: Determine whether the optimal model is a scene instance segmentation model.

[0194] In this embodiment, the server can determine which instance segmentation model has the best performance based on the first evaluation metric, the second evaluation metric, and the third evaluation metric. Specifically, if the first evaluation metric is better than both the second and third evaluation metric, the server can determine that the optimal model is the scene instance segmentation model and proceed to step S513.

[0195] If the second evaluation metric is better than the first evaluation metric and better than the third evaluation metric, the server can determine that the optimal model is not the scene instance segmentation model, and can determine the first instance segmentation model as the optimal instance segmentation model, and jump to step S514.

[0196] If the third evaluation metric is better than both the first and second evaluation metric, the server can determine that the optimal model is not the scene instance segmentation model, and can determine the second instance segmentation model as the optimal instance segmentation model, and then proceed to step S514.

[0197] Step S513: Increase the amount of simulation data.

[0198] In this embodiment, the server can increase the amount of simulation data generated to obtain an updated amount of simulation data, and then proceed to step S508 to generate an incremental simulation dataset based on the updated amount of simulation data. The incremental simulation dataset is then used to train the first instance segmentation model and the second instance segmentation model respectively, until the second evaluation metric corresponding to the trained first instance segmentation model or the third evaluation metric corresponding to the trained second instance segmentation model is better than the first evaluation metric. The trained first instance segmentation model or the trained second instance segmentation model is then determined as the optimal instance segmentation model. The specific implementation process of step S513 can also be found in [reference needed]. Figure 4 The relevant descriptions in step S409 of the embodiment will not be repeated here.

[0199] Step S514: Save the optimal instance segmentation model.

[0200] In this embodiment, the server can store the optimal instance segmentation model in a designated storage location. Business devices can periodically poll the designated storage location on the server. When the optimal instance segmentation model is detected in the designated storage location, it automatically downloads and replaces the scene instance segmentation model deployed in the business scenario, completing the model deployment. In other words, subsequent business devices can use the deployed optimal instance segmentation model to segment newly generated scene images in the business scenario, obtaining new business segmentation data.

[0201] Further, please see Figure 6 , Figure 6 This is a schematic diagram of a data processing apparatus provided in an embodiment of this application. The data processing apparatus 600 can be a computer program (including program code, etc.) running on a computer device; for example, the data processing apparatus 600 can be an application software. The data processing apparatus 600 can be used to execute corresponding steps in the methods provided in the embodiments of this application. Figure 6 As shown, the data processing device 600 can be used for Figure 3 , Figure 4 or Figure 5 Specifically, the computer device in the corresponding embodiment may include: a data acquisition module 11, a pose estimation module 12, a coordinate system transformation module 13, a virtual rendering module 14, a model training module 15, and a model verification module 16.

[0202] Data acquisition module 11 is used to acquire the business segmentation dataset generated by the scenario instance segmentation model during runtime in the business scenario; The pose estimation module 12 is used to perform zero-shot pose estimation on each candidate instance in the business segmentation dataset if there is abnormal segmentation data in the business segmentation dataset that can trigger the upgrade of the scene instance segmentation model, to obtain the pose parameters of each candidate instance in the business segmentation dataset, and to filter the candidate instances in each business segmentation dataset according to the preset filtering strategy to obtain the target instance in each business segmentation dataset. The coordinate system transformation module 13 is used to transform the pose parameters of each target instance according to the three-dimensional model corresponding to each target instance, so as to obtain the target pose parameters of each target instance in the world coordinate system. The virtual rendering module 14 is used to perform virtual rendering of the target pose parameters and 3D model corresponding to the target instance in each business segmentation data to obtain H simulation data; the simulation instance in a simulation data is generated based on the target instance corresponding to the business segmentation data of that simulation data; H is a positive integer; The model training module 15 is used to train the scene instance segmentation model based on H simulation data to obtain the first instance segmentation model, and to train the initial instance segmentation model based on H simulation data to obtain the second instance segmentation model; the initial instance segmentation model refers to the instance segmentation model loaded with general pre-trained weights. The model validation module 16 is used to validate the scenario instance segmentation model, the first instance segmentation model, and the second instance segmentation model based on the business segmentation dataset to obtain the optimal instance segmentation model. The optimal instance segmentation model is determined based on the first instance segmentation model or the second instance segmentation model. The optimal instance segmentation model is used to replace the scenario instance segmentation model deployed in the business scenario.

[0203] In one optional implementation, the business segmentation dataset includes M business segmentation data points, and the M business segmentation data points include business segmentation data A. i Business data segmentation A i Includes color images and depth images; M is a positive integer, and i is a positive integer less than or equal to M; The pose estimation module 12 is used to perform zero-shot pose estimation on each candidate instance in the service segmentation dataset. When obtaining the pose parameters of each candidate instance in the service segmentation dataset, the pose estimation module 12 is specifically used to perform the following operations: Obtain a generic instance segmentation model, and then segment the business data A using the generic instance segmentation model. i The color image in the image is segmented to obtain business segmentation data A. i The candidate mask for the corresponding N candidate instances; N is a positive integer; Data A is split into business segments. i The depth image in the image is back-projected to obtain the business segmentation data A. i The corresponding scene point cloud data is split according to N candidate masks to obtain the instance point cloud data corresponding to each of the N candidate masks. Based on business segmentation data A i The corresponding processing type performs pose estimation processing on the instance point cloud data corresponding to N candidate masks to obtain the pose parameters corresponding to the N candidate instances.

[0204] In one alternative implementation, business data A is split. i The corresponding processing type is regular geometry workpiece processing type; N candidate masks include candidate instance B. j The corresponding candidate mask; j is a positive integer less than or equal to N; The attitude estimation module 12 is used to segment data A based on business needs. iThe corresponding processing type involves performing pose estimation processing on the instance point cloud data corresponding to N candidate masks to obtain the pose parameters corresponding to the N candidate instances. Specifically, the pose estimation module 12 performs the following operations: Obtain the set of candidate geometry types indicated by the regular geometry workpiece processing type; Based on candidate instance B j The instance point cloud data corresponding to the candidate mask are used to perform random sampling consistency fitting for each candidate geometry type in the candidate geometry type set to obtain candidate instance B. j The target geometry corresponding to a candidate instance; the target geometry corresponding to a candidate instance includes the candidate geometry type determined from the set of candidate geometry types, and the geometric parameters corresponding to the candidate instance; Based on candidate instance B j Based on the corresponding candidate geometry type and geometric parameters, construct candidate instance B. j The corresponding pose parameters.

[0205] In one alternative implementation, business data A is split. i The corresponding processing type is arbitrary shape workpiece processing type; N candidate masks include candidate instance B. j The corresponding candidate mask; j is a positive integer less than or equal to N; The attitude estimation module 12 is used to segment data A based on business needs. i The corresponding processing type involves performing pose estimation processing on the instance point cloud data corresponding to N candidate masks to obtain the pose parameters corresponding to the N candidate instances. Specifically, the pose estimation module 12 performs the following operations: Obtain the 3D models of T candidate objects from the candidate object set indicated by the workpiece processing type of arbitrary shape; T is a positive integer. Based on candidate instance B j The instance point cloud data of the corresponding candidate mask are used to perform point cloud registration on the 3D models corresponding to the T candidate objects to obtain T initial pose parameters and the confidence scores corresponding to the T initial pose parameters respectively. The initial pose parameters with the highest confidence level are selected as candidate instance B. j The corresponding candidate pose parameters; For candidate instance B j The corresponding candidate pose parameters are corrected using an iterative nearest-point algorithm to obtain candidate instance B. j The corresponding pose parameters.

[0206] In one optional implementation, the business segmentation dataset includes M business segmentation data points, and the M business segmentation data points include business segmentation data A. iM is a positive integer, and i is a positive integer less than or equal to M; business segmentation data A i Includes N candidate instances, where N is a positive integer; The attitude estimation module 12 is used to filter candidate instances in each business segmentation data according to a preset filtering strategy. When obtaining the target instance in each business segmentation data, the attitude estimation module 12 is specifically used to perform the following operations: From the pose parameters corresponding to N candidate instances, obtain the translation parameters corresponding to N candidate instances respectively; the translation parameters corresponding to a candidate instance include coordinate values ​​in multiple coordinate dimensions; Based on the translation parameters corresponding to the N candidate instances, determine the business segmentation data A. i Outlier detection range values ​​in each coordinate dimension; Among the N candidate instances, excluding the abnormal candidate instances, the candidate instances are determined as business segmentation data A. i The corresponding target instance; an anomaly candidate instance refers to a candidate instance in which any coordinate value in the translation parameter does not belong to the outlier detection range value in the coordinate dimension corresponding to that coordinate value.

[0207] In one optional implementation, the business segmentation dataset includes M business segmentation data points, and the M business segmentation data points include business segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; business segmentation data A i Includes N candidate instances, where N is a positive integer; The attitude estimation module 12 is used to filter candidate instances in each business segmentation data according to a preset filtering strategy. When obtaining the target instance in each business segmentation data, the attitude estimation module 12 is specifically used to perform the following operations: Based on the candidate masks corresponding to the N candidate instances, the business segmentation data A is processed. i Feature extraction is performed on the color image and depth image to obtain a multi-dimensional feature vector corresponding to each candidate instance; the multi-dimensional feature vector corresponding to a candidate instance is used to indicate the multimodal attribute features that the candidate instance is associated with the color image and depth image. Obtain the feature detection model, and perform feature analysis on the multi-dimensional feature vector corresponding to each candidate instance through the feature detection model to obtain the feature analysis results corresponding to each candidate instance. Candidate instances whose feature analysis results indicate anomalies are identified as anomalous candidate instances. The candidate instances other than the anomalous candidate instances among the N candidate instances are identified as business segmentation data A. i The corresponding target instance.

[0208] In one optional implementation, the business segmentation dataset includes M business segmentation data points, and the M business segmentation data points include business segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; business segmentation data A i Includes N candidate instances, where N is a positive integer; The attitude estimation module 12 is used to filter candidate instances in each business segmentation data according to a preset filtering strategy. When obtaining the target instance in each business segmentation data, the attitude estimation module 12 is specifically used to perform the following operations: Obtain the mask area of ​​the candidate mask corresponding to each of the N candidate instances, perform statistical processing on the N mask areas, and generate the lower quartile and upper quartile for the N mask areas; Using the lower quartile and upper quartile of N mask areas, a first mask area threshold for indicating the lower limit of the mask area and a second mask area threshold for indicating the upper limit of the mask area are generated. Candidate instances with a mask area smaller than the first mask area threshold or a mask area larger than the second mask area threshold are identified as abnormal candidate instances. The candidate instances other than the abnormal candidate instances among the N candidate instances are identified as business segmentation data A. i The corresponding target instance.

[0209] In one optional implementation, the coordinate system transformation module 13 is used to perform coordinate system transformation on the pose parameters of each target instance based on the 3D model corresponding to each target instance. When obtaining the target pose parameters of each target instance in the world coordinate system, the coordinate system transformation module 13 is specifically used to perform the following operations: Point cloud sampling is performed on the 3D model corresponding to each target instance to obtain the model point cloud data corresponding to each target instance. Based on the pose parameters of each target instance, the model point cloud data corresponding to each target instance is transformed to obtain the transformed scene point cloud data corresponding to each target instance. Based on the transformed scene point cloud data corresponding to each target instance, determine the target height value of each target instance on the target coordinate axis; the target height value of a target instance is used to indicate the bottom position of the object corresponding to that target instance; From the target height value of each target instance on the target coordinate axis, determine the camera installation height value, obtain the preset orientation adjustment parameter, and generate a coordinate system transformation matrix based on the camera installation height value and the preset orientation adjustment parameter; the camera installation height value is used to indicate the vertical distance of the virtual camera from the ground in the world coordinate system; the preset orientation adjustment parameter is used to correct the orientation of the target instance in the world coordinate system to the preset direction; The pose parameters of each target instance are transformed using the coordinate system transformation matrix to obtain the target pose parameters of each target instance in the world coordinate system.

[0210] In one optional implementation, the business segmentation dataset includes M business segmentation data points, and the M business segmentation data points include business segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; business segmentation data A i Includes R target instances; R is a positive integer; The virtual rendering module 14 is used to perform virtual rendering of the target pose parameters and 3D model corresponding to the target instance in each business segmentation data. When H simulation data are obtained, the virtual rendering module 14 is specifically used to perform the following operations: Obtain randomized rendering parameters and segment data A according to business logic. i The target pose parameters and 3D models of R target instances are used to place R target instances in the initial virtual scene, resulting in a first virtual instance scene containing R simulation instances. Randomized rendering parameters are then used to render the first virtual instance scene, yielding business segmentation data A. i The corresponding first simulation dataset; one of the R simulation instances is determined based on one of the R target instances; Based on the target pose parameters and 3D models of D simulation instances, D simulation instances are placed in the initial virtual scene to obtain the second virtual instance scene. Randomized rendering parameters are used to render the second virtual instance scene to obtain business segmentation data A. i The corresponding second simulation dataset; the D simulation instances are obtained by randomly deleting some target instances from R target instances according to the preset instance discard probability; D is a positive integer less than R; Based on the target pose parameters and 3D models of F simulation instances, F simulation instances are placed in the initial virtual scene to obtain the third virtual instance scene. Randomized rendering parameters are used to render the third virtual instance scene to obtain the business segmentation data A. i The corresponding third simulation dataset; F simulation instances include R target instances and Q augmented instances; the Q augmented instances are generated by randomly adjusting the target pose parameters of the Q target instances among the R target instances within a preset adjustment range; Q is a positive integer less than or equal to R, and F is the sum of R and Q; When the first simulation dataset, second simulation dataset, and third simulation dataset corresponding to each business segmentation data are obtained, the first simulation dataset, second simulation dataset, and third simulation dataset corresponding to each business segmentation data are combined into H simulation datasets.

[0211] In one alternative implementation, the scene instance segmentation model includes a backbone network and a detection head network; the H simulation data include sample data and sample annotations; Model training module 15 is used to train the scene instance segmentation model based on H simulation data. When the first instance segmentation model is obtained, model training module 15 is specifically used to perform the following operations: The network parameters of the backbone network in the scene instance segmentation model are frozen. H simulation data are input into the scene instance segmentation model. The backbone network with frozen network parameters is used to extract features from the sample data in the H simulation data to obtain the feature vectors corresponding to the H simulation data respectively. H feature vectors are input into the detection head network of the scene instance segmentation model to obtain the first detection result. Based on the first detection result and sample labels, the network parameters of the detection head network are adjusted until the network parameters of the detection head network converge to obtain the first instance segmentation model.

[0212] In one alternative implementation, the H simulation data include sample data and sample annotations; Model training module 15 is used to train the initial instance segmentation model based on H simulation data. When obtaining the second instance segmentation model, model training module 15 is specifically used to perform the following operations: H simulation data are input into the initial instance segmentation model. The backbone network in the initial instance segmentation model is used to extract features from the sample data of the H simulation data to obtain the feature vectors corresponding to the H simulation data respectively. H feature vectors are input into the detection head network in the initial instance segmentation model to obtain the second detection result. Based on the second detection result and sample labels, the network parameters of the backbone network and the network parameters of the detection head network in the initial instance segmentation model are adjusted until the network parameters in the initial instance segmentation model converge to obtain the second instance segmentation model.

[0213] In one optional implementation, each business segmentation data in the business segmentation dataset includes a color image, a depth image, a first instance mask, a first instance category, and a first pose parameter corresponding to the business segmentation data output by the scene instance segmentation model. Model validation module 16 is used to validate the scene instance segmentation model, the first instance segmentation model, and the second instance segmentation model based on the business segmentation dataset. When the optimal instance segmentation model is obtained, model validation module 16 is specifically used to perform the following operations: Based on the target candidate mask, target instance category, and target pose parameters corresponding to the target instance in each business segmentation data, a pseudo-true annotation set corresponding to the business segmentation dataset is generated; the target candidate mask and target instance category corresponding to each target instance are determined during the zero-shot pose estimation process for each candidate instance in the business segmentation dataset. The first instance mask, first instance category, and first pose parameter corresponding to each business segmentation data output by the scene instance segmentation model are used as the first inference result. The second inference result of the first instance segmentation model for the business segmentation dataset and the third inference result of the second instance segmentation model for the business segmentation dataset are obtained. Based on the pseudo-true annotation set, the first inference result, the second inference result and the third inference result are evaluated respectively to obtain the first evaluation index, the second evaluation index, and the third evaluation index corresponding to the scene instance segmentation model. The optimal instance segmentation model is determined based on the first evaluation metric, the second evaluation metric, and the third evaluation metric.

[0214] In one optional implementation, when the model validation module 16 determines the optimal instance segmentation model based on the first evaluation metric, the second evaluation metric, and the third evaluation metric, the model validation module 16 is specifically used to perform the following operations: If the second evaluation metric is better than the first evaluation metric and better than the third evaluation metric, then the first instance segmentation model is determined as the optimal instance segmentation model. If the third evaluation metric is better than both the first and second evaluation metrics, then the second instance segmentation model is determined as the optimal instance segmentation model. If the first evaluation metric is better than the second evaluation metric and the third evaluation metric, then increase the number of simulation data generated to obtain the number of updated simulation data, and generate an incremental simulation dataset based on the number of updated simulation data. The first instance segmentation model and the second instance segmentation model are trained using an incremental simulation dataset until the second evaluation index corresponding to the trained first instance segmentation model or the third evaluation index corresponding to the trained second instance segmentation model is better than the first evaluation index. The trained first instance segmentation model or the trained second instance segmentation model is then determined as the optimal instance segmentation model.

[0215] In one optional implementation, the business segmentation dataset includes normal segmentation data output by the scenario instance segmentation model deployed in the business scenario, and abnormal segmentation data indicated by the scenario instance segmentation model reaching an abnormal triggering condition; the abnormal triggering condition includes at least one of the following: the number of instances detected by the scenario instance segmentation model in the business scenario is less than a preset number threshold, the area deviation between the instance mask area detected by the scenario instance segmentation model and the historical average mask area is greater than a preset deviation threshold, and the instance pose parameters detected by the scenario instance segmentation model indicate that the instance is in an abnormal capture state.

[0216] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 7 As shown, the computer device 700 in this embodiment may include a processor 701, a network interface 704, and a memory 705. Furthermore, the computer device 700 may also include a user interface 703 and at least one communication bus 702. The communication bus 702 is used to enable communication between these components. The user interface 703 may include a display screen and a keyboard; optionally, the user interface 703 may also include a standard wired interface or a wireless interface. The network interface 704 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 705 may be a high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 705 may also be at least one storage device located remotely from the processor 701. Figure 7 As shown, the memory 705, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application.

[0217] Network interface 704 provides network communication elements; user interface 703 is mainly used to provide an input interface for users; and processor 701 can be used to call the device control application stored in memory 705 to perform the following operations: Obtain the business segmentation dataset generated by the scenario instance segmentation model during runtime in the business scenario; If there are abnormal segmentation data in the business segmentation dataset that can trigger the upgrade of the scenario instance segmentation model, then zero-shot pose estimation is performed on the candidate instances in each business segmentation dataset to obtain the pose parameters of the candidate instances in each business segmentation dataset. The candidate instances in each business segmentation dataset are then filtered according to the preset filtering strategy to obtain the target instance in each business segmentation dataset. Based on the 3D model corresponding to each target instance, the pose parameters of each target instance are transformed to obtain the target pose parameters of each target instance in the world coordinate system. For each target instance in the business segmentation data, the target pose parameters and 3D model are virtually rendered to obtain H simulation data. The simulation instance in a simulation data is generated based on the target instance corresponding to the business segmentation data of that simulation data. H is a positive integer. The scene instance segmentation model is trained using H simulation data to obtain the first instance segmentation model. The initial instance segmentation model is then trained using H simulation data to obtain the second instance segmentation model. The initial instance segmentation model refers to the instance segmentation model loaded with general pre-trained weights. The scenario instance segmentation model, the first instance segmentation model, and the second instance segmentation model are validated based on the business segmentation dataset to obtain the optimal instance segmentation model. The optimal instance segmentation model is determined based on either the first instance segmentation model or the second instance segmentation model. The optimal instance segmentation model is used to replace the scenario instance segmentation model deployed in the business scenario.

[0218] Furthermore, it should be noted that embodiments of this application also provide a computer-readable storage medium storing a computer program adapted to be loaded and executed by the processor. Figure 3 , Figure 4 or Figure 5 For details on the methods provided in each step, please refer to the document. Figure 3 , Figure 4 or Figure 5 The implementation methods provided for each step are not repeated here. Furthermore, the beneficial effects of using the same method are also not repeated. For technical details not disclosed in the computer-readable storage medium embodiments involved in this application, please refer to the description of the method embodiments of this application. As an example, a computer program may be deployed to execute on a single computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed across multiple locations and interconnected via a communication network.

[0219] The computer-readable storage medium can be the apparatus provided in any of the foregoing embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the computer device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0220] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform... Figure 3 , Figure 4 or Figure 5 The methods provided are among the various optional methods available in the code, so they will not be elaborated upon here.

[0221] The terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.

[0222] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0223] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0224] The methods and related apparatus provided in this application are described with reference to the method flowcharts and / or structural diagrams provided in this application. Specifically, each block of the method flowcharts and / or structural diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by a computer program. These computer programs can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable device to create a machine, such that the computer program, executed by the processor of the computer or other programmable device, produces a mechanism for implementing the process... Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1 The computer program may be a means for performing the functions specified in one or more boxes. These computer programs may also be stored in a computer-readable storage medium that can direct a computer or other programmable device to function in a particular manner, causing the computer program stored in the computer-readable storage medium to produce an article of manufacture including the program means, or to be transmitted via a computer-readable storage medium. The computer program can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The program means is implemented in the process. Figure 1 A schematic diagram of one or more processes and / or structures. Figure 1 The functions specified in one or more boxes. These computer programs may also be loaded onto a computer or other programmable device to cause a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing the computer program executing on the computer or other programmable device with the means to implement the process. Figure 1 A process or multiple processes and / or structures illustrate the steps of the functions specified in one or more boxes.

[0225] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.

[0226] The modules in the device of this application embodiment can be merged, divided, and deleted according to actual needs.

[0227] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. A data processing method, characterized in that, The method includes: Obtain the business segmentation dataset generated by the scenario instance segmentation model during runtime in the business scenario; If there is abnormal segmentation data in the business segmentation dataset that can trigger the upgrade of the scene instance segmentation model, then zero-shot pose estimation is performed on the candidate instances in each business segmentation dataset to obtain the pose parameters of the candidate instances in each business segmentation dataset. The candidate instances in each business segmentation dataset are then filtered according to a preset filtering strategy to obtain the target instance in each business segmentation dataset. Based on the 3D model corresponding to each target instance, coordinate system transformation is performed on the pose parameters of each target instance to obtain the target pose parameters of each target instance in the world coordinate system. Virtual rendering is performed on the target pose parameters and 3D model corresponding to the target instance in each business segmentation data to obtain H simulation data; the simulation instance in a simulation data is generated based on the target instance corresponding to the business segmentation data of that simulation data; H is a positive integer; The scene instance segmentation model is trained based on the H simulation data to obtain a first instance segmentation model. The initial instance segmentation model is then trained based on the H simulation data to obtain a second instance segmentation model. The initial instance segmentation model refers to an instance segmentation model loaded with general pre-trained weights. The scenario instance segmentation model, the first instance segmentation model, and the second instance segmentation model are validated based on the business segmentation dataset to obtain the optimal instance segmentation model. The optimal instance segmentation model is determined based on the first instance segmentation model or the second instance segmentation model. The optimal instance segmentation model is used to replace the scenario instance segmentation model deployed in the business scenario.

2. The method according to claim 1, characterized in that, The service segmentation dataset includes M service segmentation data points, and the M service segmentation data points include service segmentation data A. i The business segmentation data A i Includes color images and depth images; M is a positive integer, and i is a positive integer less than or equal to M; The step of performing zero-shot pose estimation on candidate instances in each service segmentation data set in the service segmentation dataset to obtain the pose parameters of candidate instances in each service segmentation data set includes: Obtain a general instance segmentation model, and segment the business data A using the general instance segmentation model. i The color image in the image is segmented to obtain business segmentation data A. i The candidate mask for the corresponding N candidate instances; N is a positive integer; For the business segmentation data A i The depth image in the image is back-projected to obtain the business segmentation data A. i The corresponding scene point cloud data is split according to N candidate masks to obtain the instance point cloud data corresponding to each of the N candidate masks. Based on the business segmentation data A i The corresponding processing type performs pose estimation processing on the instance point cloud data corresponding to N candidate masks to obtain the pose parameters corresponding to the N candidate instances.

3. The method according to claim 2, characterized in that, The business segmentation data A i The corresponding processing type is the regular geometry workpiece processing type; the N candidate masks include candidate instance B. j The corresponding candidate mask; j is a positive integer less than or equal to N; The data based on the business segmentation data A i The corresponding processing type performs pose estimation on the instance point cloud data corresponding to N candidate masks to obtain the pose parameters corresponding to the N candidate instances, including: Obtain the set of candidate geometry types indicated by the regular geometry workpiece processing type; According to the candidate instance B j The instance point cloud data corresponding to the candidate mask are used to perform random sampling consistency fitting processing on each candidate geometry type in the candidate geometry type set to obtain the candidate instance B. j The corresponding target geometry; the target geometry corresponding to a candidate instance includes the candidate geometry type determined from the set of candidate geometry types, and the geometric parameters corresponding to the candidate instance; According to the candidate instance B j Based on the corresponding candidate geometry type and geometric parameters, the candidate instance B is constructed. j The corresponding pose parameters.

4. The method according to claim 2, characterized in that, The business segmentation data A i The corresponding processing type is the arbitrary shape workpiece processing type; the N candidate masks include candidate instance B. j The corresponding candidate mask; j is a positive integer less than or equal to N; The data based on the business segmentation data A i The corresponding processing type performs pose estimation on the instance point cloud data corresponding to N candidate masks to obtain the pose parameters corresponding to the N candidate instances, including: Obtain the 3D models corresponding to T candidate objects in the candidate object set indicated by the arbitrary shape workpiece processing type; T is a positive integer; According to the candidate instance B j The instance point cloud data of the corresponding candidate mask are used to perform point cloud registration on the 3D models corresponding to the T candidate objects to obtain T initial pose parameters and the confidence scores corresponding to the T initial pose parameters respectively. The initial pose parameters with the highest confidence level are determined as the candidate instance B. j The corresponding candidate pose parameters; For the candidate instance B j The corresponding candidate pose parameters are corrected using an iterative nearest-point algorithm to obtain the candidate instance B. j The corresponding pose parameters.

5. The method according to claim 1, characterized in that, The service segmentation dataset includes M service segmentation data points, and the M service segmentation data points include service segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; the business segmentation data A i Includes N candidate instances, where N is a positive integer; The step of filtering candidate instances in each business segmentation data according to a preset filtering strategy to obtain target instances in each business segmentation data includes: From the pose parameters corresponding to the N candidate instances, obtain the translation parameters corresponding to the N candidate instances respectively; The translation parameters corresponding to a candidate instance include coordinate values ​​in multiple coordinate dimensions; Based on the translation parameters corresponding to the N candidate instances, the business segmentation data A is determined. i Outlier detection range values ​​in each coordinate dimension; The candidate instances, excluding the abnormal candidate instances, are selected from the N candidate instances as the business segmentation data A. i The corresponding target instance; the abnormal candidate instance refers to a candidate instance in which any coordinate value in the translation parameter does not belong to the outlier detection range value in the coordinate dimension corresponding to that coordinate value.

6. The method according to claim 1, characterized in that, The service segmentation dataset includes M service segmentation data points, and the M service segmentation data points include service segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; the business segmentation data A i Includes N candidate instances, where N is a positive integer; The step of filtering candidate instances in each business segmentation data according to a preset filtering strategy to obtain target instances in each business segmentation data includes: Based on the candidate masks corresponding to the N candidate instances, the business segmentation data A is processed. i Feature extraction is performed on the color image and depth image to obtain a multi-dimensional feature vector corresponding to each candidate instance; the multi-dimensional feature vector corresponding to a candidate instance is used to indicate the multimodal attribute features that the candidate instance is associated with the color image and the depth image. A feature detection model is obtained, and feature analysis is performed on the multi-dimensional feature vector corresponding to each candidate instance through the feature detection model to obtain the feature analysis result corresponding to each candidate instance. Candidate instances whose feature analysis results indicate anomalies are identified as anomalous candidate instances. The candidate instances other than the anomalous candidate instances among the N candidate instances are identified as the business segmentation data A. i The corresponding target instance.

7. The method according to claim 1, characterized in that, The service segmentation dataset includes M service segmentation data points, and the M service segmentation data points include service segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; the business segmentation data A i Includes N candidate instances, where N is a positive integer; The step of filtering candidate instances in each business segmentation data according to a preset filtering strategy to obtain target instances in each business segmentation data includes: Obtain the mask area of ​​the candidate mask corresponding to each of the N candidate instances, perform statistical processing on the N mask areas, and generate the lower quartile and upper quartile for the N mask areas; Using the lower quartile and upper quartile of the N mask areas, a first mask area threshold for indicating the lower limit of the mask area and a second mask area threshold for indicating the upper limit of the mask area are generated. Candidate instances with a mask area smaller than the first mask area threshold or a mask area larger than the second mask area threshold are identified as abnormal candidate instances. The candidate instances other than the abnormal candidate instances among the N candidate instances are identified as the service segmentation data A. i The corresponding target instance.

8. The method according to claim 1, characterized in that, The step of performing coordinate system transformation on the pose parameters of each target instance based on the 3D model corresponding to each target instance to obtain the target pose parameters of each target instance in the world coordinate system includes: Point cloud sampling is performed on the 3D model corresponding to each target instance to obtain the model point cloud data corresponding to each target instance. Based on the pose parameters of each target instance, the model point cloud data corresponding to each target instance is transformed to obtain the transformed scene point cloud data corresponding to each target instance. Based on the transformed scene point cloud data corresponding to each target instance, the target height value of each target instance on the target coordinate axis is determined; the target height value of a target instance is used to indicate the bottom position of the object corresponding to that target instance; From the target height values ​​of each target instance on the target coordinate axis, determine the camera installation height value, obtain the preset orientation adjustment parameter, and generate a coordinate system transformation matrix based on the camera installation height value and the preset orientation adjustment parameter; the camera installation height value is used to indicate the vertical distance of the virtual camera from the ground in the world coordinate system; the preset orientation adjustment parameter is used to correct the orientation of the target instance in the world coordinate system to a preset direction; The pose parameters of each target instance are transformed according to the coordinate system transformation matrix to obtain the target pose parameters of each target instance in the world coordinate system.

9. The method according to claim 1, characterized in that, The service segmentation dataset includes M service segmentation data points, and the M service segmentation data points include service segmentation data A. i M is a positive integer, and i is a positive integer less than or equal to M; the business segmentation data A i Includes R target instances; R is a positive integer; The virtual rendering of the target pose parameters and 3D model corresponding to the target instance in each business segmentation data is performed to obtain H simulation data, including: Obtain randomized rendering parameters and segment the data A according to the business logic. i The target pose parameters and 3D models of R target instances are used to place the R target instances in the initial virtual scene, resulting in a first virtual instance scene containing R simulation instances. The first virtual instance scene is then rendered using the randomized rendering parameters to obtain business segmentation data A. i The corresponding first simulation dataset; one of the R simulation instances is determined based on one of the R target instances; Based on the target pose parameters and 3D models of D simulation instances, the D simulation instances are placed in the initial virtual scene to obtain a second virtual instance scene. The second virtual instance scene is then rendered using the randomized rendering parameters to obtain business segmentation data A. i The corresponding second simulation dataset; the D simulation instances are obtained by randomly deleting some target instances from the R target instances according to a preset instance discard probability; D is a positive integer less than R; Based on the target pose parameters and 3D models of F simulation instances, the F simulation instances are placed in the initial virtual scene to obtain a third virtual instance scene. The third virtual instance scene is then rendered using the randomized rendering parameters to obtain business segmentation data A. i The corresponding third simulation dataset; the F simulation instances include the R target instances and Q augmented instances; the Q augmented instances are generated by randomly adjusting the target pose parameters of the Q target instances among the R target instances within a preset adjustment range; Q is a positive integer less than or equal to R, and F is the sum of R and Q; When the first simulation dataset, second simulation dataset, and third simulation dataset corresponding to each business segmentation data are obtained, the first simulation dataset, second simulation dataset, and third simulation dataset corresponding to each business segmentation data are combined into H simulation datasets.

10. The method according to claim 1, characterized in that, The scene instance segmentation model includes a backbone network and a detection head network; the H simulation data include sample data and sample annotations; The step of training the scene instance segmentation model based on the H simulation data to obtain the first instance segmentation model includes: The network parameters of the backbone network in the scene instance segmentation model are frozen, the H simulation data are input into the scene instance segmentation model, and the sample data in the H simulation data are extracted by the backbone network with frozen network parameters to obtain the feature vectors corresponding to the H simulation data respectively. H feature vectors are input into the detection head network in the scene instance segmentation model to obtain a first detection result. Based on the first detection result and the sample label, the network parameters of the detection head network are adjusted until the network parameters of the detection head network converge to obtain the first instance segmentation model.

11. The method according to claim 1, characterized in that, The H simulation data include sample data and sample annotations; The step of training the initial instance segmentation model based on the H simulation data to obtain the second instance segmentation model includes: The H simulation data are input into the initial instance segmentation model. The backbone network in the initial instance segmentation model is used to extract features from the sample data in the H simulation data to obtain the feature vectors corresponding to the H simulation data respectively. H feature vectors are input into the detection head network of the initial instance segmentation model to obtain a second detection result. Based on the second detection result and the sample label, the network parameters of the backbone network and the network parameters of the detection head network in the initial instance segmentation model are adjusted until the network parameters in the initial instance segmentation model converge to obtain the second instance segmentation model.

12. The method according to claim 1, characterized in that, Each segmentation data in the segmentation dataset includes a color image, a depth image, a first instance mask, a first instance category, and a first pose parameter corresponding to the segmentation data output by the scene instance segmentation model. The step of validating the scene instance segmentation model, the first instance segmentation model, and the second instance segmentation model based on the business segmentation dataset to obtain the optimal instance segmentation model includes: Based on the target candidate mask, target instance category, and target pose parameters corresponding to the target instance in each service segmentation data, a pseudo-true annotation set corresponding to the service segmentation dataset is generated; the target candidate mask and target instance category corresponding to each target instance are determined during the zero-shot pose estimation process for each candidate instance in the service segmentation dataset. The first instance mask, first instance category, and first pose parameter corresponding to each business segmentation data output by the scene instance segmentation model are used as the first inference result. The second inference result of the first instance segmentation model for the business segmentation dataset and the third inference result of the second instance segmentation model for the business segmentation dataset are obtained. Based on the pseudo-true annotation set, the first inference result, the second inference result and the third inference result are evaluated respectively to obtain the first evaluation index, the second evaluation index and the third evaluation index corresponding to the scene instance segmentation model; The optimal instance segmentation model is determined based on the first evaluation index, the second evaluation index, and the third evaluation index.

13. The method according to claim 12, characterized in that, The step of determining the optimal instance segmentation model based on the first evaluation index, the second evaluation index, and the third evaluation index includes: If the second evaluation metric is better than the first evaluation metric and better than the third evaluation metric, then the first instance segmentation model is determined as the optimal instance segmentation model. If the third evaluation metric is better than both the first and second evaluation metric, then the second instance segmentation model is determined as the optimal instance segmentation model. If the first evaluation index is better than the second evaluation index and better than the third evaluation index, then the number of simulation data generated is increased to obtain the number of updated simulation data, and an incremental simulation dataset is generated based on the number of updated simulation data. The first instance segmentation model and the second instance segmentation model are trained using the incremental simulation dataset until the second evaluation index corresponding to the trained first instance segmentation model or the third evaluation index corresponding to the trained second instance segmentation model is better than the first evaluation index. The trained first instance segmentation model or the trained second instance segmentation model is then determined as the optimal instance segmentation model.

14. The method according to claim 1, characterized in that, The business segmentation dataset includes normal segmentation data output by the scene instance segmentation model deployed in the business scenario, and abnormal segmentation data indicated by the scene instance segmentation model reaching an abnormal triggering condition; the abnormal triggering condition includes at least one of the following: the number of instances detected by the scene instance segmentation model in the business scenario is less than a preset number threshold, the area deviation between the instance mask area detected by the scene instance segmentation model and the historical average mask area is greater than a preset deviation threshold, and the instance pose parameters detected by the scene instance segmentation model indicate that the instance is in an abnormal capture state.

15. A data processing apparatus, characterized in that, The device includes: The data acquisition module is used to acquire the business segmentation dataset generated by the scenario instance segmentation model during runtime in the business scenario; The pose estimation module is used to perform zero-shot pose estimation on each candidate instance in the business segmentation dataset if there is abnormal segmentation data in the business segmentation dataset that can trigger the upgrade of the scene instance segmentation model, to obtain the pose parameters of each candidate instance in the business segmentation dataset, and to filter the candidate instances in each business segmentation dataset according to a preset filtering strategy to obtain the target instance in each business segmentation dataset. The coordinate system transformation module is used to perform coordinate system transformation on the pose parameters of each target instance according to the three-dimensional model corresponding to each target instance, so as to obtain the target pose parameters of each target instance in the world coordinate system. The virtual rendering module is used to virtually render the target pose parameters and 3D model corresponding to the target instance in each business segmentation data to obtain H simulation data; the simulation instance in a simulation data is generated based on the target instance corresponding to the business segmentation data of that simulation data; H is a positive integer; The model training module is used to train the scene instance segmentation model based on the H simulation data to obtain a first instance segmentation model, and to train the initial instance segmentation model based on the H simulation data to obtain a second instance segmentation model; the initial instance segmentation model refers to an instance segmentation model loaded with general pre-trained weights. The model validation module is used to validate the scenario instance segmentation model, the first instance segmentation model, and the second instance segmentation model based on the business segmentation dataset to obtain the optimal instance segmentation model. The optimal instance segmentation model is determined based on the first instance segmentation model or the second instance segmentation model. The optimal instance segmentation model is used to replace the scenario instance segmentation model deployed in the business scenario.

16. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method of any one of claims 1 to 14.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1 to 14.

18. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 14.

Citation Information

Patent Citations

  • Workpiece positioning method and device based on simulation point cloud data and terminal

    CN117710460A

  • Method and apparatus with object pose estimation

    US20220198707A1