An Automatic Generation and Labeling Method for AI Training Datasets for Human Pose Multi-view Visual Recognition Based on Simulated Environment
By constructing a multi-view vision recognition AI training dataset in a simulation environment using digital twin technology, the problem of insufficient datasets in industrial environments is solved, generating an efficient and accurate dataset suitable for machine vision recognition training in complex industrial scenarios.
Patent Information
- Application Number
- CN202411850276.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-12-16
AI Technical Summary
There is a lack of AI training datasets for human posture machine vision recognition in industrial environments. Existing datasets are insufficient in content, accuracy and efficiency cannot meet the requirements, and traditional datasets have large annotation errors, making them difficult to apply to complex industrial scenarios.
By using digital twin technology, a multi-view visual recognition AI training dataset is built in a simulation environment. By precisely configuring camera parameters, simulating industrial scenes and lighting, and customizing digital characters and movements, the dataset is automatically output with annotation files and images, ensuring its authenticity and richness.
It generates high-quality, high-efficiency datasets, improves the accuracy and processing speed of machine vision recognition AI, reduces costs, and is suitable for complex industrial scenarios.
Smart Images

Figure CN119888024B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for automatically generating and labeling AI training datasets for multi-view visual recognition of human postures in simulated environments, particularly for the automatic generation and labeling of human postures and scenes in industrial environments. This method can be used in digital twin environments to address the problem of a lack of simulated human posture datasets in industrial scenarios. Background Technology
[0002] With the rapid development of information technology, digital twin technology has become a cutting-edge research topic in academia and industry. Digital twins use digital means to establish dynamic virtual models of physical entities that are multi-dimensional, multi-temporal scale, multi-disciplinary, and multi-physical quantity, in order to simulate and characterize the attributes, behaviors, and rules of physical entities in the real environment.
[0003] Against this backdrop, "virtual digital humans," as an important branch of digital twin technology, are receiving increasing attention from the industry. A "virtual digital human," or simply a "digital human," refers to a digital avatar of a real-world physical person generated in a virtual world using various technological means. Currently, digital human technology has been widely applied in multiple fields such as gaming, entertainment, and sports.
[0004] In the field of machine vision, training human pose recognition AI requires a large amount of multi-view, multi-scene human motion data and key point annotations. In recent years, with technological advancements, significant progress has been made in training machine vision for human pose recognition, and the required datasets have become more comprehensive and extensive. However, due to the specificities of the industrial sector and the complexity of human movements in industrial environments, traditional datasets in this area are severely lacking in content, resulting in limited scalability and hindering the development of visual recognition training. Specifically, this manifests in the following aspects:
[0005] First, human postures in industrial environments are more complex and diverse, human movements are more likely to reach extreme states, and are more professional, which is significantly different from human movements in daily life.
[0006] Secondly, the production and working environment in the industrial sector is relatively complex, and obtaining relevant data is time-consuming and laborious, with some data also having confidentiality issues.
[0007] Finally, the annotation information of human body key points needs to be more accurate, but the current key point annotation of datasets is mostly done manually, which has large errors and low efficiency, making it difficult to meet the requirements of industrial fields.
[0008] Existing datasets are primarily derived from living environments, and their scenes, lighting, and human movements differ significantly from those in industrial environments. Using these datasets for visual recognition training in industrial settings will lead to severe biases. Furthermore, the accuracy and speed of annotation information in existing datasets cannot meet the high-precision and high-efficiency requirements of industrial environments. Therefore, establishing a human pose dataset suitable for industrial scenarios is particularly urgent.
[0009] In the industrial sector, acquiring datasets regarding human posture, lighting conditions, and environmental factors is often challenging due to its specific nature. However, digital twin technology offers a solution to this problem. Digital twin technology can simulate near-real-world effects, allowing for the free combination of elements such as human models, movements, environment, lighting, and camera angles. Existing research has shown that datasets generated in this way can rival real-world training sets in realism and diversity. Furthermore, using these datasets for machine vision recognition training achieves the same level of performance as real-world training sets.
[0010] Meanwhile, the industrial sector has much stricter requirements for the accuracy and efficiency of keypoint annotation. Current datasets primarily employ top-down keypoint detection algorithms for multi-person keypoint detection. This approach first determines the approximate location of a single person using object detection algorithms, then performs keypoint detection on that person within the image detection bounding box, thus transforming the multi-person keypoint detection problem into a single-person keypoint detection problem. While this method performs well in terms of accuracy, its processing speed is relatively slow. Although it meets the needs of the current dataset, it still lags significantly behind the high standards of speed and accuracy required in the industrial sector. Therefore, the industrial sector needs more efficient and accurate algorithms to improve keypoint detection performance and meet its dual requirements of high efficiency and high accuracy.
[0011] Given the unique environment of industrial settings, the complexity and specialization of human movements, and the requirements for accuracy and efficiency in labeling key human body points, training data for machine vision recognition in industrial fields is relatively scarce, limiting the in-depth development of visual recognition training in this area. Therefore, this paper proposes an automatic generation and labeling method for simulation environments of AI training data for human posture machine vision recognition. This method is applicable to digital twin environments and aims to address the problem of insufficient human posture simulation datasets in industrial scenarios. Summary of the Invention
[0012] The purpose of this invention is to address the problem of the lack of AI training datasets for human posture machine vision recognition in industrial environments, and to develop a method for automatically generating and labeling AI training datasets for human posture multi-view vision recognition based on a simulated environment.
[0013] The technical solution of this invention is:
[0014] A method for automatically generating and labeling AI training datasets for human pose multi-view visual recognition based on a simulated environment, characterized by:
[0015] First, camera parameter matching is performed. The cameras within the digital twin system are precisely configured based on the key parameters of the actual visual inspection camera (including focal length, field of view, and resolution). This step ensures that the cameras in the twin system maintain performance consistency with the actual cameras, making the output images and data more closely resemble the real world. Furthermore, the number and location of cameras within the twin system can be flexibly adjusted according to the actual application scenario, comprehensively covering detailed attention to different distances and angles to achieve complete correspondence with the real world.
[0016] Second, industrial scene modeling and lighting simulation. The digital twin system accurately models and constructs scenes based on actual industrial environments, providing a variety of different industrial scenarios to meet training needs. The twin system can also freely set lighting conditions to realistically simulate the color and intensity of light on-site, enhancing the realism of the simulated environment.
[0017] Third, digital character customization and motion simulation are performed. In setting up the digital characters, meticulous adjustments are made to their height, body type, gender, age, and clothing to ensure a more realistic appearance and increase character diversity. Data from datasets or annotation files from visual inspection outputs is acquired and imported into the digital twin system. Corresponding code is written to enable the characters within the twin system to move according to the set parameters. The data input interval is adjusted to ensure the continuity and smoothness of the character's movements. For occlusion phenomena between people and between people and objects in the real world, specialized code is written and appropriate settings are implemented in the twin system to ensure accurate judgment of occlusion relationships.
[0018] Fourth, automatic output and labeling of training data. The twin system outputs the labeled files and images required for machine vision training. The labeled files include descriptions of the person's identity features (such as height, body type, gender, age, occupation, etc.), joint data, joint occlusion relationships, bounding boxes of the person, scene features, screen area, and the camera's intrinsic and extrinsic parameter matrices. By writing specific code, the twin system can automatically generate this data and allows users to adjust the output time interval to optimize output efficiency. During data output, parameters such as scene settings, camera viewpoint, person model, and clothing features can be autonomously adjusted or automatically changed, ensuring a richer and more diverse output dataset. Furthermore, the system can automatically identify and filter out highly similar actions by calculating the similarity between preceding and following actions, while ensuring that the differences between actions are not too large. This helps improve the generalization ability of the dataset.
[0019] The beneficial effects of this invention are:
[0020] This invention addresses the challenge of insufficient training datasets for human posture simulation in industrial scenarios within a digital twin environment. It proposes an automatic generation and labeling method for training AI training datasets for human posture multi-view visual recognition based on a simulated environment. This method utilizes digital twin technology, synchronizing with real camera settings to simulate real-world industrial scenes and lighting conditions, simulating human movements, and automatically generating information-rich annotation files and images. This provides a large amount of high-quality, high-efficiency datasets for machine vision AI training of human posture in industrial scenarios. This method has the following significant advantages:
[0021] First, this method can construct various industrial scenarios, create diverse human body models, simulate a wide range of human movements, set different lighting conditions, and vary the number, distance, and angle of cameras. These factors can be freely combined to enrich the dataset content and continuously increase the amount of data, ensuring the robustness, scalability, and generalization ability of the dataset. This effectively solves the problem of obtaining human posture simulation datasets in industrial scenarios.
[0022] Second, the datasets generated through digital twin technology are no less realistic than existing datasets. In a twin environment, the industrial environment and its lighting conditions are recreated based on real-world scenarios, and a wide variety of human models are customized to simulate human movements smoothly and coherently. These steps make the datasets more closely resemble the real world and enhance their applicability in training AI for machine vision recognition.
[0023] Third, the digital twin system can automatically output information such as the annotation data of human joints, occlusion relationships, human body bounding boxes, and screen area. Furthermore, during data output, it automatically identifies and filters out highly similar actions by calculating similarity, while ensuring that the differences between actions are not too large. This optimized dataset not only improves data accuracy and processing efficiency but also significantly enhances the performance of machine learning models compared to traditional methods, ensuring the adaptability and reliability of the models in practical applications.
[0024] Fourth, in the twin system, the digital human moves based on the data output by visual detection, generating more precise key point motion data. This data can be input into the twin system again to make the digital human move, undergoing data iteration. The twin system can select the output data precision; through this iterative process, the accuracy of the digital human's motion data is further improved.
[0025] Fifth, digital twin technology reduces reliance on physical experiments and on-site data collection through virtual simulation, thereby saving significant time and money. It obtains high-quality, high-volume, and high-efficiency datasets at a lower cost, accelerating the development of visual recognition training in industry. Attached Figure Description
[0026] Figure 1 Images showcasing some industrial scenarios within the twin system.
[0027] Figure 2 A schematic diagram of the digital human skeleton in twin software.
[0028] Figure 3 A schematic diagram illustrating the iterative motion data of a digital human in a digital twin environment.
[0029] Figure 4 A schematic diagram illustrating the determination of occlusion relationships in a digital twin environment.
[0030] Figure 5 A diagram showing the skeletal joints during the similarity calculation process.
[0031] Figure 6 A diagram illustrating angle calculation during the similarity calculation process.
[0032] Figure 7 Similarity algorithm flowchart.
[0033] Figure 8 Flowchart of a method for automatically generating and labeling simulation environments for AI training data in human-machine posture machine vision recognition. Detailed Implementation
[0034] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0035] like Figure 1-8 As shown.
[0036] A method for automatically generating and labeling simulation environments for AI training data in human-machine posture recognition machine vision, the implementation of which includes the following steps:
[0037] Step 1 involves configuring the camera settings within the digital twin system. First, based on the relevant information of the actual visual inspection camera, internal parameters and external parameters such as position and angle are configured to ensure that the camera in the digital twin environment can accurately simulate and reflect the characteristics of the actual camera. Then, the physical camera settings are enabled for the twin camera. Physical camera parameters include focal length, field of view, sensor type and size, lens offset, etc. These internal parameter settings allow for a better simulation of the imaging effects of a real camera.
[0038] Furthermore, the number and position of cameras within a digital twin system can be flexibly adjusted according to the needs of the actual application scenario. This adjustment helps ensure the realism and accuracy of the output images, enabling the digital twin system to provide more reliable data support when simulating industrial scenarios. For ease of management, all cameras in the twin system are placed under a common parent object. This layout not only allows individual cameras to move and rotate independently but also greatly simplifies the overall layout and adjustment of the camera group. During the output dataset process, the position and angle of the cameras can be adjusted, making their shooting range and perspective more diverse, more realistically simulating the camera state in real-world scenes, and enriching the image content in the dataset.
[0039] In machine vision deep learning, especially human pose recognition, understanding the camera's intrinsic and extrinsic parameter matrices is crucial. These parameters play a decisive role in correcting geometric distortion and restoring depth information, directly affecting the accuracy of pose recognition and the effectiveness of model training. To better support human pose machine vision recognition, the annotation files output by the Siamese dataset must include the camera's intrinsic and extrinsic parameter matrices.
[0040] The camera's intrinsic parameter matrix describes its internal parameters, including focal length, principal point coordinates, and image distortion. It is typically a 3x3 matrix, commonly denoted as K. The intrinsic parameter matrix maps 3D points in the camera coordinate system to 2D pixel coordinates on the image plane. It enables operations such as camera calibration, image correction, and projection of 3D point clouds into the image. The expression for the intrinsic parameter matrix is:
[0041]
[0042] Among them, f x f y The length of the focal length in the x and y axes is described using pixels. (C x C y The coordinates () represent the center point coordinates of the image, specifying the origin of the image plane and used to remove translational distortion. In a twin system, given the camera's focal length (f), sensor size, and field of view (FOV), f can be calculated. x f y First, the horizontal and vertical field of view (FOV) of the camera were calculated. x and FOV y Then, f can be calculated using the following formula. x f y .
[0043]
[0044] Where imageWidth and imageHeight are the image resolution (in pixels), S x S y This refers to the physical dimensions of the sensor (in millimeters). C x C y The center point is obtained by calculating the image size.
[0045] The extrinsic parameter matrix of a camera describes the camera's position and orientation in the world coordinate system. It is typically a 4x4 matrix, consisting of a 3x3 rotation matrix and a 3x1 translation vector. The rotation matrix represents the rotation of the camera coordinate system relative to the world coordinate system, with each column representing the representation of an axis (X, Y, Z axis) in the camera coordinate system in the world coordinate system. The translation vector represents the position of the origin of the camera coordinate system in the world coordinate system. These two elements together define the pose of the camera coordinate system relative to the world coordinate system. The expression for the extrinsic parameter matrix is:
[0046]
[0047] Simplified expression:
[0048]
[0049] Where R is the rotation matrix and t is the translation vector. In a twin system, a 4x4 matrix can be constructed to obtain the camera's rotation matrix and translation vector and combine them.
[0050] Step 2 involves accurately modeling and constructing the actual industrial scenario within the twin system. For scenarios involving confidentiality, the models are anonymized before use to ensure data security. The models within the twin system are designed to match the actual scenario in terms of color, size, position, and lighting variations.
[0051] In a digital twin system, materials and shaders are used to achieve the desired object shape and display effect. Materials describe the surface details of 3D objects, including color, smoothness, transparency, and metallic material settings; shaders define the various codes (such as vertex shaders and fragment shaders), attributes (such as textures), and instructions (rendering and label settings) required for rendering. The process includes creating materials, creating shaders, assigning shaders to the materials created in the previous step, assigning materials to the objects to be rendered, and adjusting the shader properties in the material panel to achieve satisfactory results. By combining materials and shaders, digital twin systems can provide more realistic, flexible, and efficient rendering solutions, helping to improve the accuracy and realism of simulations. Figure 1 This showcases some industrial scenarios within the twin system.
[0052] Setting up lighting in a twin system is a crucial step in achieving realistic rendering effects. The twin system offers various light source types and lighting settings. Here are the steps: First, add a light source. The twin system supports three types of light sources: simulated sunlight, directional lights with parallel rays, point lights simulating rays emanating from a single point in all directions, and spotlights simulating cone-shaped rays emanating from a single point. Second, adjust the light source properties, setting the color, brightness, influence range, light attenuation, and shadows. Next, configure the lighting mode. The twin system provides two lighting modes: real-time global illumination and baked global illumination. Select and set them accordingly based on the actual situation. Finally, adjust materials, use light probes, set reflections, and optimize lighting performance as needed. Through these steps, various lighting effects can be created in the twin system to meet different needs and more closely resemble the real world.
[0053] During the output of the dataset, corresponding code (which can be automatically generated by the system or written manually, and can be implemented using existing technologies, the same below) is written and related operations are performed to automatically or manually set the lighting effects in the scene to conform to the natural laws of lighting changes. At the same time, the appearance of the model can be changed by altering the materials and shaders, or different industrial scenes can be switched to enrich the scene settings in the dataset and enhance the diversity and practicality of the dataset.
[0054] Step 3 involves customizing and simulating the motion of the digital human model within the digital twin system. In this system, the digital human consists of a skin and a skeleton. The skin covers the skeleton and is responsible for shaping the digital human's form and appearance, while the skeleton drives its movement. Through precise movement of the skeleton, the skin deforms accordingly, giving the digital human a lifelike appearance and fluid dynamic performance.
[0055] The digital human's appearance is configured, including height, body shape, gender, age, and clothing. Materials and shaders are used to ensure a more realistic character design. During the output dataset process, the digital human's appearance can be modified based on the written code and changes to materials and shaders, enriching the character representation within the dataset.
[0056] A skeletal digital human model can be directly imported into a twin system, but it's necessary to ensure that the model's skeletal points match the predefined skeletal points in the twin system. The skeleton is organized in a tree-like hierarchical structure, such as... Figure 2 As shown, the entire skeletal structure has a root bone, and all other bones are directly or indirectly connected to the root bone, forming the entire skeletal framework of the character model. Generally, each bone has two matrices: one is the initial transformation matrix representing the initial position of the bone; the other is the combined transformation matrix used to perform various transformations on the bone, thereby achieving character movement.
[0057] In a twin system, the translational transformation of the digital human is represented by the coordinate changes of its joints, while the rotational transformation is controlled and represented by quaternions. A quaternion is a mathematical tool for representing rotation in three-dimensional space. It consists of one real part and three imaginary parts, in the form w + xi + yj + zk, where w is the real part and x, y, and z are the imaginary parts. The main advantage of quaternions is that they can concisely and efficiently represent three-dimensional rotations and avoid the gimbal locking problem, ensuring the accuracy and stability of the physical simulation.
[0058] Digital human motion simulation is a key step in this method. By reproducing the actions in the input file, the digital human can move smoothly in a twin environment, ultimately outputting accurate joint motion data. This transforms the low-accuracy data in the input file into high-accuracy data, meeting the accuracy requirements of industrial datasets. Furthermore, by continuously improving the accuracy requirements, the dataset is iterated. The flowchart is shown below. Figure 3 As shown.
[0059] Taking the twin software Unity as an example, the specific steps are as follows:
[0060] (3.1) Use the SQLite4Unity3d plugin to obtain data from the dataset or the annotation file output by visual inspection, and import it into the Siamese system. This annotation file includes the location data of each keypoint of the digital human and the quaternion data between the keypoints;
[0061] (3.2) Map the keypoint data in the file to the joints of the digital human. Drive the digital human's movement using the quaternion rotation method `Quaternion.Slerp(QuaternionfromRotation,QuaterniontoRotation,floatt)` to ensure the continuity and smoothness of the rotation. Here, `fromRotation` is the starting quaternion, `toRotation` is the ending quaternion, and `t` is the interpolation ratio, ranging from [0,1]. To ensure the digital human's movements are smooth and natural, the motion effect can be optimized by adjusting the time interval of the data input. Note that not all keypoints are equipped with quaternions; quaternions are only applied when a keypoint has child objects. This is because in a digital twin system, the movement of the parent object automatically drives its child objects, while the movement of the child objects does not affect the parent object.
[0062] (3.3) During the movement of the digital human, situations such as mutual occlusion and being occluded by objects may occur. It is necessary to determine the occlusion situation in order to output it to the annotation file. The flowchart for determining the occlusion relationship is as follows. Figure 4 As shown.
[0063] In Unity, a ray is emitted to detect whether an object has a collider or trigger. Objects without a collider component cannot be detected. The detection function is...
[0064] `Physics.Raycast(Rayray, outhit, float maxDistance, layerMask)`, where `ray` is the ray beam, whose emission position and direction can be set; `hit` is the object information hit by the ray; `maxDistance` is the ray distance, defaulting to infinite distance; and `layerMask` is the ray mask, indicating which layer was detected, defaulting to detecting all layers. Considering real-world scenarios, factory workshops often involve collaborative work by multiple people, hence the frequent movement of varying numbers of people in dummy scenes. Multi-person collision detection includes interactions between people and between people and objects, requiring detection of each joint of each person. Therefore, each person is on a different layer, while objects are on the same layer. Colliders need to be set on the dummy skin and objects to ensure proper ray detection. Occlusion detection can be simplified to a ray-object intersection test, with the following detection formula:
[0065] Ray = StartPoint + t·Direction
[0066] Where StartPoint is the starting point of the ray, Direction is the direction of the ray (which needs to be standardized), and t is the distance from the starting point to the intersection point. If t is between 0 and 1, it means that the ray intersects the object between the starting and ending points, i.e., occlusion exists. If t equals 1, it means that the ray successfully detected the ending point, i.e., there is no occlusion. The occlusion relationship in the dataset is represented as follows: 1 represents a joint being occluded, and 2 represents a joint not being occluded, expressed as...
[0067]
[0068] (3.4) After the digital human moves smoothly, its clothing can be manually or automatically changed by writing code and modifying materials and shaders. Appropriate time intervals and precision are set to output the coordinates of each joint and the occlusion relationship. This process not only provides crucial data support for multi-person keypoint detection algorithms but also enables dataset iteration through continuous improvement in output precision. By analyzing the digital human's joint data, we can determine the maximum and minimum values of the data on the X, Y, and Z axes, thereby obtaining the bounding box of the person and outputting it to the annotation file, greatly facilitating the implementation of multi-person detection algorithms.
[0069] Current datasets primarily employ top-down keypoint detection algorithms for multi-person keypoint detection. This approach first defines the bounding box of a single person and then performs keypoint detection on that person within the image detection bounding box, transforming the multi-person keypoint detection problem into a single-person keypoint detection problem. While this method achieves good keypoint accuracy, its processing speed is relatively slow, and it is prone to missed or duplicate detections. Therefore, in twin software, using more accurate data to quickly calculate the human body bounding box not only solves the problem of missed or false detections when people occlude each other or have complex poses, but also significantly improves processing efficiency.
[0070] Step 4 involves automatically outputting the annotation files and images required for machine vision training via the twin system. The annotation files include detailed information such as character features (height, build, gender, age, occupation, etc.), joint data, joint occlusion relationships, character bounding boxes, scene features, screen area, and camera intrinsic and extrinsic parameter matrices. The annotation files and images can be automatically generated by writing corresponding code within the twin system, and the output time interval can be adjusted to improve efficiency. During the output process, the scene, camera viewpoint, character model, and clothing features can be modified autonomously or automatically, resulting in a richer and more diverse output dataset.
[0071] In the field of machine vision recognition, similarity detection of digital human actions is a crucial task. If a dataset contains a large number of highly similar actions, it can lead to overfitting during machine vision recognition. In this case, the model may focus excessively on the detailed features of these specific actions, neglecting broader, more generalized action features, thus affecting its performance and generalization ability on new data. Conversely, if the actions in the dataset are too diverse, the model may lack sufficient feature information to learn effectively, resulting in underfitting. In this situation, the model's performance on both the training and test sets may be unsatisfactory because it cannot capture enough features to represent all actions. Therefore, to improve the model's generalization ability and performance, it is necessary to balance the similarity and diversity of actions in the dataset output and model training, performing similarity detection during dataset output to ensure that the model can learn features that are both specific and generalized.
[0072] During the output process, similarity detection is performed using the cumulative sum of angles from preceding and following actions. The smaller the weighted cumulative sum of angles from preceding and following actions, the higher the similarity between the two actions. There are a total of 19 skeletal joints, such as... Figure 5 As shown, its information is denoted as J, and expressed as J = (j0, j1, ..., j...). 18 The position information of each key point includes three dimensions, denoted as j. i , is represented as:
[0073] j i =(x i y i , z i ), i∈[0,18.
[0074] Because the lengths of different segments in a digital human skeleton are inconsistent, the weight of each joint should be different when calculating similarity. Therefore, a weight-adaptive skeleton similarity algorithm is proposed. The specific process is as follows:
[0075] (4.1) Plan the skeleton segments. There are 18 skeleton segments in total, with 19 joints. Since the order of the start and end points of the skeleton segments affects the algorithm results, the start and end points of each segment are fixed. The number of the start joint of each segment is denoted as s, the number of the end joint is denoted as e, k represents the k-th segment, and the length is denoted as l. k (p s ,p e ),k∈
[0076] [0,17], s∈[0,18], e∈[0,18], detailed information on the starting and ending points of the skeleton line segments is shown in Table 1.
[0077] l k (p s ,p e The calculation formula is:
[0078]
[0079] Table 1. Numbering of the start and end points of each line segment in the skeleton.
[0080]
[0081] (4.1) Define the weight W of each joint. There are 19 weights for each of the 19 joints: W = (w0, w1, w2, ..., w 17 ,w 18 ).
[0082] (4.2) Adaptive Weight Determination. The principle behind this is that the tail joints of longer bone sides in the template information have larger weights, while the tail joints of shorter bone sides have smaller weights. If the hip joint is selected as the reference point, its corresponding weight is 0, based on the joint coordinates of the previous movement. For example, specifically:
[0083]
[0084] (4.3) Calculate action similarity. First, calculate the similarity of the previous action using vector coordinates. and the next action The angle between the corresponding bones Figure 6 This is a schematic diagram for angle calculation.
[0085]
[0086] Then, calculate the total angle θ(P,Q) according to the above weights for each included angle value:
[0087]
[0088] The weighted total angle θ(P,Q) is a number greater than or equal to 0, and the larger the value, the lower the similarity of the corresponding actions. A large weighted total angle indicates low similarity between consecutive actions, possibly due to excessively large intervals between action outputs, leading to significant differences between actions. In this case, shortening the interval between action outputs can reduce the differences between actions. Conversely, a small weighted total angle indicates high similarity between consecutive actions, possibly due to excessively small intervals between action outputs or insufficient amplitude of the character's movements. In this case, appropriately increasing the interval between action outputs or skipping the current action and selecting the next action for output can be considered. The similarity algorithm flowchart is as follows: Figure 7 As shown.
[0089] The parts not covered in this invention are the same as or can be implemented using existing technologies.
Claims
1. A method for automatically generating and labeling AI training datasets for multi-view visual recognition of human pose based on a simulation environment, characterized by: First, camera parameter matching is performed; the cameras in the digital twin system are configured according to the key parameters of the actual visual inspection cameras; this ensures that the cameras in the twin system are consistent with the actual cameras in terms of performance, so that the output images and data are closer to the real world. In addition, the number and location of cameras in the twin system should be flexibly adjusted according to the actual application scenario to fully cover detailed attention to different distances and angles in order to achieve a complete correspondence with the real world; Second, industrial scene modeling and lighting simulation; the digital twin system accurately models and constructs scenes based on actual industrial scenarios, providing a variety of different industrial scenarios to meet training needs; the twin system can also freely set lighting conditions to realistically simulate the color and intensity of on-site light, enhancing the realism of the simulated environment; Third, digital character customization and motion simulation are performed. In setting up the digital characters, their height, body type, gender, age, and clothing are meticulously adjusted to ensure a more realistic appearance and increase diversity. Data from datasets or annotation files from visual inspection outputs are acquired, imported into the digital twin system, and corresponding code is written to enable the characters within the twin system to move according to the settings. The data input interval is adjusted to ensure the continuity and smoothness of the character's movement. For occlusion phenomena between people and between people and objects in the real world, specialized code is written and corresponding settings are made in the twin system to ensure accurate judgment of occlusion relationships. Fourth, automatic output and labeling of training data; the twin system outputs the labeled files and images required for machine vision training. The labeled files include descriptions of the person's identity features, joint data, joint occlusion relationships, person bounding boxes, scene features, screen range, and information required for the camera's intrinsic and extrinsic parameters. By writing specific code, the twin system can automatically generate this data and allows users to adjust the output time interval to optimize output efficiency. During data output, scene settings, camera perspective, person models, and their clothing features can be autonomously adjusted or automatically modified, ensuring that the output dataset is richer and more diverse in content. In addition, the system can automatically identify and filter out highly similar actions by calculating the similarity between actions before and after, while ensuring that the differences between actions are not too large, which helps improve the generalization ability of the dataset.
2. The method according to claim 1, characterized in that: The key parameters of the camera include focal length, field of view, and resolution.
3. The method according to claim 1, characterized in that: The description of the person's identity includes height, build, gender, age, and occupation.
4. The method according to claim 1, characterized in that: The number and location of cameras within the digital twin system should be flexibly adjusted according to the needs of the actual application scenario. All cameras in the twin system should be placed under a common parent object. During the output dataset process, the position and angle of the cameras should be adjusted to make their shooting range and perspective more diverse, more realistically simulating the camera state in the real scene, and enriching the image content in the dataset. The annotation file output by the twin dataset should include the camera's intrinsic parameter matrix and extrinsic parameter matrix. The camera's intrinsic parameter matrix is used to describe the camera's internal parameters, including the camera's focal length, principal point coordinates, and image distortion information. The intrinsic parameter matrix is in the form of a 3x3 matrix, denoted as K. The intrinsic parameter matrix maps three-dimensional points in the camera coordinate system to two-dimensional pixel coordinates on the image plane. The intrinsic parameter matrix enables camera calibration, image correction, and projection operations from 3D point clouds to images; the expression for the intrinsic parameter matrix is: Among them, f x f y The length of the focal length in the x and y axes is described using pixels. (C x C y () represents the coordinates of the center point of the image. These coordinates specify the origin of the image plane and are used to remove translational distortion. In a twin system, f is calculated from the camera's focal length f, sensor size, and field of view (FOV). x f y First, the horizontal and vertical field of view (FOV) of the camera are calculated. x and FOV y Then, calculate f using the following formula. x f y ; Where imageWidth and imageHeight are the image resolution, in pixels, S x S y It refers to the physical dimensions of the sensor, measured in millimeters, C. x C y The center point is obtained by calculating the image dimensions. The extrinsic parameter matrix of the camera describes the position and orientation of the camera in the world coordinate system. It is a 4x4 matrix consisting of a 3x3 rotation matrix and a 3x1 translation vector. The rotation matrix represents the rotation of the camera coordinate system relative to the world coordinate system, with each column representing the representation of one of the X, Y, or Z axes in the camera coordinate system in the world coordinate system. The translation vector represents the position of the origin of the camera coordinate system in the world coordinate system. These two elements together define the pose of the camera coordinate system relative to the world coordinate system. The expression for the extrinsic parameter matrix is: Simplified expression: Where R is the rotation matrix and t is the translation vector; in the twin system, a 4x4 matrix is constructed, and the camera's rotation matrix and translation vector are combined.
5. The method according to claim 1, characterized in that: When modeling industrial scenarios, for scenarios involving confidentiality, the models should be anonymized before use to ensure data security. In the twin system, the model's color, size, position, and lighting changes should be consistent with the actual scene. In the twin system, the required object shape display effect is achieved through materials and shaders. Materials describe the surface details of 3D objects, including color, smoothness, transparency, and metal material settings. Shaders define various codes, attributes, and instructions required for rendering. The process includes creating materials, creating shaders, and assigning the shaders to the materials created in the previous step. Then, the materials are assigned to the objects to be rendered, and the shader attributes are adjusted in the material panel. The twin system provides multiple light source types and lighting settings. The implementation steps are as follows: First, add light sources. The twin system supports three types of light sources: simulating sunlight, directional light with parallel rays, point light simulating light emanating from a single point in all directions, and spotlight simulating conical light emanating from a single point. Secondly, adjust the light source properties, setting the color, brightness, influence range, light attenuation, and shadows. Next, configure the lighting mode; the twin system provides two lighting modes: real-time global illumination and baked global illumination, which should be selected and set according to the actual situation. Finally, adjust the lighting materials, use lighting probes, set reflections, and optimize lighting performance as needed. During the output dataset process, write corresponding code and perform related operations to automatically or manually set the lighting effects in the scene to conform to the natural laws of lighting changes. At the same time, change the model appearance by modifying materials and shaders, or switch between different industrial scenes to enrich the scene settings in the dataset and enhance its diversity and practicality.
6. The method according to claim 1, characterized in that: in In a twin system, a digital human consists of a skin and a skeleton; the skin covers the skeleton and is responsible for shaping the shape and appearance of the digital human, while the skeleton is responsible for driving the movement of the digital human; through the precise movement of the skeleton, the skin will deform accordingly. The digital human's appearance is configured, including height, body type, gender, age, and clothing. Materials and shaders are used to ensure a more realistic character design. During dataset output, the digital human's appearance should be modified based on the written code and changes in materials and shaders to enrich the character representation in the dataset. Digital human models with skeletons can be directly imported into the twin system, but it's necessary to ensure that the model's skeletal points match the predefined skeletal points in the twin system. The skeleton is organized in a tree-like hierarchical structure, with a root bone in the entire structure, and all other bones directly or indirectly connected to the root. The skeleton forms the entire skeletal framework of the character model; each bone has two matrices: an initial transformation matrix representing the initial position of the bone, and a transformation matrix used to transform and combine the bones, thereby realizing the character's movement; in the twin system, the translation transformation of the digital human is represented by the coordinate values of each joint, and the rotation transformation of the digital human is controlled and represented by quaternions; a quaternion is a mathematical tool used to represent rotation in three-dimensional space, which consists of one real part and three imaginary parts, in the form w+xi+yj+zk, where w is the real part and x, y, and z are the imaginary parts.
7. The method according to claim 1, characterized in that: The aforementioned digital human motion simulation reproduces the actions in the input file, enabling the digital human to move smoothly in a twin environment. It ultimately outputs precise joint motion data, transforming low-accuracy data from the input file into high-accuracy data to meet the precision requirements of industrial datasets. Furthermore, by continuously improving precision requirements, the dataset is iterated. The specific steps are as follows: 7.1 Use the SQLite4Unity3d plugin to obtain data from the dataset or the annotation file output by visual inspection, and import it into the twin system; the annotation file includes the location data of each key point of the digital human and the quaternion data between the key points; 7.2 The key point data in the file are mapped to the joints of the digital human. The digital human is driven to move by the quaternion rotation Quaternion.Slerp(QuaternionfromRotation,QuaterniontoRotation,floatt) to ensure the continuity and smoothness of the rotation. Here, fromRotation is the starting quaternion, toRotation is the ending quaternion, and t is the interpolation ratio, which takes values in the range [0,1]. In order to ensure that the digital human's movements are smooth and natural, the motion effect is optimized by adjusting the time interval of the data input. Quaternions are only applied when there are sub-objects at the joints. 7.3 During the movement of the digital human, there will be situations where they occlude each other or are occluded by objects. It is necessary to determine the occlusion situation in order to output it to the annotation file; In Unity, a ray is emitted to detect whether an object has a collider or trigger. Objects without collision components cannot be detected. The detection function is `Physics.Raycast(Rayray, outhit, float maxDistance, layerMask)`, where `ray` is the ray, setting its emission position and direction, `hit` is the object information hit by the ray, `maxDistance` is the ray distance (default is infinite), and `layerMask` is the ray mask, indicating which layer was detected (default is all layers). Multi-person collision detection includes collisions between people and between people and objects. It requires detection of each joint of each person, so each person is on a different layer while objects are on the same layer. Colliders need to be set on the digitized human skin and objects to ensure the ray can detect them correctly. Occlusion detection is simplified to a ray-object intersection test, and the detection formula is as follows: Ray = StartPoint + t·Direction Where StartPoint is the starting point of the ray, Direction is the direction of the ray to be standardized, and t is the distance from the starting point to the intersection point; if t is between 0 and 1, it means that the ray intersects the object between the starting point and the ending point, i.e., there is occlusion; if t equals 1, it means that the ray successfully detected the ending point, i.e., there is no occlusion; the occlusion relationship of the dataset is represented as: Where: shelter=1 indicates that the joint is covered, shelter=2 indicates that the joint is not covered; 7.4 After the digital human moves smoothly, the clothing of the digital human can be changed manually or automatically by writing code and changing materials and shaders. Appropriate time intervals and precision are set to output the coordinates of each joint point of the digital human and the occlusion relationship. By analyzing the joint data of the digital human, the maximum and minimum values of the data on the X, Y and Z axes of the human are determined, and then the bounding box of the human is obtained and output to the annotation file.
8. The method according to claim 1, characterized in that: in During the output process, similarity detection is performed using the cumulative sum of angles from preceding and following actions; the smaller the weighted cumulative sum of angles from preceding and following actions, the higher the similarity between the two actions; there are a total of 19 skeletal joints, whose information is denoted as J, represented as J = (j0, j1, ..., j...). 18 The position information of each key point includes three dimensions, denoted as j. i , is represented as: j i (x i y i ,z i ), i∈[0.18]: Because the lengths of different segments in a digital human skeleton are inconsistent, the weight of each joint should be different when calculating similarity. Therefore, a weight-adaptive skeleton similarity algorithm is proposed, the specific process of which is as follows: 8.1 Planning the skeleton segments: There are 18 skeleton segments in total, with 19 joints. Since the order of the starting and ending points of the skeleton segments affects the algorithm results, the starting and ending points of each segment are fixed. The starting joint of each segment is denoted as 's', the ending joint as 'e', 'k' represents the k-th segment, and its length is denoted as 'l'. k (p s ,p e ), k∈[0,17], s∈[0,18], e∈[0,18]; l k (p s ,p e The calculation formula is: 8.2 Define the weight W of each joint; 19 joints correspond to 19 weights: W=(w0,w1,w2,…,w 17 ,In 18 ); 8.3 Adaptive Weight Determination; The principle behind this is that the tail joints of longer bone sides in the template information have larger weights, while the tail joints of shorter bone sides have smaller weights. If the hip joint is selected as the reference point, its corresponding weight is 0. Let the joint coordinates of the previous action be... The weights are adaptively set to: 8.4 Calculate action similarity; first, calculate the similarity of the previous action using vector coordinates. and the next action The angle between the corresponding bones Then, calculate the total angle θ(P,Q) according to the above weights for each included angle value: The weighted total angle θ(P,Q) is a number greater than or equal to 0, and the larger the value, the smaller the similarity of the corresponding actions. When the weighted total angle is large, it means that the similarity between the previous and subsequent actions is low, and the interval between action outputs should be shortened to reduce the difference between actions. Conversely, when the weighted total angle is small, it indicates that the similarity between the previous and subsequent actions is high, and the interval between action outputs should be appropriately increased, or the current action should be skipped and the next action should be selected for output.
Citation Information
Patent Citations
Method for constructing shuttle tanker output operation scene data set based on virtual simulation
CN115511923A
Visual algorithm simulation data set generation and verification method based on twinborn scene
CN115903541A