An Assembly Component Recognition Method Based on Self-Processing of HoloLens 2
By directly processing images on HoloLens 2 and performing real-time detection, the delay and transmission problems in remote communication mode are solved, and real-time identification and detection of complex structural components are realized.
Patent Information
- Application Number
- CN202211325834.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-10-27
AI Technical Summary
In the prior art, HoloLens 2, as a remote communication mode between the client and the workstation, has high time delay and multi-threaded device transmission problems, resulting in the inability to realize real-time identification of complex structural components.
By directly processing images on HoloLens 2 and performing real-time detection, using industrial cameras to acquire image data, perform image complementation processing and data enhancement, build a YOLOv5 network data set, and superimpose the recognition results into real-time scenarios, and use Windows ML API to integrate the ONNX model for real-time identification.
Real-time detection of complex structural components is realized, the information delay and multi-device transmission problems during AR remote identification is solved, and the recognition efficiency and accuracy are improved.
Smart Images

Figure CN115713701B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and particularly relates to an assembly component recognition method based on self - processing of HoloLens 2, which realizes real - time part recognition by directly using a machine learning model on HoloLens 2. Background Art
[0002] Under the background of the era of intelligent manufacturing, the high - performance assembly of major complex products such as aero - engines, which is digital, information - based, and intelligent, has attracted extensive research by domestic and foreign scholars. Currently, the assembly of fixed rocket engine core molds is mainly manual, and it has the characteristics of a large number of components and redundant information in paper process documents. To ensure the quality and efficiency of engine assembly, the assembly training and cognitive operations for new employees are very important. At present, the cognitive training of aero - engines mainly adopts the traditional mode of "mentoring the new by the old", which requires a large amount of labor cost and learning cycle.
[0003] Chinese Patent Publication No. CN114648900A discloses an immersive interactive training method and system combined with VR. This invention patent has the technical effect of ensuring the safety of employees and reducing the loss of actual operation equipment by constructing a virtual digital platform. However, this virtual training system constructs a completely digital virtual environment, lacks the perception of the real scene, there are significant differences between the digital models used and the physical models of assembly components, and there will be problems such as dizziness and discomfort in interaction during the immersive learning process.
[0004] Currently, the remote communication method with HoloLens 2 as the client and the workstation as the server effectively solves the above - mentioned problems. HoloLens 2 is used as the device for image acquisition and visualization, and the workstation is responsible for processing the deep - learning module, meeting the training needs of on - site perception and visualization. However, the remote mode of "client - server" has problems such as high time delay and multi - device multi - thread transmission, resulting in the inability to achieve real - time part recognition. Summary of the Invention
[0005] To solve the problems existing in the prior art, based on image recognition and machine learning, the present invention proposes an assembly component recognition method based on self - processing of HoloLens 2, which directly applies the built - in processor of HoloLens 2 for image acquisition and target recognition, real - time detects the natural features of parts, and superimposes the output results onto the real - world scene to intelligently update the assembly guidance information interface, breaking through the limitations and drawbacks of the "client - server" remote communication, solving the information delay and multi - device transmission problems in the AR remote recognition process, and realizing real - time detection of complex - structure components.
[0006] The technical solution of the present invention is as follows:
[0007] A method for identifying assembly parts based on self - processing of HoloLens 2, comprising the following steps:
[0008] Step 1: Automatically and continuously photograph on - site parts using an industrial camera to obtain the original pictures required for image recognition;
[0009] Step 2: Perform fuzzy threshold judgment on the original pictures collected in Step 1 and perform image frame filling processing to obtain an effective and available picture set;
[0010] Step 3: Perform one - time multiple - image storage, arrangement, scaling, and random cropping processing on the effective and available picture set obtained in Step 2 to achieve data augmentation of the picture set;
[0011] Step 4: Mark the positions and names of objects in the picture set to complete the construction of the data set;
[0012] Step 5: Extract the data set obtained in Step 4 and randomly divide it into a test set and a validation set;
[0013] Step 6: Use the pre - trained YOLO weight model to perform transfer learning on the data set, and train the model by setting the model iteration times, step size, learning rate, and number of categories to obtain a TensorFlow model containing several assembly part categories;
[0014] Step 7: Convert the TensorFlow model trained in Step 6 into an ONNX - format model;
[0015] Step 8: Use the Window ML API to integrate the ONNX - format model into a Windows application;
[0016] Step 9: Integrate the application program code built after integrating into the Windows application in Step 8 onto the UWP platform of the HoloLens device and deploy it on the HoloLens device;
[0017] Step 10: Turn on the research mode of HoloLens 2, create and initialize a media capturer and a media frame reader, capture the current image of the assembly parts in real - time at the required resolution through HoloLens 2, and convert the obtained media bitmap into the BRGA8 format;
[0018] Step 11: Call the TryAcquireLatestFrame method to obtain the latest video frame from the media frame reader and use it as the input data of the network model; at the same time, wait to release resources when the media capture stops;
[0019] Step 12: Create a new instance of the network model, which includes: a model file, a label file, a network model input and output, and a detection threshold;
[0020] Step 13: Parse the tag file, read the string, and split the original string to obtain a string array that is convenient for subsequent indexing;
[0021] Step 14: Set the network model input size and call step 11 to obtain the latest frame suitable for the network model input size;
[0022] Step 15: Use the latest video frame obtained in step 14 as the input of the network model created in step 12, use the input data tensor, cache output and time operation to perform network model reasoning, and obtain the network model detection results, including the coordinates of the detection bounding box: X, Y, Width, Height, label and confidence;
[0023] Step 16: Based on the detection bounding box coordinates in the detection result of step 15, a bounding box with a label and confidence is constructed around the target object;
[0024] Step 17: Use the extracted tag information to retrieve the component information in the assembly parts material library, obtain the relevant information of the currently detected component, and update it to display it on the UI interface of the HoloLens 2 holographic terminal.
[0025] Furthermore, in step 1, a high-resolution RGBD industrial-grade camera is installed at the assembly position where the assembly parts are located, and the camera can rotate in a circle and adjust its posture; an automatic image acquisition program is set up on the back-end processing computer, and the automatic image acquisition program has the function of automatic adjustment of the on-site environment and manual fine-tuning, and can adapt to assembly image acquisition in different on-site environments; the assembly parts are in the assembly material area, and the RGBD industrial-grade camera continuously shoots video frames through the automatic image acquisition program, transmits the images to the back-end computer, and stores each frame of the video stream in the form of a picture at a specified position on the computer.
[0026] Furthermore, in step 2, the interpolation algorithm provided by ffmpeg is used to perform frame filling processing; in step 3, the Mosaic data enhancement method is used to perform data enhancement to obtain an enhanced picture set; in step 4, the Labelimg image annotation tool is used to select regions and manually calibrate the pictures according to the parts required for the assembly process; in step 5, the holdout method is used to customize the proportion of extracted pictures, and the test set and validation set are randomly divided according to the proportion value.
[0027] Furthermore, in step 8, the integration process includes: loading the network model, creating a session, binding the model, and evaluating the model:
[0028] Step 8.1: Load the ONNX model obtained in Step 7 and place the ONNX model file into the APPX package of the Windows application;
[0029] Step 8.2: Create a session through the LearningModelSession class and bind the ONNX model to the current device;
[0030] Step 8.3: Bind the ONNX model to the Windows application through the session created in Step 8.2;
[0031] Step 8.4: Call the asynchronous evaluation method EvaluateAsync class to evaluate the input of the bound model in Step 8.3 and output the prediction result.
[0032] Furthermore, the specific steps of Step 10 include:
[0033] Step 10.1: Create a media capturer MediaCapture and a media frame reader; The MediaCapture class allows starting the built-in camera application of HoloLens2 and receiving the captured photo or video file; Check the current state of the media capturer, and judge that the condition union is: whether the media capture is empty, whether the camera is currently streaming, and whether the camera stream has been closed; If the judgment result is true, go to Step 10.2, otherwise it means that the current media frame is running normally, go to Step 11;
[0034] Step 10.2: Initialize the capturer and frame reader in Step 10.1: including: 1: Judge whether the media capture is empty, if not, destroy it; 2: Create a MediaCaptureInitializationSettings object for initializing MediaCapture, use the DeviceInformation class to list all VideoCapture devices, and then view their EnclosureLocation to obtain the corresponding device ID; 3: Initialize the MediaCapture object, instantiate the frame reader, and use the static method Convert to convert the software bitmap to the BRGA8 format.
[0035] Furthermore, in Step 12, set the detection threshold to 0.5, and create a new instance of the custom network model by asynchronously loading the ONNX machine learning model and the label file "Labels.json" in Step 7.
[0036] Furthermore, the specific process of Step 13 is:
[0037] Step 13.1: Use Resources.Load to read the Labels.json file to obtain an Object object; and convert the Object object into a TextAsset format file;
[0038] Step 13.2: Use TextReader to read a string from the TextAsset format file in Step 13.1; and use the Trim function to remove specified characters from the string; subsequently, split the original string according to the ":" character through the Split method and return the resulting string array for subsequent string indexing operations.
[0039] Further, the specific process of Step 14 is as follows: Determine the current media capture status according to Step 10 and perform initialization operations on the video stream; after the video stream works properly, receive the return value of the TryAcquireLatestFrame method in Step 11 to obtain the latest frame of the video stream.
[0040] Further, the specific process of Step 15 is as follows;
[0041] Step 15.1: Check the status of the current network model and the input frame of the network model. If they are empty, return a detection result with no target detection and a confidence level of 0; if they are not empty, go to Step 15.2;
[0042] Step 15.2: Use the latest frame of Step 14 as the input of the network model, and use input data tensors, cached outputs, and time operations to perform network model inference. Wait for and receive the network output result of the asynchronous evaluation in Step 8.4, and use the Stopwatch function to perform system timing on the network model inference process and store it as the detection time;
[0043] Step 15.3: Convert the network output result in Step 15.2 into a data type through the GetAsVectorView function. This data type includes: the number corresponding to the highest probability, the upper left corner and width and height of the detection box; and determine whether the highest probability is greater than the set threshold in Step 12. If it is greater than the set threshold, go to Step 15.4;
[0044] Step 15.4: Retrieve the label file of the string array obtained in Step 13 through the LINQ query method to obtain the index label corresponding to the maximum probability value;
[0045] Step 15.5: Return the detection result after asynchronous evaluation, including the coordinates of the detection bounding box: X, Y, Width, Hight, label, and confidence level, as well as the detection time in Step 15.2.
[0046] Further, the specific process of Step 16 is as follows;
[0047] Step 16.1: Perform boundary check on the detection box in Step 15.5 to obtain the boundary box information suitable for the current canvas, including the upper left coordinate points x1, y1 and the lower right coordinate points x2, y2.
[0048] Step 16.2: Draw a rectangle on the Texture2D canvas according to the coordinates of x1, y1, x2, and y2 obtained in Step 16.1:
[0049] (1) Draw a line between the specified two vectors p1 and p2: Use the Unity Vector2.Lerp method to perform linear interpolation between two vectors at the specified increment frac; the expression of the increment frac is as follows:
[0050]
[0051] (2) Use the frac value in the while loop; at each loop increment, increase the interpolation fraction to change the value of the linear interpolation incrementally until the final boundary condition is reached; in each iteration, set the pixels in the selected local area to the required color on the Texture2D.
[0052] (3) Perform coordinate transformation on the upper left coordinate points x1, y1 and the lower right coordinate points x2, y2 in Step 16.1 to obtain the four vector coordinates of the rectangle, and sequentially use Unity Vector2.Lerp for interpolation to draw lines between the four vectors, thereby drawing a completely connected rectangle.
[0053] Step 16.3: Create a label with a class and a confidence value: Instantiate a prefab at the world coordinates corresponding to the upper left coordinate point, transfer the pixel coordinates to the local camera coordinates through the canvas scale, and at the same time set the label and color of the boundary box component, and then update the text information of the boundary box component according to the label and confidence in Step 15.5.
[0054] Advantageous Effects
[0055] Based on image recognition and machine learning, the present invention directly performs real-time detection on the natural features of parts through HoloLens 2, and superimposes the output result onto the real scene to intelligently update the assembly guidance information interface, solving the information delay and multi-device transmission problems in the AR remote recognition process, and realizing real-time detection of complex structure parts.
[0056] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0057] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0058] Figure 1 : Schematic diagram of the present invention;
[0059] Figure 2 : Schematic diagram of how to draw a detection box on HoloLens 2. DETAILED DESCRIPTION
[0060] In view of the problems of high time delay and device multi-threaded transmission in the current remote mode part recognition method using HoloLens 2 as the client and workstation as the server, the present invention proposes an assembly parts recognition method based on HoloLens 2 self-processing, which directly detects the natural features of the parts in real time through HoloLens 2, and superimposes the output results into the real scene to intelligently update the assembly guidance information interface. The method specifically includes the following steps:
[0061] Step 1: Build an automatic image acquisition program based on existing public technology, use an industrial camera to continuously shoot on-site parts, and automatically store the video frames to a designated remote computer location to obtain the original images of the data set required for image recognition.
[0062] Step 2: Perform fuzzy threshold judgment on the original image collected in step 1, and perform image interpolation processing to obtain an effective and usable image set; for example, the interpolation algorithm provided by the ffmpeg software can be directly used to perform image interpolation processing.
[0063] Step 3: The valid and available picture set obtained in step 2 is processed by arranging, scaling, and randomly cropping multiple stored images at one time, thereby enriching the background of the detected object and improving the reliability of the picture set; for example, the Mosaic data enhancement method is used to arrange, scale, and randomly crop four stored images at one time for the valid and available picture set.
[0064] Step 4: Label the location and name of the objects in the image set to complete the construction of the data set. For example, the labelimg tool can be used to label the location and name of the objects in the image set to construct the YOLOv5 network data set.
[0065] Step 5: Extract the data set obtained in step 4, and randomly divide it into a test set and a validation set. For example, you can use the holdout method to customize the image ratio value, and then randomly divide it into a test set and a validation set.
[0066] Step 6: Perform transfer learning on the dataset using the pre-trained YOLO weights in the existing technology, and train the model by setting relevant parameters such as the number of model iterations, step size, learning rate, and number of classes to obtain a TensorFlow model containing several assembly component categories.
[0067] Step 7: Convert the TensorFlow model trained in Step 6 into the ONNX format to obtain an ONNX machine learning model. Here, the tf2onnx tool can be used for conversion.
[0068] Step 8: Integrate the ONNX-format model into the Windows application using the Window ML API. The integration process includes: loading the network model, creating a session, binding the model, and evaluating the model.
[0069] Step 9: Integrate the program code of the Windows application in Step 8 onto the UWP platform of the HoloLens device and deploy it on the HoloLens device; specifically, place the program code of the Windows application in Step 8 under the ENABLE_WINMD_SUPPORT definition and integrate it onto the UWP platform of the HoloLens device for deployment.
[0070] Step 10: Turn on the research mode of HoloLens 2, initialize the media capturer and media frame reader, capture the current image of the assembly components in real time at the required resolution through HoloLens 2, and convert the obtained media bitmap into the BRGA8 format.
[0071] Step 11: Call the TryAcquireLatestFrame method to obtain the latest video frame from the media frame reader and use it as the input data for the network model; at the same time, wait for the resources to be released when the media capture stops.
[0072] Step 12: Create a new instance of the network model, and the new instance includes: model file, label file, network model input and output, and detection threshold.
[0073] Step 13: Parse the label file, read the string, and split the original string through the Split method to obtain a string array that is convenient for subsequent indexing.
[0074] Step 14: Set the input size of the network model and call Step 11 to obtain the latest frame suitable for the input size of the network model.
[0075] Step 15: Use the latest video frame obtained in Step 14 as the input of the network model created in Step 12, and perform network model inference using the input data tensor, cached output, and time operations to obtain the network model detection results, including the coordinates of the detection bounding box: X, Y, Width, Height, label, and confidence level.
[0076] Step 16: Construct a bounding box with a label and confidence level around the target object according to the coordinates of the detection bounding box in the detection results of Step 15.
[0077] Step 17: Use the extracted label information to retrieve the component information in the assembly part material library, obtain the relevant information of the currently detected component, and update and display it on the UI interface in the HoloLens 2 holographic terminal.
[0078] Based on the above steps, the embodiments of the present invention will be described in detail below. The embodiments are exemplary and are intended to explain the present invention, but should not be construed as a limitation of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0079] As shown in the Figure 1 appendix, there are problems such as high information delay, unstable network connection, and data loss and leakage in the current remote network transmission process of "client-server". Therefore, in this embodiment, in the Universal Windows Platform (UWP), the WinML API is used to run the machine learning model locally and perform real-time evaluation, realizing the real-time recognition and evaluation processing of assembly components directly on the HoloLens 2 head-mounted display device:
[0080] Step 1: On-site image acquisition. Install a high-resolution RGBD industrial camera at the assembly position of the on-site assembly components. The position of the camera can rotate circumferentially and adjust its attitude. Build an automatic image acquisition program on the backend processing computer. The automatic image acquisition program has functions of automatic on-site environment adjustment and manual fine-tuning, and can adapt to the acquisition of assembly images in different on-site environments. The assembly components are in the assembly material area. The RGBD industrial camera continuously captures video frames through the automatic image acquisition program, transmits the images to the backend computer, and stores each frame of the video stream in the specified position of the computer in the form of a picture.
[0081] Step 2: Image frame processing. Due to the limitations of on-site working conditions and the jitter and acquisition delay during the attitude adjustment and rotation in image acquisition, some images after continuous video frame acquisition will have frame skipping and image blurring. Based on the above problems, the specific processing methods in this embodiment are as follows:
[0082] Step 2.1: For the image blurring problem in the above problem, in the automatic image acquisition program in the back-end processing computer in Step 1, set the image data blurring threshold, solve and judge the blurring value of the acquired image, and collect and store the video frame acquisition images that meet the image threshold requirements to obtain analyzed valid images.
[0083] Step 2.2: For the frame skipping phenomenon in the above problem, use the frame interpolation algorithm built in ffmpeg to perform frame filling on the valid images in Step 2.1 to obtain a complete set of valid and available video stream pictures.
[0084] Step 3: Mosaic data augmentation.
[0085] Regarding the limitations of the pose and quantity of the images collected on site, Mosaic data augmentation can enrich the background of the detected objects after processing, and can also expand the small target data of the components at the assembly site to improve the detection effect of small target components. Use Mosaic data augmentation to process the complete set of video stream pictures obtained in Step 2.2 to obtain an augmented picture set. The specific implementation plan is as follows: Arrange, scale, and randomly crop four stored images at one time, which enriches the data set and improves the robustness of the YOLO network training model.
[0086] Step 4: Perform data annotation on the Mosaic data-augmented image set in Step 3. The specific implementation plan is as follows: Import the data-augmented image set into the Labelimg image annotation tool, select regions and manually calibrate according to the components required for the assembly process. After annotation, the data is saved in YOLO format to the specified location, thus completing the construction of the YOLOv5 network data set.
[0087] Step 5: Divide the training set and validation set for the YOLOv5 network data set constructed in Step 4. The specific implementation plan is as follows: Use the hold-out method, set the custom extraction picture ratio value according to the constructed network data set and the on-site picture collection situation, and then randomly divide the test set and validation set according to the extraction picture ratio value to provide data input support for subsequent network model training.
[0088] Step 6: Train the YOLOv5 object detection model using the dataset obtained in Step 5. The YOLO network is a neural network that can process more than 60 frames per second, featuring fast recognition efficiency and high recognition accuracy. Specifically, transfer learning is performed using the dataset obtained in Step 5. Transfer learning uses a pre-trained network model as a basis to train a model for related tasks, and can adapt to the scene recognition and component feature mapping in this embodiment. In this embodiment, an officially pre-trained YOLOv5 weight model containing 80 classes is used, and then the training set and validation set in Step 5 are loaded. The model is trained by setting relevant parameters such as the number of model iterations, step size, learning rate, and number of classes, to obtain a TensorFlow network model containing 20 components.
[0089] Step 7: Since Windows ML only allows developers to add ONNX files to their UWP applications, this invention patent converts the model using the tf2onnx tool, converting the TensorFlow model in Step 6 into the ONNX format.
[0090] Step 8: Integrate the ONNX format model in Step 7 into the Windows application using the Window ML API. The integration process includes: loading the network model, creating a session, binding the model, and evaluating the model.
[0091] Step 8.1: Load the ONNX model obtained in Step 7 and place the ONNX model file into the APPX package of the Windows application; the model can be loaded by using the static method of the LearningModel class.
[0092] Step 8.2: Create a session through the LearningModelSession class and bind the ONNX model to the current device. In this embodiment, the Qualcomm Snapdragon 850 computing module of the HoloLens2 device is powerful enough to meet the computing requirements, so this computing platform is used to create a session.
[0093] Step 8.3: Bind the ONNX model to the Windows application through the session created in Step 8.2. The machine learning ONNX model has input and output features for passing information in and out. This invention patent uses the ImageFeatureValue function to set the expected input and output features of the machine learning model, thereby constructing an application based on the ONNX model in Step 7.
[0094] Step 8.4: Call the asynchronous evaluation method EvaluateAsync class to evaluate the input of the bound model in Step 8.3 and output the prediction result.
[0095] Step 9: Integrate the Windows program in Step 8 onto the UWP platform of the HoloLens device. The specific operations are as follows. To directly use the Win ML API in the unity script, it is necessary to enable the Windows runtime support in unity. In this embodiment, the #define directive is used to check whether the runtime support is enabled in the project. All the Windows application code obtained in Step 8 is placed under the ENABLE_WINMD_SUPPORT definition and deployed on the HoloLens device.
[0096] Step 10: Use HoloLens 2 to capture the current image of the assembled parts in real time at the required resolution.
[0097] In the research mode of HoloLens 2, researchers can access the original image data stream of the device, and this data stream can be locally processed or stored by the device, providing reliable image input for directly reading and locally processing images on the HoloLens 2 device.
[0098] Step 10.1: Create a MediaCapture and a media frame reader. The MediaCapture class allows starting the built-in camera application of HoloLens2 and receiving the captured photo or video file. Check the current status of the media capture, and the judgment condition union is: whether the media capture is null, whether the camera is currently streaming, and whether the camera stream has been closed. If the judgment result is true, go to Step 10.2; otherwise, it means the current media frame is running normally, and go to Step 11.
[0099] Step 10.2: Initialize the capturer and the frame reader in Step 10.1. 1. Judge whether the media capture is null. If it is not null, destroy it. 2. Create a MediaCaptureInitializationSettings object for initializing MediaCapture. Use the DeviceInformation class to list all VideoCapture devices, and then view their EnclosureLocation to obtain the corresponding device ID. 3. Initialize the MediaCapture object, instantiate the frame reader, and use the static method Convert to convert the software bitmap to the BRGA8 format.
[0100] Step 11: Call the TryAcquireLatestFrame method to obtain the latest video frame from the media frame reader, use it as the input data of the network model, and return a VideoFrame object and the software bitmap of the current frame; asynchronously stop the media capture and release the resources. When the media reading frame stops, destroy the frame reader and set the media capture value to null.
[0101] Step 12: Create a new instance of the custom network model in the application. The new instance mainly includes: a model file, a label file, network model inputs and outputs, and a detection threshold. In the implementation of this embodiment, the detection threshold is set to 0.5, and the ONNX machine learning model and the label file "Labels.json" in Step 7 are asynchronously loaded to create a new instance of the custom network model.
[0102] Step 13: Parse the imagenet labels from the Labels.json file in Step 12 to obtain a label file in the form of a string array. The specific implementation is as follows:
[0103] Step 13.1: Use Resources.Load to read the Labels.json file to obtain an Object object. To obtain the characters in this object, convert the Object object to the TextAsset format.
[0104] Step 13.2: Use TextReader to read the string from the TextAsset in Step 13.1. And use the Trim function to remove specified characters such as commas, spaces, and line breaks from the current string. Subsequently, split the original string according to the ":" character through the Split method and return the resulting string array, which can be used for subsequent string indexing operations.
[0105] Step 14: Set the input size of the network model and call Step 11 to obtain the latest frame suitable for the input size of the network model.
[0106] The specific implementation is as follows: 1. Judge the current media capture status according to Step 10 and perform the initialization operation of the video stream. 2. After the video stream works properly, receive the return value of the TryAcquireLatestFrame method in Step 11 to obtain the latest frame of the video stream.
[0107] Step 15: Continuously receive the latest frame obtained in Step 14 using the input interface of the custom network model in Step 12, and perform asynchronous evaluation through the EvaluateAsync method in Step 8.4 to obtain the network model prediction result. The specific implementation is as follows:
[0108] Step 15.1: Check the status of the current network model and the network model input frame. If they are empty, return a detection result of no target detection and a confidence level of 0. If neither is empty, go to Step 15.2.
[0109] Step 15.2: Use the latest video frame from Step 14 as the input to the network model. Perform network model inference using the input data tensor, cached output, and time operations. Wait for and receive the network output result from the asynchronous evaluation in Step 8.4, and use the Stopwatch function to time the network model inference process systemically and store it as the detection time.
[0110] Step 15.3: Convert the network output result in Step 15.2 to a data type using the GetAsVectorView function. The data type includes: the number corresponding to the highest probability, the top-left corner, width, and height of the detection box. Then, determine whether the highest probability is greater than the set threshold of 0.5 in Step 12. If it is greater than the set threshold, proceed to Step 15.4.
[0111] Step 15.4: Retrieve the label file of the string array in Step 13 through the LINQ query method to obtain the index label corresponding to the maximum probability value.
[0112] Step 15.5: Return the detection result after asynchronous evaluation, including the coordinates of the detection bounding box: X, Y, Width, Hight, label, and confidence, as well as the detection time in Step 15.2.
[0113] Step 16: Draw a detection box on the HoloLens 2 side, as Figure 2 shown. Since there is no suitable method in Unity to draw a rectangle (or bounding box) on the canvas. Therefore, in this embodiment, according to the detection box coordinates in the detection result returned by Step 15.5, the Texture2D method in Unity is extended to implement drawing lines between specified 2D coordinates on the canvas and constructing a bounding box with a label and confidence around the target object. The specific implementation scheme is as follows:
[0114] Step 16.1: Perform a boundary check on the detection box in Step 15.5 to obtain the boundary box information suitable for the current canvas, mainly including the top-left coordinate points x1, y1, and the bottom-right coordinate points x2, y2. The specific method is to use the ternary operator to determine whether X in Step 15.5 is greater than 0. If it is greater than 0, perform an int type coercion on X to obtain the top-left coordinate point x1; otherwise, add 3 extra pixels to the boundary, that is, x1 is 3. Similarly, obtain the y1, x2, and y2 after boundary check.
[0115] Step 16.2: According to the coordinates of x1, y1, x2, and y2 obtained in Step 16.1, draw a rectangle on the Texture2D canvas. This invention patent creates a Texture2Dextension method that can quickly and automatically draw a rectangle when given the top-left and bottom-right coordinates. The specific implementation scheme is as follows:
[0116] (1) Draw a line between the specified two vectors p1 and p2. Specifically, use the Unity Vector2.Lerp method to perform linear interpolation between the two vectors at the specified increment frac. Among them, the expression of the increment frac is as follows:
[0117]
[0118] (2) Use the frac value in the while loop. At each loop increment, increase the interpolation fraction to incrementally change the value of the linear interpolation until the final boundary condition is reached. In each iteration, the pixels within the selected local area will be set to the desired color on the Texture2D.
[0119] (3) Perform coordinate transformation on the upper left corner coordinate points x1, y1 and the lower right corner coordinate points x2, y2 in step 16.1 to obtain the four vector coordinates of the rectangle, and sequentially use Unity Vector2.Lerp for interpolation drawing between the four vectors to draw a completely connected rectangle.
[0120] Step 16.3: Create a label with a class and a confidence value. Instantiate a prefab at the world coordinate corresponding to the upper left corner coordinate point, and transfer the pixel coordinates to the local camera coordinate through the canvas scale, and at the same time set the label and color of the bounding box component. Then update the text information of the bounding box component according to the label and confidence in step 15.5.
[0121] Step 17: Build a UI interface, pick up the recognition result, and update the content of the UI interface.
[0122] Retrieve the assembly part material library according to the label information in step 15.5, obtain the relevant information of the currently detected component, and update and display it on the holographic end of HoloLens 2.
[0123] The construction process of the assembly part material library is as follows: In view of the complex assembly process of solid rocket engines, with numerous assembly parts and mainly manual assembly, combined with the digital requirements of the part recognition and training system of the virtual reality platform, the part material library is constructed. 1. Use Solidworks to perform physical modeling on the assembly parts and auxiliary tooling. The physical model mainly considers the assembly dimensions, assembly tolerance zones, mating accuracy, and the disassembly method of the tooling, and is converted into the fbx format through 3dmax to provide an available and effective physical model for subsequent three-dimensional visualization. 2. According to the assembly process documents, visually express the original procedural process through multi-sensory fusion such as vision, touch, and hearing, specifically including assembly animations, text information, and voice prompts, so as to establish a three-dimensional process instruction library. 3. Classify and integrate the constructed physical model and process instruction library, and build an assembly part material library, which mainly includes three parts: part attribute information, assembly guidance, and associated parts, providing data support for information retrieval after on-site part recognition.
[0124] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and purposes of the present invention.
Claims
1. An assembly component recognition method based on self - processing of HoloLens 2, characterized in that: It includes the following steps: Step 1: Use an industrial camera to automatically and continuously capture on-site components to obtain the original pictures required for image recognition; Step 2: Perform fuzzy threshold judgment on the original pictures collected in Step 1 and perform image frame filling processing to obtain an effective and available picture set; Step 3: Perform one-time multi-image storage image arrangement, scaling, and random cropping processing on the effective and available picture set obtained in Step 2 to achieve data augmentation of the picture set; Step 4: Annotate the object positions and names in the picture set to complete the construction of the data set; Step 5: Extract the data set obtained in Step 4 and randomly divide it into a test set and a validation set; Step 6: Use the pre-trained YOLO weight model to perform transfer learning on the data set, and train the model by setting the model iteration times, step size, learning rate, and number of categories to obtain a TensorFlow model containing several assembly component categories; Step 7: Convert the TensorFlow model trained in Step 6 into an ONNX format model; Step 8: Use the Window ML API to integrate the ONNX format model into a Windows application; Step 9: Integrate the application code built after integrating into the Windows application in Step 8 onto the UWP platform of the HoloLens device and deploy it on the HoloLens device; Step 10: Turn on the research mode of HoloLens 2, create and initialize a media capturer and a media frame reader, capture the current image of the assembly components in real time at the required resolution through HoloLens 2, and convert the obtained media bitmap into the BRGA8 format; Step 11: Call the TryAcquireLatestFrame method to obtain the latest video frame from the media frame reader and use it as the input data of the network model; at the same time, wait for the resources to be released when the media capture stops; Step 12: Create a new instance of the network model, and the new instance includes: a model file, a label file, the input and output of the network model, and a detection threshold; Step 13: Parse the label file, read the string, and split the original string to obtain a string array that is convenient for subsequent indexing; Step 14: Set the input size of the network model and call Step 11 to obtain the latest frame suitable for the input size of the network model; Step 15: Use the latest video frame obtained in Step 14 as the input of the network model created in Step 12, and use the input data tensor, cached output, and time operation to perform network model inference to obtain the network model detection results, including the coordinates of the detection bounding box: X, Y, Width, Hight, label, and confidence; Step 16: Build a bounding box with a label and confidence around the target object according to the detection bounding box coordinates in the detection results in Step 15; Step 17: Use the extracted label information to retrieve the component information in the assembly part material library, obtain the relevant information of the currently detected component, and update and display it on the UI interface in the HoloLens 2 holographic terminal.
2. The assembly part recognition method based on self - processing of HoloLens 2 according to claim 1, wherein: In step 1, a high-resolution RGBD industrial-grade camera is set up at the assembly position where the assembly parts are located. The camera can rotate in a circle and adjust its posture; an automatic image acquisition program is set up on the back-end processing computer. The automatic image acquisition program has the functions of automatic adjustment of the on-site environment and manual fine-tuning, and can adapt to assembly image acquisition in different on-site environments; the assembly parts are in the assembly material area, and the RGBD industrial-grade camera continuously shoots video frames through the automatic image acquisition program, transmits the images to the back-end computer, and stores each frame of the video stream in the form of a picture at a specified position on the computer.
3. The method for identifying assembled parts based on self - processing of HoloLens 2 according to claim 1, wherein: In step 2, the built-in interpolation algorithm of ffmpeg is used to perform frame filling processing; in step 3, the Mosaic data enhancement method is used to perform data enhancement to obtain an enhanced picture set; in step 4, the Labelimg image annotation tool is used to select regions and manually calibrate the pictures according to the parts required for the assembly process; in step 5, the holdout method is used to customize the image ratio, and the test set and validation set are randomly divided according to the ratio value.
4. The method for identifying assembled parts based on self - processing of HoloLens 2 according to claim 1, wherein: In step 8, the integration process includes: loading the network model, creating a session, binding the model, and evaluating the model: Step 8.1: Load the ONNX model obtained in step 7, and put the ONNX model file into the APPX package of the Windows application; Step 8.2: Create a session through the LearningModelSession class to bind the ONNX model to the current device; Step 8.3: Bind the ONNX model to the Windows application through the session created in step 8.2; Step 8.4: Call the asynchronous evaluation method EvaluateAsync class to evaluate the input of the bound model in step 8.3 and output the prediction result.
5. The method for identifying assembled parts based on self - processing of HoloLens 2 according to claim 1, wherein: The specific steps of step 10 include: Step 10.1: Create a media capturer MediaCapture and a media frame reader; the MediaCapture class allows you to start the HoloLens2 built-in camera application and receive captured photos or video files; check the current state of the media capturer, and the judgment conditions are: whether the media capture is empty, whether the camera is currently streaming, and whether the camera stream is closed; if the judgment result is true, go to step 10.2, otherwise it means that the current media frame is running normally, go to step 11; Step 10.2: Initialize the capturer and frame reader in Step 10.1, including:
1. Determine whether the media capture is empty. If it is not empty, destroy it.
2. Create a MediaCaptureInitializationSettings object for MediaCapture initialization. Use the DeviceInformation class to list all VideoCapture devices, then view their EnclosureLocation to obtain the corresponding device ID.
3. Initialize the MediaCapture object, instantiate the frame reader, and use the static method Convert to convert the software bitmap to the BRGA8 format.
6. The assembly component recognition method based on self - processing of HoloLens 2 according to claim 1, wherein: In Step 12, set the detection threshold to 0.5, and create a new instance of the custom network model by asynchronously loading the ONNX machine learning model and the label file "Labels.json" in Step 7.
7. The assembly part recognition method based on self - processing of HoloLens 2 according to claim 6, wherein: The specific process of Step 13 is as follows: Step 13.1: Use Resources.Load to read the Labels.json file to obtain an Object object, and convert the Object object to a TextAsset format file. Step 13.2: Use TextReader to read the string from the TextAsset format file in Step 13.1, and use the Trim function to delete the specified characters from the string. Then, split the original string according to the ":" character through the Split method and return the resulting string array for subsequent string indexing operations.
8. The method for identifying assembled parts based on self - processing of HoloLens 2 according to claim 5, wherein: The specific process of Step 14 is as follows: Determine the current media capture status according to Step 10 and perform the initialization operation of the video stream. After the video stream works properly, receive the return value of the TryAcquireLatestFrame method in Step 11 to obtain the latest frame of the video stream.
9. The assembly part recognition method based on self - processing of HoloLens 2 according to claim 4, wherein: The specific process of Step 15 is as follows; Step 15.1: Check the status of the current network model and the input frame of the network model. If it is empty, return the detection result of no target detection and a confidence level of 0. If both are not empty, go to Step 15.2; Step 15.2: Use the latest frame in Step 14 as the input of the network model, and use the input data tensor, cached output, and time operation to perform network model inference. Wait for and receive the network output result of the asynchronous evaluation in Step 8.4, and use the Stopwatch function to perform system timing on the network model inference process and store it as the detection time. Step 15.3: Convert the network output result in Step 15.2 to a data type through the GetAsVectorView function. The data type includes: the number corresponding to the highest probability, the upper left corner, width, and height of the detection box. And determine whether the highest probability is greater than the set threshold in Step 12. If it is greater than the set threshold, go to Step 15.4; Step 15.4: Retrieve the label file of the string array obtained in Step 13 through the LINQ query method to obtain the index label corresponding to the maximum probability value. Step 15.5: Return the detection result after asynchronous evaluation, including the coordinates of the detected bounding box: X, Y, Width, Height, label, and confidence, as well as the detection time in Step 15.
2.
10. The method for identifying assembled parts based on self - processing of HoloLens 2 according to claim 9, wherein: The specific process of Step 16 is as follows: Step 16.1: Perform a boundary check on the detection box in Step 15.5 to obtain the bounding box information suitable for the current canvas, including the upper-left coordinate points x1, y1, and the lower-right coordinate points x2, y2. Step 16.2: Draw a rectangle on the Texture2D canvas according to the coordinates of x1, y1, x2, and y2 obtained in Step 16.1: (1) Draw a line between the specified two vectors p1 and p2: Use the Unity Vector2.Lerp method to perform linear interpolation between two vectors at the specified increment frac; the expression of the increment frac is as follows: (2) Use the frac value in the while loop; at each loop increment, increase the interpolation fraction to change the value of the linear interpolation incrementally until the final boundary condition is reached; in each iteration, the pixels in the selected local area will be set to the required color on the Texture2D. (3) Perform coordinate transformation on the upper-left coordinate points x1, y1 and the lower-right coordinate points x2, y2 in Step 16.1 to obtain the four vector coordinates of the rectangle, and sequentially use Unity Vector2.Lerp for interpolation to draw lines between the four vectors, thereby drawing a completely connected rectangle. Step 16.3: Create a label with the class and confidence values: Instantiate a prefab at the world coordinates corresponding to the upper-left coordinate point, transfer the pixel coordinates to the local camera coordinates through the canvas scale, and at the same time set the label and color of the bounding box component, and then update the text information of the bounding box component according to the label and confidence in Step 15.5.
Citation Information
Patent Citations
Immersive interactive training method and system under VR (Virtual Reality) combination
CN114648900A
Small target intelligent identification system based on remote video monitoring and identification method thereof
CN111339977A
WebAR processing method based on 3D detection recognition and moving target tracking
CN112633145A