Method and device for detecting apparent defects in real time in mixed reality environment

By combining augmented reality technology and deep learning methods in a mixed reality environment, and combining YOLOv8 detection model with mixed reality equipment, the shortcomings of traditional appearance defect detection technology in real time and accuracy are solved, and efficient and real-time appearance defect detection effects are achieved.

CN120102564APending Publication Date: 2025-06-06STATE GRID SHANDONG ELECTRIC POWER CO CONSTR CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510076965.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional appearance defect detection technology has shortcomings in real-time, accuracy and portability, especially in field environments where complex lighting, angle and location changes are difficult to meet detection needs.

Method used

Combining augmented reality (AR) technology and deep learning methods, the YOLOv8 detection model is combined with mixed reality devices to achieve real-time detection of appearance defects. The method includes activating the defect detection function of the mixed reality device, initializing the camera, establishing a network connection, transmitting the image to the local server for detection, and returning the detection result to the device through the network, and finally projecting the detection result in the mixed reality view.

Benefits of technology

It realizes real-time detection of appearance defects in a mixed reality environment, improves real-time and accuracy of detection, reduces delay problems in traditional methods, improves user experience and detection efficiency, and is suitable for a variety of complex detection scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120102564A_ABST
    Figure CN120102564A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for detecting appearance defects in real time in a mixed reality environment, and belongs to the technical field of mixed reality and image processing.The method comprises the steps that the defect detection function of mixed reality equipment is started, and a system camera of the mixed reality equipment is initialized; establishing network connection between the mixed reality equipment and a local server; controlling a mixed reality equipment system camera to collect and encode an image, and transmitting the encoded image to a local server; after receiving the image, the local server decodes the image, inputs the image into a YOLOv8 model for defect detection, and returns a detection result to the mixed reality equipment through the network; and the mixed reality equipment carries out 2D-3D coordinate conversion on the bounding box of the defect in the detection result, and projects the detection result into a mixed reality view. According to the invention, real-time processing of defect detection and instant feedback of results are realized, efficient front-end display and interaction are realized, and user experience and detection efficiency are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and a device for real-time detection of appearance defects in a mixed reality environment, belonging to the technical field of mixed reality and image processing. Background Art

[0002] The surface of the product is inspected by a machine vision inspection system to identify surface defects that do not meet quality requirements. This inspection method improves the overall quality of the product and customer satisfaction while ensuring that the appearance quality and functionality of the product meet customer requirements. Industrial appearance defect inspection covers a variety of industries, including but not limited to the inspection of spots, pits, scratches, defects and other defects on the surface of the inspection object. This inspection usually uses machine vision inspection technology to obtain the surface image of the product and use the corresponding image processing algorithm to extract the feature information of the image. Subsequently, the surface defects are located and identified based on these feature information.

[0003] Traditional appearance defect detection technology mainly relies on fixed cameras and background processing equipment, which has great limitations in terms of real-time and convenience. In modern construction acceptance scenarios, such as power transmission and transformation projects, it is usually necessary to identify and analyze the appearance quality of accessories, component surfaces, and text materials in real time. Traditional methods cannot effectively deal with complex factors such as lighting, angles, and positions on site, and cannot meet the high requirements for real-time, accuracy, and portability during on-site acceptance.

[0004] In order to solve these shortcomings, the current augmented reality (AR) technology combines with deep learning methods to provide a new solution. By using artificial neural networks (ANN) for reasoning and prediction, AR can perform real-time detection of the appearance of components and materials at the construction site, and identify appearance quality defects or text material information. This real-time reasoning technology relies on fully collected or augmented image data, especially representative on-site data of power transmission and transformation project construction acceptance, and adjusts or customizes the network model to make the model more suitable for complex acceptance environments. Therefore, the present invention proposes a method for appearance defect detection based on computer vision and deep learning. Summary of the invention

[0005] In order to solve the above problems, the present invention proposes a method and device for real-time detection of appearance defects in a mixed reality environment.

[0006] The technical solution adopted by the present invention to solve the technical problem is: In a first aspect, an embodiment of the present invention provides a method for real-time detection of appearance defects in a mixed reality environment, comprising the following steps: Step S1, starting the defect detection function of the mixed reality device and initializing the system camera of the mixed reality device; Step S2, establishing a network connection between the mixed reality device and the local server; Step S3, controlling the camera of the mixed reality device system to collect and encode images, and transmitting the encoded images to the local server; Step S4, after receiving the image, the local server decodes and inputs the image into the YOLOv8 model for defect detection, and returns the detection result to the mixed reality device through the network; the YOLOv8 model outputs the bounding box, defect category and confidence of each defect; In step S5, the mixed reality device performs 2D to 3D coordinate conversion on the boundary box of the defect in the detection result, and projects the detection result into the mixed reality view using the internal parameters of the mixed reality device system camera and the depth sensor data.

[0007] As a possible implementation of this embodiment, the step S1, starting the defect detection function of the mixed reality device and initializing the system camera of the mixed reality device, includes: Step S11, starting the defect detection function of the mixed reality device and checking the first-person perspective function of the mixed reality head display; Step S12, initializing a system camera of the mixed reality device, wherein the system camera of the mixed reality device is a camera facing the external environment, and the mixed reality device can access and control the use and stop of the camera; Step S13, check whether the mixed reality device is started normally.

[0008] As a possible implementation of this embodiment, the step S2, establishing a network connection between the mixed reality device and the local server, includes: Use the functions in the System.Net and System.Net.Sockets libraries to establish a TCP connection between the mixed reality device and the local server.

[0009] As a possible implementation of this embodiment, the use of functions in the System.Net and System.Net.Sockets libraries to establish a TCP connection between the mixed reality device and the local server includes: Define the IP address and port number of the local server; Create a Socket instance so that the mixed reality device and the local server can communicate using the TCP protocol. The mixed reality device attempts to connect to the local server and enters the communication state after the connection is successfully established.

[0010] As a possible implementation of this embodiment, the step S3, controlling the camera of the mixed reality device system to capture and encode images, and transmitting the encoded images to the local server, includes: Click the camera button on the control panel of the mixed reality device to start the camera to capture the image in the current field of view. The captured image data is stored in texture format. Read the current field of view image data from the texture data buffer of the mixed reality device; Use the JPEG encoding library to encode texture format image data into JPEG format and generate a byte stream; Creating a broadcast packet containing meta information for the byte stream image data, the broadcast packet including data type, image size information, a timestamp and image data; Send the generated advertising packet to the API endpoint specified by the local server.

[0011] As a possible implementation of this embodiment, in step S4, after receiving the image, the local server decodes and inputs the image into the YOLOv8 model for defect detection, and returns the detection result to the mixed reality device through the network, including: The local server decodes the received encoded image data into an image matrix in a standard RGB format; The decoded image is input into the YOLOv8 model for defect detection. The YOLOv8 model analyzes each detection area in the image, identifies the existing appearance defects and generates the corresponding bounding box location, defect category and confidence level. All detection results are integrated into a detection result data packet, and the detection result data packet is transmitted back to the mixed reality device through a TCP connection.

[0012] As a possible implementation of this embodiment, the output of the YOLOv8 model includes: The bounding box position of each detected defect area is expressed as ( , , , ) format, where ( , ) is the coordinate of the upper left corner of the defect area, ( , ) is the coordinate of the lower right corner of the defect area; Defect category information corresponding to each bounding box; Confidence: The confidence score indicates the reliability of the test result. A test result with a higher confidence score means that the model is more certain about the defect. The detection result data packet includes: Bounding Box Position:( , , , ); Defect categories: including but not limited to category labels for scratches and cracks; Confidence Score: Indicates the confidence level.

[0013] As a possible implementation of this embodiment, in step S5, the mixed reality device performs 2D to 3D coordinate conversion on the boundary box of the defect in the detection result, and projects the detection result into the mixed reality view using the camera intrinsic parameters and depth sensor data of the mixed reality device system, including: Extracting the 2D bounding box coordinates, category label, and confidence information of each defect from the received inspection result data packet; Using the camera intrinsic matrix K Convert the bounding box from 2D pixel coordinates to 3D coordinates; Use the camera's inverse projection formula to map the 2D bounding box coordinates to 3D space points in the camera's coordinate system; The depth sensor is used to obtain the depth data of each pixel, the depth information is matched with the 3D coordinates in the camera coordinate system, and the 3D position in the camera coordinate system is converted into the world coordinate system of the mixed reality device, so that the defect can be aligned with the real position of the inspected object in the actual space.

[0014] As a possible implementation of this embodiment, the camera intrinsic parameter matrix K for: , in, and The camera is Axis and The focal length along the axis, and are the principal point coordinates of the camera respectively; The inverse projection formula of the camera is: , in,( , ) is the 2D coordinate in the bounding box, point Z Represents the depth value, obtained from the depth sensor data.

[0015] In a second aspect, an embodiment of the present invention provides a device for real-time detection of appearance defects in a mixed reality environment, comprising: An initialization module, used to start the defect detection function of the mixed reality device and initialize the system camera of the mixed reality device; A network connection module, used to establish a network connection between the mixed reality device and the local server; An image encoding module is used to control the camera of the mixed reality device system to collect and encode images, and transmit the encoded images to a local server; A defect detection module is used for decoding the image after the local server receives it and inputting the image into the YOLOv8 model for defect detection, and returning the detection result to the mixed reality device through the network; the YOLOv8 model outputs the bounding box, defect category and confidence of each defect; The detection result projection module is used for the mixed reality device to perform 2D to 3D coordinate conversion on the boundary box of the defect in the detection result, and to project the detection result into the mixed reality view using the internal parameters of the mixed reality device system camera and the depth sensor data.

[0016] The beneficial effects of the technical solution of the embodiment of the present invention are as follows: The present invention combines the YOLOv8 detection model with mixed reality devices to achieve real-time processing of defect detection and instant feedback of results; users can intuitively see the detection results in a mixed reality environment, and the spatial location and category information of the defects are directly superimposed on the actual object, reducing the delay problem in traditional detection methods. Compared with systems that rely on background processing, the present invention achieves efficient front-end display and interaction, greatly improving user experience and detection efficiency.

[0017] Traditional detection systems usually require fixed cameras and detection devices, which limits the application scope of the detection scene. However, the mixed reality device of the present invention can move freely and quickly locate, which is not only suitable for fixed scenes, but also can be applied to dynamic and complex industrial and construction sites, so that the device can adapt to the defect detection needs of various angles and positions.

[0018] Due to the limited computing power of mixed reality devices, the present invention adopts the design of local server for YOLOv8 model inference, transferring high-load detection tasks from the device side to the local server, significantly reducing the computing pressure of the device. This lightweight end-side solution of the present invention enables the detection model to have higher complexity and accuracy without affecting the performance of the device, thereby extending the battery life and usage time of the device.

[0019] The present invention has high adaptability and scalability, and can be applied to a variety of detection scenarios, such as industrial quality inspection, building inspection, and equipment maintenance. Through simple configuration and equipment movement, the system can adapt to the requirements of different application scenarios, reducing the installation and maintenance costs of detection equipment, and providing a highly flexible solution for detection work in various industries. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1is a flow chart of a method for real-time detection of appearance defects in a mixed reality environment according to an exemplary embodiment; Figure 2 It is a schematic structural diagram of a device for real-time detection of appearance defects in a mixed reality environment according to an exemplary embodiment. DETAILED DESCRIPTION

[0021] In order to more clearly illustrate the technical features of the solution of the present invention, the present invention is described in detail below through specific implementation methods and in conjunction with the accompanying drawings.

[0022] like Figure 1 As shown, a method for real-time detection of appearance defects in a mixed reality environment provided by an embodiment of the present invention includes the following steps: Step S1, starting the defect detection function of the mixed reality device and initializing the system camera of the mixed reality device; Step S2, establishing a network connection between the mixed reality device and the local server; Step S3, controlling the camera of the mixed reality device system to collect and encode images, and transmitting the encoded images to the local server; Step S4, after receiving the image, the local server decodes and inputs the image into the YOLOv8 model for defect detection, and returns the detection result to the mixed reality device through the network; the YOLOv8 model outputs the bounding box, defect category and confidence of each defect; In step S5, the mixed reality device performs 2D to 3D coordinate conversion on the boundary box of the defect in the detection result, and projects the detection result into the mixed reality view using the internal parameters of the mixed reality device system camera and the depth sensor data.

[0023] As a possible implementation of this embodiment, the step S1, starting the defect detection function of the mixed reality device and initializing the system camera of the mixed reality device, includes: Step S11, starting the defect detection function of the mixed reality device and checking the first-person perspective function of the mixed reality head display; Step S12, initializing a system camera of the mixed reality device, wherein the system camera of the mixed reality device is a camera facing the external environment, and the mixed reality device can access and control the use and stop of the camera; Step S13, check whether the mixed reality device is started normally.

[0024] As a possible implementation of this embodiment, the step S2, establishing a network connection between the mixed reality device and the local server, includes: Use the functions in the System.Net and System.Net.Sockets libraries to establish a TCP connection between the mixed reality device and the local server.

[0025] As a possible implementation of this embodiment, the use of functions in the System.Net and System.Net.Sockets libraries to establish a TCP connection between the mixed reality device and the local server includes: Define the IP address and port number of the local server; Create a Socket instance so that the mixed reality device and the local server can communicate using the TCP protocol. The mixed reality device attempts to connect to the local server and enters the communication state after the connection is successfully established.

[0026] As a possible implementation of this embodiment, the step S3, controlling the camera of the mixed reality device system to capture and encode images, and transmitting the encoded images to the local server, includes: Click the camera button on the control panel of the mixed reality device to start the camera to capture the image in the current field of view. The captured image data is stored in texture format. Read the current field of view image data from the texture data buffer of the mixed reality device; Use the JPEG encoding library to encode texture format image data into JPEG format and generate a byte stream; Creating a broadcast packet containing meta information for the byte stream image data, the broadcast packet including data type, image size information, a timestamp and image data; Send the generated advertising packet to the API endpoint specified by the local server.

[0027] As a possible implementation of this embodiment, in step S4, after receiving the image, the local server decodes and inputs the image into the YOLOv8 model for defect detection, and returns the detection result to the mixed reality device through the network, including: The local server decodes the received encoded image data into an image matrix in a standard RGB format; The decoded image is input into the YOLOv8 model for defect detection. The YOLOv8 model analyzes each detection area in the image, identifies the existing appearance defects and generates the corresponding bounding box location, defect category and confidence level. All detection results are integrated into a detection result data packet, and the detection result data packet is transmitted back to the mixed reality device through a TCP connection.

[0028] As a possible implementation of this embodiment, the output of the YOLOv8 model includes: The bounding box position of each detected defect area is expressed as ( , , , ) format, where ( , ) is the coordinate of the upper left corner of the defect area, ( , ) is the coordinate of the lower right corner of the defect area; Defect category information corresponding to each bounding box; Confidence: The confidence score indicates the reliability of the test result. A test result with a higher confidence score means that the model is more certain about the defect. The detection result data packet includes: Bounding Box Position:( , , , ); Defect categories: including but not limited to category labels for scratches and cracks; Confidence score: Usually between 0 and 1, indicating a confidence level of 0~100%. For example, 0.8 indicates an 80% confidence level.

[0029] As a possible implementation of this embodiment, in step S5, the mixed reality device performs 2D to 3D coordinate conversion on the boundary box of the defect in the detection result, and projects the detection result into the mixed reality view using the camera intrinsic parameters and depth sensor data of the mixed reality device system, including: Extracting the 2D bounding box coordinates, category label, and confidence information of each defect from the received inspection result data packet; Using the camera intrinsic matrix K Convert the bounding box from 2D pixel coordinates to 3D coordinates; Use the camera's inverse projection formula to map the 2D bounding box coordinates to 3D space points in the camera's coordinate system; The depth sensor is used to obtain the depth data of each pixel, the depth information is matched with the 3D coordinates in the camera coordinate system, and the 3D position in the camera coordinate system is converted into the world coordinate system of the mixed reality device, so that the defect can be aligned with the real position of the inspected object in the actual space.

[0030] As a possible implementation of this embodiment, the camera intrinsic parameter matrix K for: , in, and The camera is Axis and The focal length along the axis, and are the principal point coordinates of the camera respectively; The inverse projection formula of the camera is: , in,( , ) is the 2D coordinate in the bounding box, point Z Represents the depth value, obtained from the depth sensor data.

[0031] like Figure 2 As shown, an embodiment of the present invention provides a device for real-time detection of appearance defects in a mixed reality environment, comprising: An initialization module, used to start the defect detection function of the mixed reality device and initialize the system camera of the mixed reality device; A network connection module, used to establish a network connection between the mixed reality device and the local server; An image encoding module is used to control the camera of the mixed reality device system to collect and encode images, and transmit the encoded images to a local server; A defect detection module is used for decoding the image after the local server receives it and inputting the image into the YOLOv8 model for defect detection, and returning the detection result to the mixed reality device through the network; the YOLOv8 model outputs the bounding box, defect category and confidence of each defect; The detection result projection module is used for the mixed reality device to perform 2D to 3D coordinate conversion on the boundary box of the defect in the detection result, and to project the detection result into the mixed reality view using the internal parameters of the mixed reality device system camera and the depth sensor data.

[0032] The specific implementation process of the present invention for real-time detection of appearance defects in a mixed reality environment is as follows.

[0033] 1. Start the defect detection function and initialize the mixed reality device system camera.

[0034] Enable the mixed reality defect detection function to view the first-person perspective function of the mixed reality headset; Initialize the mixed reality device system camera, which is a camera facing the external environment. The mixed reality device can access and control the use and stop of the camera; Check whether the mixed reality device is started normally and is ready to connect to the local server for image acquisition.

[0035] 2. After the system is initialized, a network connection is established with the local server to realize the transmission of image data and the reception of detection results.

[0036] After the system is initialized, it uses the functions in the System.Net and System.Net.Sockets libraries to establish a TCP connection with the local server to achieve reliable transmission and reception of data.

[0037] First, create a Socket instance and use the TCP protocol for communication. Define the IP address and port number of the local server to ensure that the device can correctly connect to the local server.

[0038] After initialization is completed, the system attempts to connect to the local server. After the connection is successfully established, the device and the local server can enter the communication state and prepare for subsequent data transmission.

[0039] 3. The user device triggers image acquisition through the photo button on the virtual control panel and transmits the image to the local server.

[0040] The user clicks the camera button on the control panel of the mixed reality device, and the system starts the camera to collect the image in the current field of view. After the image is collected, the following preprocessing steps are performed: After the user clicks the photo button, the system starts the camera to capture the image of the current field of view, and the generated image data is stored in texture format. Since directly transmitting the texture format is not convenient for the local server to process, we need to encode the image and create broadcast information to transmit it to the local server.

[0041] Read texture data: Read the current field of view image data from the texture data buffer of the mixed reality device.

[0042] Encode to JPEG format: Use the JPEG encoding library to encode texture format image data into JPEG format to compress the image size.

[0043] Generate JPEG byte stream: The encoded JPEG format image will generate a byte stream (byte array), which is the image data transmitted to the local server.

[0044] A broadcast packet containing meta information is then created for the JPEG-encoded image data so that the local server can identify and process the transmitted data.

[0045] Data Type: Add an identifier to indicate that the data type is a JPEG image.

[0046] Image size information: includes image resolution information (such as 640x640), so that the local server can correctly restore the image size during decoding.

[0047] Timestamp: Add an acquisition timestamp so that the backend can record the image acquisition time to ensure the timeliness of the data.

[0048] Image data: JPEG-encoded byte stream data is attached as the main content of the broadcast packet.

[0049] The generated broadcast information is sent to the API endpoint specified by the local server. After receiving the data, the local server can decode the JPEG image and perform subsequent YOLOv8 detection processing.

[0050] 4. After the local server receives the image, it decodes it and inputs the image into the YOLOv8 model for defect detection. The model outputs the bounding box, defect category and confidence of each defect, and finally returns the standardized detection results to the device through the network.

[0051] After receiving the image data transmitted by the device, the local server decodes it and inputs the image into the YOLOv8 model for defect detection. The model outputs the bounding box, defect category, and confidence level of each defect. After the detection is completed, the local server standardizes the results and returns them to the device via the network. The specific implementation process includes the following steps: The image data received by the local server is usually a byte stream in JPEG format. The local server first decodes the byte stream data into an image matrix in standard RGB format to meet the input requirements of the YOLOv8 model. The decoding process ensures that the clarity and resolution of the image meet the detection requirements of the model.

[0052] The decoded and adjusted image is fed into the YOLOv8 model for defect detection. The YOLOv8 model analyzes each detection area in the image, identifies possible appearance defects, and generates the corresponding bounding box location, defect category, and confidence.

[0053] The output of the YOLOv8 model contains the bounding box locations of each detected defect region. Bounding boxes are usually given in the form of ( , , , ) format, ( , ) represents the coordinate of the upper left corner of the defect area, ( , ) represents the coordinates of the lower right corner of the defect area. These coordinates indicate the specific location of the defect in the image, which is convenient for subsequent coordinate conversion and display.

[0054] The defect category information corresponding to each bounding box (such as honeycomb) helps to classify different types of defects. These category labels are distinguished by color in the device's mixed reality view, improving the user's recognition efficiency.

[0055] The YOLOv8 model outputs a confidence score to indicate the reliability of the detection result. A detection result with a higher confidence score means that the model is more certain about the defect. The detection results are filtered to eliminate results with a confidence score below a certain threshold to avoid false positives.

[0056] All test results are combined into a single data package, including the following: Bounding Box Position:( , , , ); Defect category: category label, such as scratches, cracks, etc.; Confidence score: usually between 0 and 1, such as 0.8 means 80% confidence.

[0057] The detection result data packet is transmitted back to the mixed reality device through the TCP connection. After receiving the data packet, the device unpacks and parses the result to prepare for the subsequent 3D display and user interface display.

[0058] 5. After receiving the detection results, the mixed reality device converts the coordinates of the defect bounding box from 2D to 3D, and projects the results into the mixed reality view using the device's camera intrinsics and depth sensor data.

[0059] After receiving the detection results sent back by the local server, the mixed reality device converts the coordinates of the defect bounding box in the 2D image into 3D space coordinates so as to accurately display the defect location in the mixed reality view.

[0060] The specific conversion and projection process includes the following steps.

[0061] After receiving the detection result data packet, the mixed reality device parses the bounding box position, defect category and confidence of each defect. In order to display the detection results in the mixed reality view, the system converts the pixel coordinates of the 2D bounding box into 3D coordinates in the device space.

[0062] The system first extracts the 2D bounding box coordinates, category label, and confidence information of each defect from the received inspection result data packet. The bounding box coordinates are ( , , , ), indicating the pixel location of the defect area; the category label and confidence information will be used for subsequent display and distinction in the mixed reality view.

[0063] Camera intrinsic matrix for mixed reality devices K Contains the focal length and principal point position of the camera, which is used to convert from 2D pixel coordinates to 3D coordinates. Camera intrinsic matrix K It is expressed as: , in, and The camera is Axis and The focal length along the axis, and are the principal point coordinates of the camera respectively.

[0064] Use the camera's inverse projection formula to map the 2D bounding box coordinates to 3D space points in the camera coordinate system. , )Use the formula: , in, Z Represents the depth value, obtained from the depth sensor data.

[0065] The depth sensor of the mixed reality device provides depth data for each pixel. The system combines the depth information with the 3D coordinates in the camera coordinate system to obtain the actual spatial position of each defect. The depth data 𝑍 enables the projected position to have accurate distance perception in the mixed reality view.

[0066] The 3D position in the camera coordinate system is converted to the world coordinate system of the mixed reality device, so that the defect can be aligned with the real position of the detected object in real space. This conversion can be achieved by the spatial positioning module of the device to adjust the camera coordinate position to the reference coordinate system of the device.

[0067] Pseudo code: SetResult method flow (the 3D position in the camera coordinate system is converted to the world coordinate system of the mixed reality device and displayed on the mixed reality device): Function SetResult(width,height,x,y,imgbytes): # Check if x and y contain enough coordinate information If (x.Length<2 OR y.Length is null): Return # Get the camera's world matrix and projection matrix cameraToWorldMatrix=photoCaptureFrame.TryGetCameraToWorldMatrix() projectionMatrix=photoCaptureFrame.TryGetProjectionMatrix(Camera.main.nearClipPlane,Camera.main.farClipPlane) # Calculate the 3D world coordinates of the center point of the bounding box centerPoint=LocatableCameraUtils.PixelCoordToWorldCoord(cameraToWorldMatrix,projectionMatrix,cameraResolution,(x[0]+x[1] / 2,y[0]+y[1] / 2)) # Initialize the ShapeData array of 4 points to store the 3D world coordinates of the four corners of the bounding box ShapeData[0]=LocatableCameraUtils.PixelCoordToWorldCoord(cameraToWorldMatrix,projectionMatrix,cameraResolution,(x[0],y[0])) ShapeData[1]=LocatableCameraUtils.PixelCoordToWorldCoord(cameraToWorldMatrix,projectionMatrix,cameraResolution,(x[0]+x[1],y[0])) ShapeData[2]=LocatableCameraUtils.PixelCoordToWorldCoord(cameraToWorldMatrix,projectionMatrix,cameraResolution,(x[0]+x[1],y[0]-y[1])) ShapeData[3]=LocatableCameraUtils.PixelCoordToWorldCoord(cameraToWorldMatrix,projectionMatrix,cameraResolution,(x[0],y[0]-y[1])) # Use Raycast to detect collision and see if any object intercepts the ray from frezeePos to the center point If(Physics.Raycast(frezeePos,centerPoint-frezeePos,hit,100,(1<<31))): # If a collision is detected, update the center point position to the collision point centerPoint_hit=hit.point # Check or create a LineRenderer object If(LineObj is null): LineObj = New GameObject() # Get or add a LineRenderer component LineRenderer LR=LineObj.GetComponent <linerenderer>() If (LR is null): LR=LineObj.AddComponent <linerenderer>() Initialize LineRenderer (material, color, width, loop, worldSpace, positionCount) # Calculate and adjust the coordinates of the 4 points of ShapeData to adapt to the collision point halfDiagonal=Distance(centerPoint,ShapeData[0]) halfDiagonal_hit=Distance(frezeePos,centerPoint_hit)* halfDiagonal / Distance(frezeePos,centerPoint) centerPoint_displacement=centerPoint_hit-centerPoint Adjust ShapeData points with centerPoint_displacement andhalfDiagonal_hit # Set the 4 position points of LineRenderer LR.SetPosition(0, ShapeData[0]) LR.SetPosition(1, ShapeData[1]) LR.SetPosition(2, ShapeData[2]) LR.SetPosition(3, ShapeData[3]) Else: # If no collision is detected, create or get a Quad to display the image If (Quad is null): Quad = CreatePrimitive(PrimitiveType.Quad) # Initialize the Quad's renderer and material Renderer quadRenderer=Quad.GetComponent <renderer>() quadRenderer.material=New Material(shader) # Create and load a Texture2D object using imgbytes data Texture2D targetTexture=New Texture2D(width,height) targetTexture.LoadImage(imgbytes) targetTexture.wrapMode=TextureWrapMode.Clamp # Set the Quad's material texture quadRenderer.sharedMaterial.SetTexture("_MainTex", targetTexture) # Calculate the position and rotation of the Quad to align it with the camera position=cameraToWorldMatrix.GetColumn(3)-cameraToWorldMatrix.GetColumn(2) rotation=Quaternion.LookRotation(-cameraToWorldMatrix.GetColumn(2),cameraToWorldMatrix.GetColumn(1)) # Resize the Quad to fit the image resolution imageCenterDirection,imageTopLeftDirection,imageTopRightDirection,imageBotLeftDirection,imageBotRightDirection=Calculate directions w=Distance(imageTopLeftDirection,imageTopRightDirection) h=Distance(imageTopLeftDirection,imageBotLeftDirection) Quad.transform.localScale=(w,h,1) # Apply the camera matrix and projection matrix to the material quadRenderer.sharedMaterial.SetMatrix("_WorldToCameraMatrix",cameraToWorldMatrix.inverse) quadRenderer.sharedMaterial.SetMatrix("_CameraProjectionMatrix",projectionMatrix) # Set the position and rotation of the Quad Quad.transform.position=position Quad.transform.rotation=rotation End Function.

[0068] Pseudocode structure explanation: Image detection and coordinate conversion: Convert 2D bounding box coordinates to 3D world coordinates using camera and projection matrices; Collision detection: determine whether there is an object blocking between frezeePos and centerPoint, and if so, adjust the display position; Draw the detection box: If there is a collision, use LineRenderer to display the bounding box; if there is no collision, create a Quad to display the image and set the material and texture.

[0069] To ensure that the detection results are accurately positioned in the user's field of view, the system will adjust the display position of the 3D coordinates according to the device's posture and the user's current viewing angle, so that the detection results can always be aligned with the actual defect location.

[0070] Using the 3D position after coordinate conversion, the system generates a labeling box for each defect in the mixed reality view, and applies different color or shape marks based on the defect category and confidence information to facilitate quick identification by users.

[0071] Based on the characteristics of the power transmission and transformation project site, the present invention adopts data augmentation technology to expand the data set to cover a variety of construction conditions and environmental changes, and utilizes or optimizes the YOLO series network model for model training to ensure that it has good accuracy while ensuring real-time performance, thereby forming a standardized model that takes into account both recognition speed and accuracy. The present invention not only improves the accuracy of detection, but also effectively meets the needs of rapid feedback at the construction site.

[0072] The present invention realizes real-time, flexible and efficient defect detection, overcomes the shortcomings of traditional methods in equipment fixity, real-time performance, computing pressure and environmental adaptability, and provides a practical and efficient technical solution for appearance defect detection. It is suitable for various scenarios such as quality inspection of auxiliary parts and component appearance at the construction site of power transmission and transformation projects, and can realize efficient and intuitive defect detection and information feedback through augmented reality equipment.

[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.< / renderer> < / linerenderer> < / linerenderer>

Claims

1. A method for real-time detection of appearance defects in a mixed reality environment, characterized in that: The steps include: Step S1, starting the defect detection function of the mixed reality device and initializing the system camera of the mixed reality device; Step S2, establishing a network connection between the mixed reality device and the local server; Step S3, controlling the camera of the mixed reality device system to collect and encode images, and transmitting the encoded images to the local server; Step S4, after receiving the image, the local server decodes and inputs the image into the YOLOv8 model for defect detection, and returns the detection result to the mixed reality device through the network; the YOLOv8 model outputs the bounding box, defect category and confidence of each defect; In step S5, the mixed reality device performs 2D to 3D coordinate conversion on the boundary box of the defect in the detection result, and projects the detection result into the mixed reality view using the internal parameters of the mixed reality device system camera and the depth sensor data.

2. The method for real-time detection of appearance defects in a mixed reality environment according to claim 1, characterized in that: The step S1 starts the defect detection function of the mixed reality device and initializes the system camera of the mixed reality device, including: Step S11, starting the defect detection function of the mixed reality device and checking the first-person perspective function of the mixed reality head display; Step S12, initializing a system camera of the mixed reality device, wherein the system camera of the mixed reality device is a camera facing the external environment, and the mixed reality device can access and control the use and stop of the camera; Step S13, check whether the mixed reality device is started normally.

3. The method for real-time detection of appearance defects in a mixed reality environment according to claim 1, characterized in that: The step S2, establishing a network connection between the mixed reality device and the local server, includes: Use the functions in the System.Net and System.Net.Sockets libraries to establish a TCP connection between the mixed reality device and the local server.

4. The method for real-time detection of appearance defects in a mixed reality environment according to claim 3, characterized in that: The use of functions in the System.Net and System.Net.Sockets libraries to establish a TCP connection between the mixed reality device and the local server includes: Define the IP address and port number of the local server; Create a Socket instance so that the mixed reality device and the local server can communicate using the TCP protocol. The mixed reality device attempts to connect to the local server and enters the communication state after the connection is successfully established.

5. The method for real-time detection of appearance defects in a mixed reality environment according to claim 1, characterized in that: The step S3, controlling the camera of the mixed reality device system to collect and encode images, and transmitting the encoded images to the local server, includes: Click the camera button on the control panel of the mixed reality device to start the camera to capture the image in the current field of view. The captured image data is stored in texture format. Read the current field of view image data from the texture data buffer of the mixed reality device; Use the JPEG encoding library to encode texture format image data into JPEG format and generate a byte stream; Creating a broadcast packet containing meta information for the byte stream image data, the broadcast packet including data type, image size information, a timestamp and image data; Send the generated advertising packet to the API endpoint specified by the local server.

6. The method for real-time detection of appearance defects in a mixed reality environment according to claim 1, characterized in that: In step S4, after receiving the image, the local server decodes and inputs the image into the YOLOv8 model for defect detection, and returns the detection result to the mixed reality device through the network, including: The local server decodes the received encoded image data into an image matrix in a standard RGB format; The decoded image is input into the YOLOv8 model for defect detection. The YOLOv8 model analyzes each detection area in the image, identifies the existing appearance defects and generates the corresponding bounding box location, defect category and confidence level. All detection results are integrated into a detection result data packet, and the detection result data packet is transmitted back to the mixed reality device through a TCP connection.

7. The method for real-time detection of appearance defects in a mixed reality environment according to claim 6, characterized in that: The outputs of the YOLOv8 model include: The bounding box position of each detected defect area is expressed as ( , , , ) format, where ( , ) is the coordinate of the upper left corner of the defect area, ( , ) is the coordinate of the lower right corner of the defect area; Defect category information corresponding to each bounding box; Confidence: The confidence score indicates the reliability of the test result. A test result with a higher confidence score means that the model is more certain about the defect. The detection result data packet includes: Bounding Box Position:( , , , ); Defect categories: including but not limited to category labels for scratches and cracks; Confidence score: Indicates the confidence level.

8. The method for real-time detection of appearance defects in a mixed reality environment according to any one of claims 1 to 7, characterized in that: In step S5, the mixed reality device performs 2D to 3D coordinate conversion on the boundary box of the defect in the detection result, and projects the detection result into the mixed reality view using the camera intrinsic parameters and depth sensor data of the mixed reality device system, including: Extracting the 2D bounding box coordinates, category label, and confidence information of each defect from the received inspection result data packet; Using the camera intrinsic matrix K Convert the bounding box from 2D pixel coordinates to 3D coordinates; Use the camera's inverse projection formula to map the 2D bounding box coordinates to 3D space points in the camera's coordinate system; The depth sensor is used to obtain the depth data of each pixel, the depth information is matched with the 3D coordinates in the camera coordinate system, and the 3D position in the camera coordinate system is converted into the world coordinate system of the mixed reality device, so that the defect can be aligned with the real position of the inspected object in the actual space.

9. The method for real-time detection of appearance defects in a mixed reality environment according to claim 8, characterized in that: The camera intrinsic parameter matrix K for: , in, and The camera is Axis and The focal length along the axis, and are the principal point coordinates of the camera respectively; The inverse projection formula of the camera is: , in,( , ) is the 2D coordinate in the bounding box, point Z Represents the depth value, obtained from the depth sensor data.

10. A device for real-time detection of appearance defects in a mixed reality environment, characterized in that: include: An initialization module, used to start the defect detection function of the mixed reality device and initialize the system camera of the mixed reality device; A network connection module, used to establish a network connection between the mixed reality device and the local server; An image encoding module is used to control the camera of the mixed reality device system to collect and encode images, and transmit the encoded images to a local server; A defect detection module is used for decoding the image after the local server receives it and inputting the image into the YOLOv8 model for defect detection, and returning the detection result to the mixed reality device through the network; the YOLOv8 model outputs the bounding box, defect category and confidence of each defect; The detection result projection module is used for the mixed reality device to perform 2D to 3D coordinate conversion on the boundary box of the defect in the detection result, and to project the detection result into the mixed reality view using the internal parameters of the mixed reality device system camera and the depth sensor data.