Measurement Method and System Based on KAN and MR

By introducing KAN network and mixed reality technology into the measurement technology, combining depth sensors and RGB cameras, the shortcomings of existing measurement technologies in accuracy, real-time and environmental adaptability are solved, and high-precision real-time scale line detection is achieved.

CN120071096BActive Publication Date: 2025-07-01UNIV OF JINAN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510543271.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-01
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing measurement technology has shortcomings in accuracy, real-time and environmental adaptability, making it difficult to achieve high-precision real-time scale line detection under complex backgrounds and lighting changes.

Method used

Using KAN (Kolmogorov-Arnold Network) and mixed reality technology measurement methods, high-precision real-time scale line detection is achieved through depth sensors, RGB cameras and interactive display modules, combined with image preprocessing, feature extraction and KAN network training.

Benefits of technology

It realizes high-precision real-time scale line detection, overcomes the insufficient accuracy and real-time performance of traditional measurement technologies, adapts to complex backgrounds and lighting changes, and improves measurement efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071096B_ABST
    Figure CN120071096B_ABST
Patent Text Reader

Abstract

A measurement method and system based on KAN and MR according to the present invention belong to the technical field of computer vision and data recognition. The method includes the steps of: collecting the scale line image on the measuring tool and performing image preprocessing; extracting features from the preprocessed image to obtain the geometric feature vector of the image; establishing a KAN network and performing network training to obtain a trained KAN network. The input vector of the KAN network consists of the geometric features of the left endpoint coordinate, the right endpoint coordinate, and the midpoint coordinate of the red scale line. The KAN network includes 3 hidden layers, with 64 nodes in each layer, the activation function is an adaptive spline, a single-node output is used to predict the scale value, and the mean square error is used as the loss function; inputting the geometric feature vector into the trained KAN network to output the predicted scale value. The present invention adopts the KAN network and the mixed reality technology to realize the real-time and high-precision reading of the scale line, improving the measurement efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a measurement method and system based on KAN (Kolmogorov - Arnold network) and MR, specifically a real - time detection method and system for scale lines based on mixed reality technology and artificial intelligence algorithms, belonging to the technical field of computer vision and data recognition. Background Art

[0002] In the existing measurement field, the reading of scale lines mainly relies on traditional manual methods or scanning and recognition technologies based on optical instruments. Scale lines are marking lines used to measure or mark the length or time of an object, widely applied in various fields, including engineering, architecture, physics, mathematics, etc. The traditional manual visual reading method is vulnerable to the influence of viewing angle, light, and human factors, resulting in low accuracy, low efficiency, and even possible incorrect readings. While the existing scanning and recognition technologies based on optical instruments can improve accuracy, they usually require additional hardware devices and are complex to operate, restricting their application scope.

[0003] With the rapid development of mixed reality technology, although some mixed reality systems have attempted to achieve real - time measurement, due to the limitations of algorithm processing, it is often difficult to adapt to complex backgrounds and light changes, resulting in the inability to meet the requirements of accuracy and real - time of the reading results in some practical applications.

[0004] As a new neural network structure, the KAN network can more precisely simulate the complex mapping relationship between input and output by introducing a non - linear function to replace the linear weights in traditional neural networks. Compared with the traditional multi - layer perceptron (MLP), KAN has stronger capabilities in function approximation and can adapt to the complex non - linear relationship between image features and scale values. Therefore, the present invention proposes a measurement method and system based on KAN and MR. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a measurement method and system based on KAN and MR, which can achieve high - precision real - time scale line detection and reading.

[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0007] In a first aspect, a measurement method based on KAN and MR provided by an embodiment of the present invention includes the following steps:

[0008] S1, collect the scale line image on the measuring tool and perform image pre - processing;

[0009] S2, extract features from the pre - processed image to obtain the geometric feature vector of the image;

[0010] S3. Establish a KAN network and perform network training to obtain a trained KAN network. The input vector of the KAN network consists of geometric features of the left endpoint coordinates, right endpoint coordinates, and the midpoint coordinates of the red scale line. The KAN network includes 3 hidden layers, with 64 nodes in each layer, and the activation function is an adaptive spline. The KAN network uses a single-node output to output the predicted scale value. Each hidden layer is composed of spline functions, and the spline functions fit the non-linear relationship of the input features through piecewise polynomials: , where is the spline function, is the training coefficient, is the number of hidden layers; The KAN network uses the mean squared error (MSE) as the loss function: , where represents the loss value, is the total number of training samples, is the true scale value of the i-th sample, is the predicted scale value of the i-th sample by the model;

[0011] S4. Input the geometric feature vector into the trained KAN network to output the predicted scale value.

[0012] As a possible implementation of this embodiment, S1 includes the following steps:

[0013] S11. Initialize the depth sensor, RGB camera, and interactive display module;

[0014] S12. Capture the scale line image and perform preprocessing and quality assessment on the captured scale line image.

[0015] As a possible implementation of this embodiment, S11 includes the following steps:

[0016] S111. Start the RGB camera and display prompt information through the interactive display module to guide the user to operate;

[0017] S112. Use the initialized camera method to check whether the camera mode has been started and configure the resolution and its parameters of the camera as needed;

[0018] S113. After the camera initialization is completed, configure the camera parameters and start the photo-taking mode.

[0019] As a possible implementation of this embodiment, S12 includes the following steps:

[0020] S121. When the user triggers the recognition logic, capture the scale line image and crop the scale line image according to a preset area. The cropping area is accurately determined by calculating the ratio of the viewport coordinates to the image size, ensuring that the cropping area includes the complete scale line area, and the cropping area also includes offset adjustment;

[0021] S122. Perform brightness equalization processing on the cropped scale line image, calculate the brightness value of each pixel in the scale line image. If there are pixels in the scale line image whose brightness exceeds the preset threshold, it is determined that the scale line image has an overexposure problem;

[0022] S123. Count the proportion of pixels in the scale line image whose brightness exceeds the preset threshold. If the pixel proportion exceeds the set threshold, the image is determined to be an overexposed image, and the user is prompted that the image quality does not meet the requirements and is advised to adjust the angle and reshoot.

[0023] As a possible implementation of this embodiment, the cropping area is accurately determined by calculating the ratio of the viewport coordinates to the image size, including the following steps:

[0024] Convert the world coordinates to the camera coordinate system;

[0025] Use the projection matrix to convert the camera coordinates to the cropping coordinates;

[0026] Convert the cropping coordinates to the normalized device coordinates;

[0027] Convert the normalized device coordinates to the viewport coordinates, and calculate the width and height of the scale line image cropping area according to the viewport coordinates.

[0028] As a possible implementation of this embodiment, the S2 includes the following steps:

[0029] S21. Apply an edge detection algorithm to extract the contours of the preprocessed scale line image;

[0030] S22. Perform judgment processing on the extracted contours, and extract the coordinate points of the bottom left, bottom right, top left, and top right that include the complete scale line information;

[0031] S23. Perform perspective transformation on the extracted coordinate points to obtain the perspective-transformed image;

[0032] S24. Perform skeletonization processing on the perspective-transformed image, and extract the coordinate values of the left and right ends of the complete scale line and the accurate coordinate values of the red scale line.

[0033] As a possible implementation of this embodiment, the S21 includes the following steps:

[0034] S211. Convert the original scale line image from the RGBA color space to the RGB color space, and then convert it to the HSV color space;

[0035] S212. Define the HSV range of the black area in the HSV color space, extract the black area in the scale line image, and generate a mask that only retains the black area part;

[0036] S213. Perform morphological operations on the extracted black area mask;

[0037] S214. Extract the contours in the morphologically processed mask and draw the extracted contours on the original scale line image; if no contours are detected, output an error message.

[0038] As a possible implementation manner of this embodiment, the S22 includes the following steps:

[0039] S221. Initialize four extreme points and save the coordinates of the top-leftmost, top-rightmost, bottom-leftmost, and bottom-rightmost points found in the contour;

[0040] S222. Traverse each contour of the scale line image and determine the extreme points according to the positions of the contour points, ignoring the pixel points close to the image edge; if no valid extreme points are found, return empty and display an error message;

[0041] S223. Check whether the four extracted extreme points can form a rectangle. If a rectangle cannot be formed, return empty and prompt that the detection fails;

[0042] S224. If the extreme points are valid and form a rectangle, perform perspective transformation to map the extracted extreme points to the area with the target width and height;

[0043] S227. If steps 221 - 226 are all successfully completed, return the updated extreme point dictionary.

[0044] S226. Calculate the direction vector and update the coordinates of the top-leftmost extreme point and the top-rightmost extreme point based on the direction vector. If the updated coordinates exceed the image boundary, return empty and prompt the user;

[0045] S227. If steps 221 - 226 are all successfully completed, return the updated extreme point dictionary.

[0046] As a possible implementation manner of this embodiment, the S23 includes the following steps:

[0047] S231. Perform perspective transformation on the four detected valid extreme points;

[0048]

[0048] Take four extreme points and the target image size as inputs, calculate a 3x3 perspective transformation matrix, and convert the points of the source image into points in the target image according to the perspective transformation matrix;

[0049]

[0049] Use the perspective transformation matrix to map the original image to the size of the target image, and transform the four extreme points into the target size and perspective, and output the image after perspective transformation processing.

[0050] As a possible implementation manner of this embodiment, the S24 includes the following steps:

[0051]

[0051] Convert the transformed image into a grayscale image, and perform an inverse binary operation to obtain a binary image, where a binary threshold is set, and the pixel values greater than the threshold are changed to 255, and the others are changed to 0;

[0052]

[0052] Use a vertical structuring element to perform multiple dilation operations on the binary image to enhance the vertical scale line area and generate a dilated image;

[0053]

[0053] Perform contour detection on the dilated image to obtain all contours, and filter out the vertical contours whose height is greater than twice the width as valid contours;

[0054]

[0054] Perform skeletonization processing on the dilated image to obtain a skeletonized image, extract the longest vertical line of the red scale line from the skeletonized image, calculate the midpoint of the red scale line, and mark the longest line segment in the image;

[0055]

[0055] Draw valid contours on a new blank image, and perform skeletonization processing on the image with valid contours drawn to obtain a skeletonized image;

[0056]

[0056] Traverse each pixel point in the skeletonized image, check the skeleton pixels, and update the positions of the leftmost skeleton point and the rightmost skeleton point;

[0057]

[0057] Return the left and right end coordinate values of the complete scale line and the accurate coordinate values of the red scale line.

[0058] In a second aspect, a measurement system based on KAN and MR provided by an embodiment of the present invention includes:

[0059] An image acquisition module, configured to acquire the scale line image on the measurement tool and perform image preprocessing;

[0060] A feature extraction module, configured to extract features from the preprocessed image to obtain the geometric feature vector of the image;

[0061] A network establishment module for establishing a KAN network, performing network training, and obtaining a trained KAN network. The input vector of the KAN network consists of geometric features of the left endpoint coordinate, right endpoint coordinate, and the midpoint coordinate of the red scale line. The KAN network includes 3 hidden layers, with 64 nodes in each layer, and the activation function is an adaptive spline. The KAN network uses a single-node output to output the predicted scale value. Each hidden layer is composed of a spline function, and the spline function fits the non-linear relationship of the input features through piecewise polynomials: , where is the spline function, is the training coefficient, is the number of hidden layers; The KAN network uses the mean squared error (MSE) as the loss function: , where represents the loss value, is the total number of training samples, is the true scale value of the i-th sample, is the predicted scale value of the i-th sample by the model;

[0062] A KAN network module for inputting the geometric feature vector into the trained KAN network and outputting the predicted scale value.

[0063] The beneficial effects of the technical solution of the embodiment of the present invention are as follows:

[0064] The present invention utilizes the computer vision, depth perception, and spatial positioning capabilities of the HoloLens 2 device, combines the KAN network with image processing algorithms, and uses the KAN network and mixed reality technology to achieve high-precision real-time scale line detection and reading, overcoming the deficiencies of traditional manual measurement and optical scanners in terms of accuracy, real-time performance, environmental adaptability, etc., and ensuring efficient and accurate reading in a dynamic environment.

[0065] The present invention realizes real-time and high-precision reading of scale lines by using the KAN network and mixed reality technology, improving the measurement efficiency and accuracy; The KAN network design of the present invention enables the model to adaptively process the non-linear relationship of input features, improving the generalization ability and adaptability of the model; The KAN network uses a spline function to replace the linear layer of the traditional neural network, achieving high-precision non-linear mapping, breaking through the linear bottleneck, and improving the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 is a flowchart of a measurement method based on KAN and MR shown according to an exemplary embodiment;

[0067] Figure 2 is a schematic structural diagram of a measurement system based on KAN and MR shown according to an exemplary embodiment. Detailed implementation manners

[0068] To more clearly illustrate the technical features of the solution of the present invention, the present invention will be elaborated in detail below through specific implementation manners and in combination with its accompanying drawings.

[0069] As shown in Figure 1 A measurement method based on KAN and MR provided by an embodiment of the present invention includes the following steps:

[0070] S1. Collect the scale line images on the measurement tool and perform image preprocessing;

[0071] S2. Extract features from the preprocessed image to obtain the geometric feature vector of the image;

[0072] S3. Establish a KAN network and perform network training to obtain a trained KAN network. The input vector of the KAN network consists of the geometric features of the left endpoint coordinates, right endpoint coordinates, and the midpoint coordinates of the red scale line. The KAN network includes 3 hidden layers, with 64 nodes in each layer, and the activation function is an adaptive spline; the KAN network uses a single-node output to output the predicted scale value; each hidden layer is composed of a spline function, and the spline function fits the nonlinear relationship of the input features through piecewise polynomials: , where is the spline function, is the training coefficient, is the number of hidden layers; the KAN network uses the mean square error (MSE) as the loss function: , where represents the loss value, is the total number of training samples, is the true scale value of the i-th sample, is the predicted scale value of the i-th sample by the model;

[0073] S4. Input the geometric feature vector into the trained KAN network and output the predicted scale value.

[0074] As a possible implementation manner of this embodiment, the S1 includes the following steps:

[0075] S11. Initialize the depth sensor, RGB camera, and interactive display module;

[0076] S12. Capture the scale line image and perform preprocessing and quality evaluation on the captured scale line image.

[0077] As a possible implementation manner of this embodiment, the S11 includes the following steps:

[0078] S111, Start the RGB camera and display a prompt message through the interactive display module to guide the user to operate;

[0079] S112, Use the initialized camera method to check whether the camera mode has been started, and configure the resolution and its parameters of the camera as needed;

[0080] S113, After the camera initialization is completed, configure the camera parameters and start the photo-taking mode.

[0081] As a possible implementation of this embodiment, the S12 includes the following steps:

[0082] S121, When the user triggers the recognition logic, capture the scale line image and crop the scale line image according to the preset area. The cropping area is accurately determined by calculating the ratio of the viewport coordinates to the image size, ensuring that the cropping area includes the complete scale line area, and the cropping area also includes offset adjustment;

[0083] S122, Perform brightness equalization processing on the cropped scale line image, calculate the brightness value of each pixel in the scale line image. If there are pixels in the scale line image whose brightness exceeds the preset threshold, it is determined that the scale line image has an overexposure problem;

[0084] S123, Count the pixel ratio of the scale line image whose brightness exceeds the preset threshold. If the pixel ratio exceeds the set threshold, it is determined that the image is an overexposed image, and the user is prompted that the image quality does not meet the requirements, and the user is advised to adjust the angle and take a new photo.

[0085] As a possible implementation of this embodiment, the cropping area is accurately determined by calculating the ratio of the viewport coordinates to the image size, including the following steps:

[0086] Convert the world coordinates to the camera coordinate system;

[0087] Use the projection matrix to convert the camera coordinates to the cropping coordinates;

[0088] Convert the cropping coordinates to the normalized device coordinates;

[0089] Convert the normalized device coordinates to the viewport coordinates, and calculate the width and height of the scale line image cropping area according to the viewport coordinates.

[0090] As a possible implementation of this embodiment, the S2 includes the following steps:

[0091] S21, Apply the edge detection algorithm to extract the contours of the preprocessed scale line image;

[0092] S22. Perform judgment processing on the extracted contour, and extract the coordinate points of the bottom-left, bottom-right, top-left, and top-right that include complete scale line information;

[0093] S23. Perform perspective transformation on the extracted coordinate points to obtain the image after perspective transformation;

[0094] S24. Perform skeletonization processing on the image after perspective transformation, and extract the coordinate values of the left and right ends of the complete scale line and the accurate coordinate values of the red scale line.

[0095] As a possible implementation manner of this embodiment, the S21 includes the following steps:

[0096] S211. Convert the original scale line image from the RGBA color space to the RGB color space, and then convert it to the HSV color space;

[0097] S212. Define the HSV range of the black area in the HSV color space, and extract the black area in the scale line image to generate a mask that only retains the black area part;

[0098] S213. Perform morphological operations on the extracted black area mask;

[0099] S214. Extract the contour in the mask after morphological processing, and draw the extracted contour on the original scale line image; if no contour is detected, output an error message.

[0100] As a possible implementation manner of this embodiment, the S22 includes the following steps:

[0101] S221. Initialize four extreme points, and save the coordinate points of the top-left, top-right, bottom-left, and bottom-right found in the contour;

[0102] S222. Traverse each contour of the scale line image, and judge the extreme points according to the positions of the contour points, ignoring the pixel points close to the image edge; if no valid extreme points are found, return empty and display an error message;

[0103] S223. Check whether the four extracted extreme points can form a rectangle. If they cannot form a rectangle, return empty and prompt that the detection fails;

[0104] S224. If the extreme points are valid and form a rectangle, perform perspective transformation to map the extracted extreme points to the area of the target width and height;

[0105] S225. Perform preliminary detection of the red scale line. If no valid red scale line is detected, return an error message and re-detect until a valid red scale line is detected;

[0106] S226. Calculate the direction vector and update the coordinates of the top - left extreme point and the top - right extreme point based on the direction vector. If the updated coordinates exceed the image boundary, return null and prompt the user.

[0107] S227. If steps 221 - 226 are all successfully completed, return the updated extreme point dictionary.

[0108] As a possible implementation of this embodiment, the S23 includes the following steps:

[0109] S231. Perform a perspective transformation on the four detected valid extreme points.

[0110] S232. Take the four extreme points and the target image size as inputs, calculate a 3x3 perspective transformation matrix, and convert the points of the source image to points in the target image according to the perspective transformation matrix.

[0111] S233. Use the perspective transformation matrix to map the original image to the target image size, and transform the four extreme points to the target size and perspective, and output the image after perspective transformation processing.

[0112] As a possible implementation of this embodiment, the S24 includes the following steps:

[0113] S241. Convert the transformed image to a grayscale image and perform an inverse binary operation to obtain a binary image. Set the binary threshold, change the pixel values greater than the threshold to 255, and the others to 0.

[0114] S242. Use a vertical structuring element to perform multiple dilation operations on the binary image to enhance the vertical scale line area and generate a dilated image.

[0115] S243. Perform contour detection on the dilated image to obtain all contours, and filter out the vertical contours whose height is greater than twice the width as valid contours.

[0116] S244. Perform skeletonization on the dilated image to obtain a skeletonized image. Extract the longest vertical line of the red scale line from the skeletonized image, calculate the mid - point of the red scale line, and mark the longest line segment in the image.

[0117] S245. Draw the valid contours on a new blank image and perform skeletonization on the image with the valid contours drawn to obtain a skeletonized image.

[0118] S246. Traverse each pixel point in the skeletonized image, check the skeleton pixels, and update the positions of the left - most skeleton point and the right - most skeleton point.

[0119] S247 returns the left and right end coordinate values of the complete scale line and the accurate coordinate values of the red scale line.

[0120] As Figure 2 shown, a measurement system based on KAN and MR provided by an embodiment of the present invention includes:

[0121] An image acquisition module for acquiring the scale line image on the measurement tool and performing image preprocessing;

[0122] A feature extraction module for extracting features from the preprocessed image to obtain the geometric feature vector of the image;

[0123] A network establishment module for establishing a KAN network and performing network training to obtain a trained KAN network. The input vector of the KAN network consists of geometric features of the left end point coordinate, the right end point coordinate, and the midpoint coordinate of the red scale line. The KAN network includes 3 hidden layers, with 64 nodes in each layer, and the activation function is an adaptive spline; the KAN network uses a single-node output to output the predicted scale value; each hidden layer is composed of a spline function, and the spline function fits the nonlinear relationship of the input features through piecewise polynomials: , where is the spline function, is the training coefficient, is the number of hidden layers; the KAN network uses the mean square error (MSE) as the loss function: , where represents the loss value, is the total number of training samples, is the true scale value of the i-th sample, is the predicted scale value of the model for the i-th sample;

[0124] A KAN network module for inputting the geometric feature vector into the trained KAN network and outputting the predicted scale value.

[0125] The specific process of real-time scale line detection in the present invention is as follows.

[0126] 1. Start the scale line detection function and initialize the depth sensor, RGB camera, and interactive display module.

[0127] The depth sensor is responsible for obtaining the spatial depth information of the scale line to ensure the three-dimensional positioning accuracy of the image; the camera is responsible for capturing the scale line image in real time; the interactive display module ensures the real-time presentation of feedback information in the HoloLens 2 device.

[0128] Specifically:

[0129] (1)Initialize the functions of the camera, depth sensor, and interactive display module, and start the photo-taking process. During this process, the system first starts the RGB camera and guides the user to operate by displaying prompt messages. The code is as follows:

[0130] / / Initialize the camera InitializeCamera();

[0131] / / Prompt the user to operate hitText.text = "Please place the complete scale line in the detection frame and click the start detection button. The photo will be taken and read after five seconds!"

[0132] (2)Initialize the camera (InitializeCamera() method) This method is used to check whether the camera mode has been started and configure the resolution and other parameters of the camera as needed. Use PhotoCapture to start the camera mode and prepare for image capture.

[0133] (3)After the camera initialization is completed (OnPhotoCaptureCreated() method) When the camera is ready, we configure the camera parameters through the OnPhotoCaptureCreated method and start the photo-taking mode. The camera resolution and other settings affect the image quality, so this step ensures the maximum resolution to obtain a clear scale line image.

[0134] II. When the user triggers the recognition logic, capture the image, crop the captured image, and then evaluate the quality of the cropped image. Use the image brightness equalization algorithm to automatically determine whether there is an overexposure problem in the image.

[0135] Preprocess the captured image, including denoising, image cropping, etc. The cropping area focuses on the area containing the scale line, removing unnecessary background information and improving the efficiency of subsequent processing. If overexposure is detected, the system automatically compensates through the image enhancement algorithm to restore image details and improve the accuracy of subsequent processing.

[0136] Specifically:

[0137] (1)When the user triggers the recognition logic, the system first captures the image and crops it according to the preset area. The cropping area is accurately determined by calculating the ratio of the viewport coordinates to the image size, ensuring that the cropping area includes the complete scale line area. The cropping area also includes offset adjustment to adapt to the actual needs in different scenarios.

[0138] (2) Perform brightness equalization processing on the cropped image and calculate the brightness value of each pixel in the image. If there are pixels in the image with brightness exceeding the preset threshold, it is determined that the image may have an overexposure problem. By counting the proportion of pixels with higher brightness (exceeding the preset brightness threshold) in the image, if it exceeds the set threshold, the image is determined to be an overexposed image.

[0139] (3) According to the brightness equalization algorithm, the system will evaluate the proportion of overexposed pixels in the total pixels and determine whether the image has an overexposure phenomenon based on this. If an overexposure area exceeding the set proportion is detected, the system will prompt the user that the image quality does not meet the requirements and recommend taking a new photo or adjusting the camera settings.

[0140] Among them, the cropping area is an adaptation rectangular frame suitable for the size of the complete scale line pre-designed in the mixed reality environment. The specific cropping logic is as follows:

[0141] Coordinate transformation: By converting the world coordinates to the viewport coordinates, ensure that the cropping area can adapt to the actual position on the display screen. First, the world coordinates need to be converted to the camera coordinate system. Assume the camera position is , the world coordinate point is , and the transformation matrix of the camera is , then the coordinates in the camera coordinate system are calculated by the following formula:

[0142] ;

[0143] The projection matrix converts the camera coordinates to the cropping coordinates. In perspective projection, the projection matrix is usually expressed as , then the coordinates in the cropping space are:

[0144] ;

[0145] The coordinates in the cropping space need to be converted to the normalized device coordinates (NDC), and the range of NDC is [−1,1]. By dividing by the component of the homogeneous coordinates, the normalized device coordinates are obtained:

[0146] ;

[0147] The normalized device coordinates need to be further converted to the viewport coordinates, and the range of the viewport coordinates is usually [0,1] (in Unity it is [0,1]). The viewport coordinates are calculated by the formula:

[0148] ,

[0149] This formula maps NDC coordinates in the range [-1, 1] to viewport coordinates in the range [0, 1]. Considering that the direction of the y-axis in Unity's coordinate system is opposite, the y-axis needs to be inverted.

[0150] Calculation of the clipping area ratio: Calculate the width and height of the clipping area based on the viewport coordinates.

[0151] Pixel coordinate calculation: Use the viewport coordinates and the total size (width and height) of the image to determine the pixel coordinates of the actual clipping area.

[0152] Clipping operation: Finally, use the calculated pixel coordinates to clip the image using an image clipping function. This part of the implementation is based on slicing operations of the image matrix.

[0153] Among them, the equalization processing and exposure judgment mainly include:

[0154] Gray conversion: First, convert the image to a grayscale image for subsequent exposure analysis;

[0155] Exposure detection: Extract pixels with too high brightness through threshold processing to create an exposure mask;

[0156] Calculate the exposure ratio: Calculate the proportion of overexposed pixels in the total pixels and determine whether it exceeds the preset ratio;

[0157] Judge the exposure problem: Judge whether the image has a serious exposure problem according to the overexposure ratio. If so, prompt the user to adjust the angle and reshoot.

[0158] Third, apply an edge detection algorithm to the preprocessed image for contour extraction.

[0159] The system identifies the contours of the scale lines and uses morphological processing for optimization to ensure that the edges of the scale lines are smooth and continuous.

[0160] Specifically:

[0161] (1) First, convert the original image (imageMat) from the RGBA color space to RGB, and then convert it to the HSV color space. This conversion helps to more accurately identify specific color regions, especially for effective color screening in terms of hue (H), saturation (S), and value (V).

[0162] (2) Use the hue, saturation, and value values in the HSV color space to define the HSV range of the black area. By calling the Core.inRange() function, extract the black area in the image and generate a mask that only retains the black area part.

[0163] (3) Perform morphological operations on the extracted black area mask. First, use the closing operation (MORPH_CLOSE) to fill small holes in the image, and then perform the dilation operation to further enhance the black area in the image for better contour detection.

[0164] (4) Use the Imgproc.findContours() function to extract contours from the morphologically processed mask, and store the contours in contourss. Subsequently, use Imgproc.drawContours() to draw the extracted contours on the original image (contoursMat). Each contour is displayed with a green line.

[0165] (5) If no contours are detected (contourss.Count == 0), output an error message to prompt the user that no contours are found and subsequent detection cannot continue.

[0166] IV. Judge and process the detected contours, and extract the coordinate points of the bottom - left, bottom - right, top - left, and top - right that include the complete scale line information for accurately positioning the scale line area in the image, and further providing a basis for subsequent perspective transformation and data processing.

[0167] Based on the detected contour information, the system further judges and filters out the areas containing complete scale lines. Extract the key coordinate points, including the bottom - left, bottom - right, top - left, and top - right corner points, to accurately calibrate the boundaries of the scale lines. By further analyzing the shape and position of the contours, ensure that the extracted coordinate points can truly reflect the actual scale line positions.

[0168] Specifically:

[0169] (1) Initialize four extreme points: topLeft, topRight, bottomLeft, and bottomRight to save the coordinate points of the top - left, top - right, bottom - left, and bottom - right found in the contour.

[0170] (2) Traverse each contour (contours) and judge the extreme points according to the positions of the contour points (point.x and point.y). To avoid noise points, ignore the pixel points near the image edge (set by the margin parameter).

[0171] (3) If no valid extreme points are found (i.e., the coordinates of the extreme points are not updated), return empty and display an error message indicating that the detection fails. At this time, ask the user to adjust the image position and re - detect.

[0172] (4) After extracting four extreme points, check whether these points can form a rectangle. If it is not a rectangle, return empty and prompt that the detection fails.

[0173] (5) Assume that the extreme points are valid and form a rectangle, then perform perspective transformation to map the extracted extreme points to the area of the target width and height, based on the normalized image.

[0174] (6) After performing perspective transformation, perform the preliminary detection of the red scale line. If no valid red scale line is detected, the system will prompt the user that the rebound value is not enabled or lower than the valid value, and require re-detection.

[0175] (7) Calculate the direction vectors and update the coordinates of topLeft and topRight based on these direction vectors, ensuring that the updated points are still within the image boundary. If the updated coordinates exceed the image boundary, return empty and prompt the user.

[0176] (8) If all steps are successfully completed, return the updated dictionary of extreme points, including the positions of four points: TopLeft, TopRight, BottomLeft, and BottomRight.

[0177] Among them, for the preliminary detection of the red scale line, the specific steps are as follows:

[0178] First, check whether the input image is empty. If it is empty, output an error message and return false.

[0179] Convert the input image to a grayscale image. After conversion, the color information of the image is removed, and only the brightness information is retained, which is convenient for subsequent processing.

[0180] Perform binarization on the grayscale image. Set the part where the pixel value is greater than the set threshold to 255, representing the foreground part (white) in the image; the part less than the threshold is set to 0, representing the background (black). The purpose of binarization is to convert the image into a black-and-white image to highlight the important areas.

[0181] Perform traditional skeletonization on the binarized image to further simplify the image and highlight the structure of the lines. This process converts the lines in the image into thinner line forms, which is convenient for subsequent line analysis.

[0182] Traverse each column of the image to detect whether there are continuous white pixels (representing the red scale line) in the vertical direction. By counting the number of continuous pixels, if the number of continuous white pixels in a certain column is greater than the set minimum length (30 pixels), it is considered that a valid vertical line is detected.

[0183] If a valid vertical line is detected, return true, indicating successful detection. If no vertical line meeting the criteria is found, return false. Any exception occurring during the process will also return false and output an error message.

[0184] V. Perform perspective transformation on the extracted coordinate points.

[0185] Use a perspective transformation algorithm (such as homography transformation) to geometrically correct the extracted coordinate points. This step ensures that regardless of the image acquisition angle, the scale lines can be presented in a standard form parallel to the image, making the readings more accurate. After perspective transformation, the scale lines in the image will be restored to their positions in the real world, eliminating the errors caused by the viewing angle.

[0186] Specifically:

[0187] (1) After detecting valid extreme points (such as four coordinate points: TopLeft, TopRight, BottomLeft, and BottomRight), the method first calls the PerformPerspectiveTransform() function to perform perspective transformation. If the perspective transformation fails, an error message will be displayed and subsequent operations will be terminated.

[0188] (2) The PerformPerspectiveTransform() function accepts four extreme points and the target image size as inputs, and the detailed steps of performing perspective transformation are as follows:

[0189] The core of perspective transformation is to describe the transformation relationship between four points through a 3x3 matrix.

[0190] For perspective transformation, we have a point on the source image and a point on the target image , which are transformed through the perspective matrix. The perspective transformation matrix is a 3x3 matrix, usually represented as:

[0191]

[0192] This matrix will transform the point on the source image into the point in the target image through the following transformation formula:

[0193]

[0194] Here, is the homogeneous coordinate, used to convert the final result from homogeneous coordinate back to ordinary coordinate. The transformed target coordinates are:

[0195]

[0196] Four points on the source image (topLeft, topRight, bottomRight, bottomLeft) and four points on the target image (the four corner points of the target image) are defined. The perspective transformation matrix is calculated based on these points.

[0197] The perspective matrix is obtained by solving the correspondence between the source points and the target points. Specifically, the elements of the perspective transformation matrix are obtained by solving the following system of equations:

[0198]

[0199] (3)Using the calculated perspective transformation matrix, the original image is mapped to the size of the target image through the warpPerspective() function. This function adjusts the perspective and size of the image according to the transformation matrix and outputs the transformed image.

[0200] (4)Return the image processed by perspective transformation. This image is transformed into the target size and perspective according to the four extreme points extracted.

[0201] VI. After the perspective transformation, the image is skeletonized to extract the complete scale line size and the accurate position of the red scale line.

[0202] Apply the skeletonization algorithm to the image after perspective transformation to extract the skeleton structure of the scale line and further optimize the accuracy of the scale line. By refining the skeleton structure, the system can accurately locate the red scale line and its position, ensuring the recognition accuracy in a complex background. By detecting the shortest path of the skeleton, the system can extract the accurate size of the scale line and its position in the image.

[0203] Specifically:

[0204] (1)Convert the transformed image to a grayscale image grayImage. Perform reverse binaryzation operation to convert the image into a binary image binaryImage. Set a certain binaryzation threshold, and all pixel values greater than this threshold become white (255), and the others become black (0).

[0205] (2)Perform dilation operation multiple times using a vertical structuring element to enhance the vertical scale line area and generate dilatedImage.

[0206] (3)Perform contour detection on the dilated image to obtain all contours. If contours are detected, by traversing all contours and calculating the minimum bounding rectangle of each contour, filter out the vertical contours whose height is more than twice the width.

[0207] (4) Skeletonize the expanded image to obtain skeletonred. Extract the longest vertical line of the red scale line from the skeletonized image. Calculate the midpoint midRedPoint of the red scale line and mark the longest line segment in the image.

[0208] (5) Draw the effective contour on a new blank image mask. Skeletonize the mask image with the effective contour drawn again to obtain skeleton.

[0209] (6) Traverse each pixel point in the skeletonized image skeleton to check which are the skeleton pixels (value is 255). Update leftMostPoint and rightMostPoint to the positions of the leftmost and rightmost skeleton points. If no valid leftmost or rightmost point is found, return the coordinate (0, 0), indicating that the scale line is not detected.

[0210] (7) Return the left endpoint coordinate ( x L , y L ), the right endpoint coordinate ( x R , y R ), and the midpoint coordinate of the red scale line ( x R ′ , y R ′ ).

[0211] VII. Predict the specific scale value of the red scale line in the full scale line size through the KAN network.

[0212] Predict the actual value of the scale line through the trained KAN network, and accurately predict the current scale value. Provide the accurate scale value to the user and display it in real time on the screen of HoloLens 2.

[0213] Specifically:

[0214] (1) First, train the KAN network model. Manually construct data or use a mathematical model to generate simulated data with noise. These data should simulate the actual application scenario. The left endpoint coordinate ( x L , y L ) and the right endpoint coordinate ( x R , y R) Generate randomly within a reasonable range. Ensure that it conforms to a reasonable value in the actual measurement environment and set a judgment mechanism to ensure that the left coordinate of the left endpoint is to the left of the left coordinate of the right endpoint.

[0215] KAN Network Design:

[0216] The KAN network is a neural network similar to a multi-layer perceptron (MLP), where the connections in each layer use non-linear transformations (spline functions) instead of traditional linear weights.

[0217] Network Structure:

[0218] Input Layer: Input feature vector X, which includes coordinates and other numerical features;

[0219] Hidden Layers: Multiple hidden layers that use spline functions for non-linear transformations to learn the complex relationship between input features and target values;

[0220] Output Layer: Output a numerical value, which is the predicted scale value .

[0221] Scale Value T It can be generated through coordinate differences, proportional relationships, or other formulas. In this example, the scale value is generated based on the relative positions between the left endpoint and the right endpoint:

[0222]

[0223] Here, max_width is the maximum width in the entire dataset, ensuring that the generated scale value is between 10 and 100.

[0224] To simulate measurement errors in the actual scenario, a certain amount of random noise can be added to the generated data. This helps improve the robustness of the model.

[0225] Data Preparation: Separate the feature vector X and the target value of each sample T as training data.

[0226] Loss Function: Use the mean squared error (MSE) as the loss function to measure the difference between the predicted value and the actual value T :

[0227] .

[0228] Optimization Method: Use the gradient descent method or an optimization algorithm (such as Adam) to train the KAN network. Adjust the parameters in the network by minimizing the loss function.

[0229] Training Iterations: Through multiple iterations, the network learns the non-linear relationship between input features and scale values.

[0230] (2) Check whether the X coordinate of the left endpoint is less than that of the right endpoint by if( x L >= x R ). If the condition is not met, throw ArgumentException to ensure that the left point is always on the left side of the right point and prevent errors in subsequent calculations.

[0231] (3) Prepare the above features into input data and construct a feature vector. The following features need to be extracted for each image: the left coordinate of the left endpoint ( x L , y L ), the left coordinate of the right endpoint ( x R , y R ), and the left coordinate of the midpoint of the red scale line ( x R ′ , y R ′ ).

[0232] (4) The feature vector input to the KAN network should contain the above features, and a vector X is obtained after merging.

[0233]

[0234] (5) Use Math.Clamp( T , 10, 100) to ensure that the scale value does not exceed the expected range (10 to 100). Even if the calculation result exceeds the range, it will be forced to be limited within this interval.

[0235] During the calculation process, if T is NaN or Infinity, check for abnormal situations through double.IsNaN( T ) and double.IsInfinity( T ), output an error log Debug.LogError("Scale value calculation failed."), and update the UI to prompt the user "Scale value calculation failed".

[0236] The finally calculated T will be displayed on the interface through hitText.text to provide the user with the actual scale value.

[0237] The present invention captures and processes images by activating a depth sensor, an RGB camera, and an interactive display module. Through steps such as image cropping, brightness equalization, edge detection, and perspective transformation, the system extracts the key coordinate points of the scale lines and applies a skeletonization algorithm to optimize the recognition accuracy. Finally, through the KAN network, the specific scale value of the red scale line is predicted, the current scale value is accurately obtained, and it is displayed in real time on the mixed reality device.

[0238] Compared with the prior art, the present invention has the following technical advantages:

[0239] 1. Improve the accuracy of scale line detection:

[0240] By combining a depth sensor and an RGB camera, precise positioning in three-dimensional space is achieved, and perspective transformation is used to eliminate the influence of the shooting angle on the recognition of scale lines, ensuring that scale lines can be accurately detected from different perspectives.

[0241] 2. Optimize image quality and enhance adaptability:

[0242] An image brightness equalization algorithm is adopted to automatically detect overexposure problems, ensuring clear scale line image feedback under lighting conditions and improving the stability of detection.

[0243] 3. Precisely extract scale line information:

[0244] Through edge detection and morphological processing, the extraction of the scale line contour is optimized, and combined with the skeletonization algorithm, the precise positioning of the complete scale line and the red scale line is achieved, avoiding false detection or missed detection caused by noise or discontinuous edges.

[0245] 4. Eliminate perspective distortion and improve measurement accuracy:

[0246] A homography transformation is used to perform perspective correction on the scale lines, making them present a standard parallel structure in the image, ensuring accurate readings without relying on a fixed-angle camera placement.

[0247] 5. Intelligently predict scale values, provide real-time feedback and optimization:

[0248] By extracting key point coordinates and other geometric information from the image, preparing feature vectors, inputting new image data, extracting features and passing them to the trained KAN network, the scale value is predicted in real time and displayed in real time on the mixed reality device, providing an intuitive and efficient measurement experience.

[0249] 6. Suitable for mixed reality environments and improve the interaction experience:

[0250] Combined with the interactive display function of the mixed reality device, users can view the detection results in real time, improve the measurement efficiency, and enhance the convenience and practicality in engineering applications.

[0251] The present invention can be widely applied to fields such as intelligent measurement, automated detection, and augmented reality (AR), and is particularly suitable for scenarios that require high-precision detection and real-time feedback.

[0252] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A measurement method based on KAN and MR, characterized in that: The steps include: S1, collecting the scale line image on the measuring tool and performing image preprocessing; S2, extracting features from the preprocessed image to obtain the geometric feature vector of the image; S3, establish a KAN network and perform network training to obtain a trained KAN network. The input vector of the KAN network consists of geometric features of the left endpoint coordinates, the right endpoint coordinates and the midpoint coordinates of the red scale line. The KAN network includes 3 hidden layers, each with 64 nodes, and the activation function is an adaptive spline. The KAN network uses a single node output and outputs a predicted scale value. Each hidden layer is composed of a spline function, and the spline function fits the nonlinear relationship of the input features through a piecewise polynomial: ,in, is the spline function, is the training coefficient, is the number of hidden layers; the KAN network uses mean square error as the loss function: ,in, represents the loss value, is the total number of training samples, is the true scale value of the i-th sample, is the predicted scale value of the model for the i-th sample; S4, input the geometric feature vector into the trained KAN network and output the predicted scale value.

2. The measurement method based on KAN and MR according to claim 1, characterized in that: The S1 comprises the following steps: S11, initializing the depth sensor, RGB camera and interactive display module; S12, capturing a scale line image and performing preprocessing and quality assessment on the captured scale line image.

3. The measurement method based on KAN and MR according to claim 2, characterized in that: The S11 comprises the following steps: S111, starting the RGB camera and displaying prompt information through the interactive display module to guide user operation; S112, using the camera initialization method to check whether the camera mode is started, and configuring the camera resolution and its parameters as needed; S113, after the camera is initialized, configure the camera parameters and start the photo taking mode.

4. The measurement method based on KAN and MR according to claim 2, characterized in that: The S12 comprises the following steps: S121, when the user triggers the recognition logic, the scale line image is captured and the scale line image is cropped according to a preset area, the cropping area is accurately determined by calculating the ratio of the viewport coordinates to the image size, ensuring that the cropping area includes a complete scale line area, and the cropping area also includes an offset adjustment; S122, performing brightness equalization processing on the cropped scale line image, calculating the brightness value of each pixel in the scale line image, and if there are pixels in the scale line image whose brightness exceeds a preset threshold, it is determined that the scale line image has an overexposure problem; S123, counting the proportion of pixels in the scale line image whose brightness exceeds a preset threshold. If the pixel proportion exceeds the set threshold, the image is determined to be an overexposed image, and the user is prompted that the image quality does not meet the requirements, and the user is advised to adjust the angle and retake the photo.

5. The measurement method based on KAN and MR according to any one of claims 1 to 4, characterized in that: The S2 comprises the following steps: S21, applying an edge detection algorithm to extract contours of the preprocessed scale line image; S22, judging and processing the extracted contour, extracting the coordinate points of the lower left, lower right, upper left and upper right that include complete scale line information; S23, performing perspective transformation on the extracted coordinate points to obtain an image after perspective transformation; S24, skeletonizing the image after the perspective transformation, and extracting the coordinate values ​​of the left and right ends of the complete scale line and the accurate coordinate value of the red scale line.

6. The measurement method based on KAN and MR according to claim 5, characterized in that: The S21 comprises the following steps: S211, converting the original scale line image from RGBA color space to RGB color space, and then converting it to HSV color space; S212, defining the HSV range of the black area in the HSV color space, extracting the black area in the scale line image, and generating a mask that only retains the black area portion; S213, performing morphological operations on the extracted black area mask; S214, extracting contours from the morphologically processed mask and drawing the extracted contours onto the original scale line image; if no contours are detected, outputting an error message.

7. The measurement method based on KAN and MR according to claim 5, characterized in that: The S22 comprises the following steps: S221, initialize four extreme points, and save the coordinates of the upper left, upper right, lower left, and lower right points found in the contour; S222, traverse each contour of the scale line image, and determine the extreme point according to the position of the contour point, ignoring the pixel points close to the edge of the image; if no valid extreme point is found, return null and display an error message; S223, checking whether the four extracted extreme value points can form a rectangle, if not, returning null and prompting detection failure; S224, if the extreme points are valid and form a rectangle, a perspective transformation is performed to map the extracted extreme points to an area of ​​the target width and height; S225, perform a preliminary detection of the red scale line, if no valid red scale line is detected, return an error message, and re-detect until a valid red scale line is detected; S226, calculating the direction vector and updating the coordinates of the upper leftmost extreme point and the upper rightmost extreme point based on the direction vector. If the updated coordinates exceed the image boundary, returning null and prompting the user; S227, if steps 221 to 226 are all completed successfully, return the updated extreme point dictionary.

8. The measurement method based on KAN and MR according to claim 5, characterized in that: The S23 comprises the following steps: S231, performing perspective transformation on the four detected valid extreme value points; S232, taking the four extreme points and the target image size as input, calculating a 3x3 perspective transformation matrix, and converting the points of the source image into points in the target image according to the perspective transformation matrix; S233, using the perspective transformation matrix to map the original image to the target image size, and transform the four extreme points into the target size and viewing angle, and output the image processed by the perspective transformation.

9. The measurement method based on KAN and MR according to claim 5, characterized in that: The S24 comprises the following steps: S241, converting the transformed image into a grayscale image, and performing an inverse binarization operation to obtain a binary image, wherein a binarization threshold is set, and pixel values ​​greater than the threshold are changed to 255, and other pixels are changed to 0; S242, using a vertical structure element to perform multiple dilation operations on the binary image, enhancing the vertical scale line area, and generating a dilated image; S243, performing contour detection on the dilated image to obtain all contours, and screening out vertical contours whose height is greater than twice the width as valid contours; S244, skeletonizing the dilated image to obtain a skeletonized image, extracting the longest vertical line of the red scale line from the skeletonized image, calculating the midpoint of the red scale line, and marking the longest line segment in the image; S245, drawing a valid contour on the new blank image, and performing skeleton processing on the image with the valid contour drawn to obtain a skeletonized image; S246, traversing each pixel point in the skeletonized image, checking the skeleton pixel, and updating the position of the leftmost skeleton point and the rightmost skeleton point; S247, returns the left and right end coordinate values ​​of the complete scale line and the exact coordinate value of the red scale line.

10. A measurement system based on KAN and MR, characterized in that: include: An image acquisition module is used to acquire the scale line image on the measuring tool and perform image preprocessing; A feature extraction module is used to extract features from the preprocessed image to obtain the geometric feature vector of the image; The network establishment module is used to establish a KAN network and perform network training to obtain a trained KAN network. The input vector of the KAN network consists of the geometric features of the left endpoint coordinates, the right endpoint coordinates and the midpoint coordinates of the red scale line. The KAN network includes 3 hidden layers, each with 64 nodes, and the activation function is an adaptive spline. The KAN network uses a single node output to output a predicted scale value. Each hidden layer is composed of a spline function, and the spline function fits the nonlinear relationship of the input features through a piecewise polynomial: ,in, is the spline function, is the training coefficient, is the number of hidden layers; the KAN network uses mean square error as the loss function: ,in, represents the loss value, is the total number of training samples, is the true scale value of the i-th sample, is the predicted scale value of the model for the i-th sample; The KAN network module is used to input the geometric feature vector into the trained KAN network and output the predicted scale value.

Citation Information

Patent Citations

  • Aircraft attitude estimation method based on BP neural network model

    CN115456171A

  • Intelligent identification method for arc-shaped display instrument of main control room of nuclear power plant

    CN117392651A