A target localization method based on monocular vision

By analyzing AR marking information and deep learning models for object segmentation and feature extraction, combined with the neural network fusion model to integrate angle information, the problem of insufficient accuracy in complex environments of traditional monocular visual target positioning methods is solved, and high-precision three-dimensional positioning is achieved.

CN118710866BActive Publication Date: 2025-07-01诚芯智联(武汉)科技技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410772135.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-16
Publication Date
2025-07-01
Estimated Expiration
2044-06-16

AI Technical Summary

Technical Problem

The traditional target positioning method based on monocular vision lacks positioning accuracy and robustness in complex environments, and it is difficult to make full use of multi-source information, which cannot meet the application needs of high-precision three-dimensional positioning.

Method used

By obtaining scene image data containing preset AR marks, the two-dimensional coordinate information, angle information and identification code information of the marked points are parsed, and the deep learning models U-Net and ResNet-50 are combined for object segmentation and feature extraction. Finally, the angle information and high-dimensional feature vectors are integrated through the neural network fusion model to obtain the three-dimensional spatial coordinates and direction information of the target object.

Benefits of technology

It improves the accuracy and efficiency of image processing and target recognition, enhances the adaptability in different scenarios, and meets the needs of high-precision three-dimensional positioning applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118710866B_ABST
    Figure CN118710866B_ABST
Patent Text Reader

Abstract

The present invention relates to a target positioning method based on monocular vision, and relates to the field of target positioning. The target positioning method based on monocular vision obtains scene image data containing preset AR markers and performs parsing and processing to obtain two-dimensional coordinate information, angle information, and identification code information of the marker points; based on the two-dimensional coordinate information and identification code information of the marker points, and using a deep learning model to perform object segmentation on the scene image data to obtain a pixel-level mask of the target object. The present invention performs image parsing by using preset AR markers, combines the deep learning models U-Net and ResNet-50 for object segmentation and feature extraction, and integrates the angle information and high-dimensional feature vectors through a neural network fusion model to finally obtain the three-dimensional spatial coordinates and orientation information of the target object, which not only improves the accuracy and efficiency of image processing and target recognition, but also meets the requirements of high-precision three-dimensional positioning applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target localization, and specifically to a target localization method based on monocular vision. Background Art

[0002] In the fields of modern computer vision and image processing, target localization is a key and important research topic with wide applications. Traditional target localization methods mainly rely on multiocular vision, such as sensors like stereo vision or lidar. However, these methods usually have disadvantages such as high equipment cost, complex calculations, and strong environmental dependence. In many practical applications, monocular vision systems have become a very attractive option due to their low cost, easy deployment, and maintenance. Although monocular vision systems have the above advantages, they face many challenges in target localization. For example, monocular vision systems cannot directly obtain depth information in the scene, which makes it difficult to accurately determine the three-dimensional position of an object based on a single frame image. The target objects in monocular images may be affected by various factors such as illumination, occlusion, and background complexity, resulting in increased difficulty in feature extraction and recognition. Traditional methods based on image processing and feature matching often perform poorly in complex environments and are difficult to meet the requirements of high-precision localization. Summary of the Invention

[0003] The present invention aims at the technical problems existing in the prior art, and provides a target localization method based on monocular vision, which solves the problems that traditional target localization methods usually have insufficient localization accuracy and robustness in complex environments, and the multi-modal data fusion processing adopted by traditional target localization methods is relatively simple, making it difficult to fully utilize multi-source information. At the same time, there are deficiencies in obtaining three-dimensional space coordinates and object direction information, and it is difficult to meet the application requirements of precise three-dimensional localization.

[0004] The technical solution of the present invention to solve the above technical problems is as follows: A target localization method based on monocular vision includes the following steps: acquiring scene image data containing preset AR markers and performing parsing processing to obtain two-dimensional coordinate information, angle information, and identification code information of the marker points; performing object segmentation on the scene image data based on the two-dimensional coordinate information and identification code information of the marker points and using a deep learning model to obtain a pixel-level mask of the target object; extracting features from the pixel-level mask of the target object based on a deep convolutional network to obtain a high-dimensional feature vector of the target object; performing semantic localization fusion processing based on the angle information of the marker points and the high-dimensional feature vector to obtain the three-dimensional space coordinates and direction information of the target object.

[0005] Further, the specific process of obtaining the two-dimensional coordinate information, angle information, and identification code information of the marked points is as follows: Perform grayscale processing on the scene image data containing the preset AR markers; perform recognition and matching on the grayscale processed scene image data with the preset AR marker library to obtain each AR marker point in each scene image data. The specific process is as follows: Identify the corner data in the grayscale processed scene image data based on the corner detection algorithm. For each corner, select the corner neighborhood within the set range. For each corner neighborhood, perform normalization processing respectively, generate descriptors for the normalized corner neighborhood based on the SIFT algorithm, and store them in the form of feature vectors; Analyze each AR marker point respectively based on the corner data to obtain the two-dimensional coordinate information and angle information of each AR marker point; Identify and locate the pattern information within each AR marker based on image processing technology; Analyze the pattern information within each AR marker respectively, and convert the analysis results into binary data; Convert the binary data into identification code information in character form.

[0006] Further, the specific steps of obtaining the two-dimensional coordinate information and angle information of each AR marker point are as follows: For each AR marker point, identify the set number of corner points respectively; Analyze the center coordinates based on the coordinates of each corner point with the set number of corner points, and record them as the two-dimensional coordinate information of the corresponding AR marker point; Analyze the included angle of each marker point relative to the horizontal axis of the scene image based on the first corner point and the second corner point. The calculation formula is as follows: where, θ i is the included angle of the i-th AR marker point relative to the horizontal axis of the scene image, (x i ′1,y i ′1) are the coordinates of the first corner point of the i-th AR marker point, (x i ′2,y i ′2) are the coordinates of the second corner point of the i-th AR marker point, i = 1, 2, 3, …, N, and N is the number of AR marker points.

[0007] Further, the specific process of obtaining the pixel-level mask of the target object is as follows: Read the two-dimensional coordinate information and identification code information of each AR marker point, and determine the position of the target object in the scene image based on the preset marker relationship between the marker points and the target object; Establish the size information and position information of the ROI based on the identification code information; Input the image data of the target object covered by the ROI into the pre-trained U-Net model for segmentation processing to obtain the pixel-level mask of the target object.

[0008] Further, the U-Net model includes an encoder for reducing the spatial size of the image, a decoder for restoring the spatial size of the image, and in each step, the decoder will perform feature fusion with the corresponding layer of the encoder through skip connections.

[0009] Further, inputting the image data of the target object covered by the ROI into the pre-trained U-Net model for segmentation processing specifically means performing forward propagation on the image data of the target object covered by the ROI through the U-Net model. At this time, the model outputs a probability map with the same size as the ROI, and converts the probability map into a binary mask based on a set threshold. Among them, the binary threshold processing formula is specifically: where M is the generated binary mask, (x, y) is the pixel coordinate, t is the set threshold, and P is the probability map generated by the U-Net model.

[0010] Further, the specific process of obtaining the high-dimensional feature vector of the target object is as follows: read the pixel-level mask of the target object and perform format conversion processing; adjust the size of the pixel-level mask of the target object after format conversion processing based on the image scaling algorithm; input the pixel-level mask of the target object after mask size adjustment into the pre-trained ResNet-50 model for feature extraction to obtain the high-dimensional feature vector of the target object.

[0011] Further, the pre-training steps of the ResNet-50 model are as follows: obtain the pixel-level mask data of several target objects and establish a pixel-level mask data set; divide the pixel-level mask data set into a mask training data set and a mask verification data set; train the ResNet-50 model based on the mask training data set and calculate the cross-entropy loss function; use the mask verification data set to evaluate the model performance and adjust the model parameters according to the verification results until the expected standard is met.

[0012] Further, the calculation formula of the cross-entropy loss function is as follows: where L is the cross-entropy loss function, r j is the true label with the target category being j, p j is the probability that the ResNet-50 model predicts the jth class, and j = 1, 2, 3,..., C is the number of classes.

[0013] Further, the specific process of obtaining the three-dimensional spatial coordinates and orientation information of the target object is: read the angular information of the marker points and the high-dimensional feature vector of the target object and perform scalar processing; fuse the angular information of the marker points and the high-dimensional feature vector of the target object after scalar processing based on the concatenation strategy; further process the fused feature vector based on multiple fully connected layers in the preset neural network fusion model, and output the three-dimensional spatial coordinates and orientation information of the target object through the output layer of the neural network fusion model.

[0014] The beneficial effects of the present invention are as follows: By using preset AR markers for image parsing, combining the deep learning models U-Net and ResNet-50 for object segmentation and feature extraction, and integrating the angle information and high-dimensional feature vectors through a neural network fusion model, the three-dimensional spatial coordinates and orientation information of the target object are finally obtained. This not only improves the accuracy and efficiency of image processing and target recognition, but also enhances the adaptability in different scenarios, meeting the requirements of high-precision three-dimensional positioning applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flowchart of a target positioning method based on monocular vision according to the present invention.

[0016] Figure 2 It is a specific step flowchart for obtaining the two-dimensional coordinate information and angle information of each AR marker point in the target positioning method based on monocular vision according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application.

[0018] In the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, "a plurality" means two or more, unless otherwise specifically defined.

[0019] In the description of the present application, the term "for example" is used to mean "used as an example, illustration, or explanation". Any embodiment described as "for example" in the present application is not necessarily construed as being more preferred or having more advantages than other embodiments. In order for any person skilled in the art to implement and use the present invention, the following description is given. In the following description, details are set forth for purposes of explanation. It should be understood that those skilled in the art can recognize that the present invention can be implemented without using these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid unnecessary details from obscuring the description of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope that conforms to the principles and features disclosed in the present application.

[0020] The general idea for the problems in the embodiments of this application is as follows:

[0021] First, use a monocular camera to obtain scene image data containing a preset AR marker, and perform grayscale processing and parsing to obtain the two-dimensional coordinate information, angle information, and identification code information of the marker points. Then, based on the two-dimensional coordinate information and identification code information of the marker points, use the deep learning model U-Net to perform object segmentation on the scene image data to obtain a pixel-level mask of the target object. Next, based on the deep convolutional network ResNet-50, perform feature extraction on the pixel-level mask of the target object to obtain a high-dimensional feature vector of the target object. Then, perform fusion processing based on the angle information of the marker points and the high-dimensional feature vector, use a neural network fusion model, and fuse these features through a concatenation strategy. Finally, based on the fused feature vector, use the fully connected layer of the neural network model for further processing, and finally output the three-dimensional spatial coordinates and orientation information of the target object.

[0022] Please refer to Figure 1 , an embodiment of the present invention provides a technical solution: a target positioning method based on monocular vision, including the following steps: using a monocular camera to obtain scene image data containing a preset AR marker and performing parsing processing to obtain the two-dimensional coordinate information, angle information, and identification code information of the marker points; based on the two-dimensional coordinate information and identification code information of the marker points and using the deep learning model U-Net to perform object segmentation on the scene image data to obtain a pixel-level mask of the target object; based on the deep convolutional network ResNet-50, perform feature extraction on the pixel-level mask of the target object to obtain a high-dimensional feature vector of the target object; perform semantic localization fusion processing based on the angle information of the marker points and the high-dimensional feature vector, specifically by using a neural network fusion model for fusion processing to obtain the three-dimensional spatial coordinates and orientation information of the target object.

[0023] Specifically, the specific process of obtaining the two-dimensional coordinate information, angle information, and identification code information of the marked points is as follows: Gray-scale the scene image data containing the preset AR markers; Identify and match the gray-scaled scene image data with the preset AR marker library to obtain each AR marker point in each scene image data. The AR marker library contains various different AR marker designs. The specific process is as follows: Based on the corner detection algorithm Harri, identify the corner data in the gray-scaled scene image data. Corners are points with large local variations in the image, usually the intersections of object edges. For each corner, select the corner neighborhood within a set range. This neighborhood contains sufficient information to describe the local features of the corner. For each corner neighborhood, perform normalization processing respectively, including rotation normalization to make the descriptor insensitive to rotation, size normalization to make the descriptor insensitive to image scaling, and brightness normalization to make the descriptor insensitive to brightness changes. Generate descriptors for the normalized corner neighborhoods based on the SIFT algorithm and store them in the form of feature vectors. The SIFT algorithm is Scale-Invariant Feature Transform: Extract the local gradient direction and magnitude of the corner neighborhood, and summarize this information into a direction histogram. In this way, each corner will be associated with a feature vector composed of the neighborhood gradient information within its set range; Analyze each AR marker point respectively based on the corner data to obtain the two-dimensional coordinate information and angle information of each AR marker point; Identify and locate the pattern information within each AR marker based on image processing techniques. Pattern location: Use computer vision techniques, such as edge detection and feature point matching, to identify the exact position and boundary of the pattern. This pattern may be a region of the marker, such as the two-dimensional matrix of a QR code. Analyze the pattern information within each AR marker respectively, specifically analyze the internal structure of the pattern, such as the black and white squares in a QR code. The layout of these squares represents the encoded information; And convert the analysis result into binary data. For example, in a QR code, black squares may represent "1" and white squares represent "0"; Convert the binary data into identification code information in character form.

[0024] In this implementation, a monocular camera is used to obtain a scene image containing a preset AR marker and grayscale it to simplify subsequent image processing steps and reduce computational complexity. The grayscale processing in this step can reduce the complexity and computational requirements of the image, improve the processing speed and efficiency, enhance the adaptability and efficiency of the algorithm in real-time applications by simplifying the image preprocessing steps. The grayscale processed image is identified and matched with a preset AR marker library to obtain each AR marker point in each scene image. The AR marker library contains various AR markers with different designs. This step can quickly match and identify the AR markers in the scene through the preset AR marker library, improve the accuracy of recognition, enhance the recognition ability of the system in different scenes and diverse markers, and solve the problem of difficult marker recognition in complex scenes. Based on the Harris corner detection algorithm, the corner data in the grayscale processed image is identified. Corners are points with large local variations in the image, usually the intersections of object edges. The Harris corner detection algorithm in this step can effectively identify the significant feature points in the image, provide a basis for subsequent feature extraction, improve the accuracy of feature detection, and ensure the reliability of subsequent processing steps. For each corner, a corner neighborhood within a set range is selected, including rotation normalization, size normalization, and brightness normalization. This step makes the feature descriptor more robust through normalization processing, adapts to different shooting angles, distances, and lighting conditions, solves the problem of unstable features caused by rotation, scaling, and brightness changes in the image, and improves the accuracy of feature matching. A descriptor is generated for the normalized corner neighborhood based on the SIFT algorithm.The SIFT algorithm extracts the local gradient direction and magnitude in the neighborhood of corner points, and summarizes this information into an orientation histogram. The SIFT descriptor in this step can effectively capture the local features in the image, has high distinctiveness, improves the accuracy of feature matching, enhances the robustness of the system in complex scenes. Analyze each AR marker point separately based on the corner point data to obtain the two-dimensional coordinate information of each AR marker point and the angle relative to the horizontal axis of the scene image. This step accurately obtains the spatial position and orientation information of the marker points, provides reliable data for subsequent positioning and navigation, solves the problem of inaccurate marker point positioning in traditional methods, and enhances the accuracy and reliability of the system. Use image processing technology to identify and locate the pattern information within each AR marker, such as QR codes, and use edge detection and feature point matching technology to identify the exact position and boundary of the pattern. This step can further improve the accuracy of the identification code information through accurate pattern recognition, solves the problem of difficult pattern positioning in complex scenes, and improves the decoding accuracy of the pattern information. Analyze the pattern information within each AR marker, convert the analysis result into binary data, and further convert it into the identification code information in character form. This step ensures the correctness and consistency of the identification code information through accurate binary conversion and decoding, enhances the system's ability to identify and decode complex encoded patterns, and improves the accuracy and reliability of the identification code information.

[0025] Specifically, as Figure 2 shown, the specific steps to obtain the two-dimensional coordinate information and angle information of each AR marker point are as follows: For each AR marker point, identify the set number of corner points respectively, usually four corner points. For square or rectangular AR markers, the four corner points are usually the four outer corners of the pattern. These corner points are easily detected by detection algorithms such as Harris corner detection due to their significant angle changes. If the AR marker is a non-standard shape, such as a circle or polygon, the corner points are selected based on the significant feature points of the shape, such as protruding or recessed points; Analyze the center coordinates based on the coordinates of each corner point with the set number of corner points, and record them as the two-dimensional coordinate information of the corresponding AR marker point. The specific formula for calculating the center coordinates is as follows: where, (x,y) is the center coordinate of the corner point, (x n ,y n ) is the coordinate of the nth corner point identified as set, and n is the set number of corner points identified; For each AR marker point, arbitrarily select one side respectively, and record the two ends of the side as the first corner point and the second corner point; Analyze the angle of each marker point relative to the horizontal axis of the scene image based on the first corner point and the second corner point. The calculation formula is as follows: where, θ i is the angle of the ith AR marker point relative to the horizontal axis of the scene image, (x i ′1,y i′1) is the coordinate of the first corner point of the i-th AR marker point, (x i ′2, y i ′2) is the coordinate of the second corner point of the i-th AR marker point, i = 1, 2, 3, …, N, where N is the number of AR marker points.

[0026] In this implementation, for each AR marker point, the set number of corner points is identified respectively, usually four corner points. For a square or rectangular AR marker, these four corner points are usually the four outer corners of the pattern. If the AR marker is a non-standard shape, such as a circle or a polygon, the selection of corner points is based on the significant feature points of the shape, such as protruding or recessed points. This step using significant feature points, such as corner points, helps to improve the accuracy and reliability of detection. The flexible handling of standard and non-standard shapes enhances the adaptability of the algorithm, solves the problem of high difficulty in corner point recognition in complex scenarios, especially under AR markers of different shapes, and ensures the accuracy of corner point detection. Based on the coordinates of each corner point corresponding to the set number of corner points, the center coordinates are analyzed and recorded as the two-dimensional coordinate information of the corresponding AR marker point. The specific formula for calculating the center coordinates provides a simple and effective method to determine the position of the AR marker. Using the average position of all corner points to calculate the center coordinates improves the stability and accuracy of positioning, solves the problem that a single corner point may be inaccurately positioned due to noise or other factors, and enhances the reliability of position calculation by using the average value of multiple corner points. For each AR marker point, an edge is arbitrarily selected respectively, and the two ends of the edge are recorded as the first corner point and the second corner point respectively. Based on the first corner point and the second corner point, the included angle of each marker point relative to the horizontal axis of the scene image is analyzed. By calculating the angle information, the direction of the marker point can be determined, providing additional geometric information. Using simple geometric calculation formulas improves the calculation efficiency and real-time performance, solves the problem of relying only on position and lacking direction information, and improves the comprehensiveness and accuracy of target positioning by adding angle information.

[0027] Specifically, the specific process of obtaining the pixel-level mask of the target object is as follows: Read the two-dimensional coordinate information and identification code information of each AR marker point, and determine the position of the target object in the scene image based on the preset marker relationship between the marker point and the target object. Usually, the AR marker is located at the center or other key positions of the target object; Based on the identification code information, establish the size information and position information of the ROI. For example, if the AR marker is always located at the center of the object, the ROI can be directly defined around the marker; If the marker is located at one end of the object, then the ROI needs to be offset accordingly; The image data of the target object covered by the ROI is input into a pre-trained U-Net model for segmentation processing to obtain the pixel-level mask of the target object.

[0028] The U-Net model includes an encoder for reducing the spatial dimensions of an image, and the encoder consists of multiple convolutional layers and max-pooling layers, while increasing the depth of the feature channels, which helps the model capture and learn abstract and complex features in the image, a decoder for restoring the spatial dimensions of the image, and at each step, the decoder performs feature fusion with the corresponding layer of the encoder through skip connections, and the decoder contains multiple upsampling layers and convolutional layers for gradually reducing the depth of the feature channels, which helps the model utilize low-level and high-level features when reconstructing a high-resolution output. The skip connections directly connect the feature maps in the encoder to the corresponding layers in the decoder, which can ensure that sufficient detail information is retained even in the deep layers of the network and avoid losing important spatial information during the upsampling process; before inputting the image data of the target object covered by the ROI into the pre-trained U-Net model, first adjust the image of the ROI region to the input size required by the U-Net model, such as 256x256 or 512x512 pixels, ensure that the ratio and quality of the image content are maintained during the adjustment process, normalize the image so that the pixel value range is usually between 0 and 1, which is usually achieved by dividing by 255, that is, the maximum possible value of the pixel value; the specific formula is as follows: T′ = resize(T, s); where T′ is the image after adjusting the image of the ROI region to the input size required by the U-Net model, resize is the function for adjusting the image size, and s is the target size, such as 256 or 512; Among them, T′ is the image after adjusting the image of the ROI region to the input size required by the U-Net model, and 255 is the maximum possible value of the original pixel value.

[0029] In this implementation, the two-dimensional coordinate information and identification code information of each AR marker point are read, and based on the preset marker relationship between the marker point and the target object, the position of the target object in the scene image is determined. Usually, the AR marker is located at the center or other key positions of the target object. Through the coordinates and identification code of the AR marker, the initial position of the target object can be accurately determined in this step, improving the accuracy of ROI definition, enhancing the precision of target object position determination, and solving the problem of inaccurate positioning in traditional methods. If the AR marker is always located at the center of the object, the ROI can be directly defined around the marker. If the marker is located at one end of the object, the position of the ROI is appropriately offset according to the identification code information. This step flexibly defines the ROI to ensure that the target object is completely contained within the ROI, improving the flexibility and adaptability of ROI definition regardless of whether the marker point is at the center or edge of the object, and solving the problem that a fixed ROI is difficult to adapt to different scenarios. The image in the ROI area is adjusted to the input size required by the U-Net model, such as 256x256 or 512x512 pixels. Through adjusting the image size and normalization processing, this step ensures that the input data is suitable for the U-Net model, enhancing the model's performance and solving the problem of inconsistent image size and pixel value range, improving the standardization and consistency of model processing. The U-Net model includes an encoder and a decoder structure. The encoder consists of multiple convolutional layers and max-pooling layers to increase the depth of the feature channels and capture the abstract and complex features in the image. The decoder is used to restore the spatial size of the image and perform feature fusion with the corresponding layers of the encoder through skip connections at each step. The decoder contains multiple upsampling layers and convolutional layers to gradually reduce the depth of the feature channels and reconstruct the high-resolution output. The structural design of U-Net ensures that sufficient detail information is retained while reducing the image spatial size, improving the segmentation accuracy and solving the problem of detail information loss in traditional segmentation methods, enhancing the accuracy and detail retention of pixel-level segmentation. The adjusted and normalized ROI image data is input into the pre-trained U-Net model for forward propagation. The model outputs a probability map with the same size as the ROI. Based on a set threshold, usually 0.5, the probability map is converted into a binary mask. Through the threshold processing of the probability map, this step accurately distinguishes the target object from the background, generating a clear binary mask and solving the problem of unclear distinction between the target object and the background in traditional segmentation methods, improving the reliability of the segmentation result. Morphological operations, such as erosion and dilation, are performed on the generated binary mask to remove small noise points and fill small holes in the target object. The morphological operations in this step further optimize the mask, improving the integrity and accuracy of the segmentation result, solving the problems of noise points and holes in the mask, and enhancing the precision and usability of the segmentation result.

[0030] Specifically, inputting the image data of the target object covered by the ROI into the pre-trained U-Net model for segmentation processing means performing forward propagation of the image data of the target object covered by the ROI through the U-Net model. At this time, the model outputs a probability map with the same size as the ROI, where the value of each pixel represents the probability that the pixel belongs to the target object. Based on a set threshold, usually 0.5, the probability map is converted into a binary mask. Pixels with a probability higher than the threshold are marked as 1, that is, the target object, and those lower than the threshold are marked as 0, that is, the background. Then, morphological operations such as erosion and dilation are performed on the binary mask to remove small noise and fill small holes in the target object. Among them, the binary threshold processing formula is specifically: Among them, M is the generated binary mask, (x, y) is the pixel coordinate, t is the set threshold, usually 0.5, and P is the probability map generated by the U-Net model.

[0031] In this implementation, the image in the ROI area is adjusted to the input size required by the U-Net model, such as 256x256 or 512x512 pixels, to ensure the consistency of the model input data. This step ensures that the image data matches the input requirements of the U-Net model, improves the processing efficiency and effect of the model, solves the input problem caused by inconsistent image sizes, and ensures the smooth progress of subsequent segmentation processing. Then, the image is normalized so that the pixel value range is usually between 0 and 1, which is achieved by dividing by 255. This step standardizes the input image data, ensures the consistency of the input data, facilitates model processing, solves the problem of inconsistent pixel value ranges, and improves the standardization and consistency of model processing. Then, the adjusted and normalized ROI image data is input into the pre-trained U-Net model for forward propagation, and the model outputs a probability map with the same size as the ROI. This step uses the pre-trained U-Net model to quickly and efficiently generate a probability map representing the probability that each pixel belongs to the target object. By using a deep learning model, it solves the problem of low segmentation efficiency in traditional methods and improves the segmentation accuracy and speed. Based on a set threshold, usually 0.5, the probability map is converted into a binary mask. This simple and effective binary processing step can clearly distinguish the target object from the background.

[0032] Specifically, the specific process of obtaining the high-dimensional feature vector of the target object is as follows: Read the pixel-level mask of the target object and perform format conversion processing, that is, convert the binary mask into a format suitable for input to ResNet-50. ResNet usually accepts inputs of three-channel RGB, so copy the single-channel mask to three color channels to make it a pseudo-color image; Based on an image scaling algorithm such as bilinear interpolation, adjust the size of the pixel-level mask of the target object after format conversion processing. ResNet-50 usually requires the input image to be 224x224 pixels in size. Use the image scaling algorithm to adjust the mask to this size to ensure that the proportion and quality of the image content are maintained during the adjustment process; Input the pixel-level mask of the target object with adjusted mask size into the pre-trained ResNet-50 model for feature extraction to obtain the high-dimensional feature vector of the target object, that is, propagate the normalized mask image through the ResNet-50 model. To extract features rather than perform classification, usually the output of a certain layer in the network is extracted. For example, the output of the layer before the last convolutional layer can be used as the feature vector. This layer is usually called the bottleneck layer before the fully connected layer, and the output is obtained from the selected layer. These outputs will serve as high-dimensional feature vectors.

[0033] In this implementation, the pixel-level mask of the target object is read. This is usually a single-channel binary image. The single-channel mask is copied to three color channels to make it a pseudo-color image, which is suitable for input into the ResNet-50 model. The format conversion in this step ensures that the single-channel mask image can adapt to the input requirements of ResNet-50, improves the compatibility of processing, solves the problem that the single-channel mask cannot be directly input into ResNet-50, and ensures the consistency and compatibility of the data format. Using an image scaling algorithm, such as bilinear interpolation, the pseudo-color image is adjusted to the input size required by ResNet-50, which is 224x224 pixels. This step ensures that the input image matches the size requirements of the ResNet-50 model, facilitates subsequent feature extraction, solves the problem of inconsistent input image sizes, and ensures the standardization and consistency of model processing, improving the processing efficiency. The resized image is normalized so that the pixel values range from 0 to 1. The normalization of the input image data in this step ensures the consistency of the data, facilitates the processing of the ResNet-50 model, solves the problem of inconsistent pixel value ranges, and improves the standardization and consistency of model processing. The normalized image is input into the pre-trained ResNet-50 model for forward propagation. In order to extract features rather than perform classification, the output of a certain layer in the network is extracted. For example, the output of the layer before the last convolutional layer is used as the feature vector. This step uses the pre-trained ResNet-50 model to quickly and efficiently extract the high-dimensional feature vector of the image. These feature vectors contain rich image information, solve the problems of low efficiency and small amount of information in traditional feature extraction methods, and improve the accuracy and efficiency of feature extraction. The output of the bottleneck layer in the ResNet-50 model is extracted as the high-dimensional feature vector. The bottleneck layer feature vector in this step contains the most important high-level image features, has a high representation ability, provides rich and accurate high-dimensional feature vectors, and provides a solid foundation for subsequent object localization and classification.

[0034] Specifically, the pre-training steps of the ResNet-50 model are as follows: Obtain the pixel-level mask data of several target objects and establish a pixel-level mask dataset; Divide the pixel-level mask dataset into a mask training dataset and a mask validation dataset; Train the ResNet-50 model based on the mask training dataset and calculate the cross-entropy loss function for classification tasks; The calculation formula of the cross-entropy loss function is as follows: where L is the cross-entropy loss function, r j is the true label of the target category j, usually one-hot encoded, p jis the probability that the ResNet-50 model predicts the j-th class, where j = 1, 2, 3, …, C is the number of classes. The model performance is evaluated using a masked validation dataset, and the model parameters are adjusted according to the validation results until they meet the expected standards; the ResNet-50 model includes an initial convolutional layer followed by multiple residual blocks, each block containing 3 convolutional layers and ending with a global average pooling layer and a fully connected layer, and the output of every two layers is added to the output of the third layer through a skip connection. The weights are initialized using He initialization or other initialization methods suitable for deep convolutional networks. The weights in the network are updated using gradient descent or other optimization algorithms, such as Adam, according to the gradients calculated from the loss function. The residual structure allows the gradients to flow directly to earlier layers, and the loss on the training data is gradually reduced through multiple epochs until the model performance stabilizes or reaches a predetermined number of iterations; among them, the calculation formula for the residual block is as follows: Sc = ReLU{F[a, (W d )] + a}; where Sc is the output of the residual block, a is the data input into the residual block, F[a, (W d )] is the function processed by a series of convolutional layers in this residual block, and these layers transform the input a using the weight set (W d ), (W d ) is the weight set of all layers in the residual block, +a is the residual connection, that is, the original input a is directly added to the output after a series of transformations, and ReLU is the activation function used to increase non-linearity; among them, the weight update formula, that is, the calculation formula for gradient descent, is as follows: where W′ is the updated weight, W is the weight before update, κ is the learning rate, which is a hyperparameter used to control the step size of weight update during the optimization process, is the gradient of the cross-entropy loss function with respect to the weight W, representing the sensitivity of the cross-entropy loss function to each weight parameter, and calculating this gradient is achieved through the backpropagation algorithm, which is used to guide how to adjust the weights to reduce the loss.

[0035] In this implementation, pixel-level mask data of several target objects is collected, which is used to train and validate the ResNet-50 model. The rich dataset in this step is the basis for model training, which can improve the generalization ability and performance of the model. By establishing a rich pixel-level mask dataset, the problem of poor training effect caused by insufficient data is solved. The pixel-level mask dataset is divided into a training dataset and a validation dataset. For example, 80% is used for training and 20% is used for validation. The dataset division helps the training and validation of the model, ensuring that the performance of the model is effectively evaluated and optimized. By reasonable dataset division, the problem of model overfitting in the training process is solved, and the generalization ability of the model is improved. For the classification task, the cross-entropy loss function is used to measure the difference between the model prediction and the true label. The cross-entropy loss function in this step is applicable to multi-classification tasks and can effectively evaluate the classification performance of the model. By using the cross-entropy loss function, the problem of inaccurate loss calculation in multi-classification tasks is solved, and the classification effect of the model is improved. The ResNet-50 model is trained based on the masked training dataset, and the cross-entropy loss function is used to optimize the model parameters. Through a large amount of training data in this step, the model can fully learn the features of the target object, improve the recognition ability of the model, solve the problem of lack of data support in the training process of the model, and ensure that the model can learn effective features. The masked validation dataset is used to evaluate the model performance, and the model parameters are adjusted according to the validation results until they meet the expected standards. The validation process in this step can effectively detect the performance of the model on new data and prevent model overfitting. Through validation and tuning, the generalization ability and practical application effect of the model are improved, and the problem of inconsistent performance of the model on different datasets is solved. The ResNet-50 model includes an initial convolutional layer, followed by multiple residual blocks. Each block contains 3 convolutional layers and ends with a global average pooling layer and a fully connected layer. The output of every two layers is added to the output of the third layer through a skip connection. The weights are initialized using He initialization or other initialization methods suitable for deep convolutional networks. The residual structure in this step can effectively solve the problem of gradient disappearance in the training of deep neural networks and improve the training efficiency of the model. By using the residual structure, the problem of difficult training of deep neural networks is solved, and the stability and performance of the model are improved. Gradient descent or other optimization algorithms, such as Adam, are used to update the weights in the network according to the gradients calculated by the loss function. Using effective optimization algorithms in this step can quickly converge to the optimal solution and improve the training efficiency. By using optimization algorithms and gradient updates, the problems of difficult adjustment of model parameters and slow convergence are solved, and the training speed and effect of the model are improved. The design of the residual block enables the model to better retain information during the training process, improves the expression ability of the model, solves the problem of difficult training of traditional deep networks, and improves the training stability and effect of the model through the residual block.

[0036] Specifically, the specific process of obtaining the three-dimensional spatial coordinates and orientation information of the target object is as follows: Read the angular information of the marker points and the high-dimensional feature vectors of the target object, and perform scalar processing; Based on the concatenation strategy, fuse the angular information of the marker points and the high-dimensional feature vectors of the target object after scalar processing; Further process the fused feature vectors based on multiple fully connected layers in the preset neural network fusion model, and output the three-dimensional spatial coordinates and orientation information of the target object through the output layer of the neural network fusion model.

[0037] In this implementation, the angular information of each AR marker point is read from the image analysis. The angular information is usually expressed as the angle relative to the horizontal axis of the image. Obtaining accurate angular information in this step provides a reliable data basis for subsequent fusion processing, solves the problem of inaccurate acquisition of angular information, and improves the accuracy of the orientation information of the target object. Extract the high-dimensional feature vectors of the target object from the pre-trained ResNet-50 model. These feature vectors contain rich image information. The high-dimensional feature vectors in this step provide detailed visual features of the target object, which helps to improve the accuracy of the three-dimensional spatial coordinates and orientation information, solves the problems of low feature extraction efficiency and small amount of information in traditional methods, and improves the accuracy and efficiency of feature extraction. Convert the angular information into sine and cosine values to capture its periodic characteristics. The angular information after scalar processing in this step is more easily fused with the high-dimensional feature vectors, ensuring the effective expression of the angular information, solving the problem of inaccurate expression that may occur when the angular information is directly used, and improving the effect of data fusion. Concatenate the angular information after scalar processing with the high-dimensional feature vectors of the target object to form a comprehensive feature vector. The concatenated comprehensive feature vector in this step contains the detailed visual features and orientation information of the target object, providing rich inputs for subsequent neural network processing. Through data fusion, the problem of limited expression ability of single features is solved, and the prediction ability of the model for the three-dimensional spatial coordinates and orientation information is improved. Input the comprehensive feature vector into the preset neural network fusion model, and further process it through multiple fully connected layers. Through the processing of multiple fully connected layers, features can be further extracted and optimized, improving the expression ability and prediction accuracy of the model, solving the problem of insufficient single-level feature processing ability, and improving the accuracy of feature expression through deep processing. Through the output layer of the neural network fusion model, predict the three-dimensional spatial coordinates and orientation information of the target object. The finally output three-dimensional spatial coordinates and orientation information have high accuracy and can meet the high-precision positioning requirements, solving the problem of insufficient positioning accuracy in traditional methods, and improving the prediction accuracy of the three-dimensional spatial coordinates and orientation information through the deep learning ability of the fusion model.

[0038] It should be noted that in the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0039] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0040] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0041] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0042] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0043] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0044] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A target positioning method based on monocular vision, characterized in that: The following steps are involved: Obtain scene image data containing preset AR markers and perform parsing to obtain two-dimensional coordinate information, angle information, and identification code information of the marker points; Based on the two-dimensional coordinate information and identification code information of the marker points, the scene image data is segmented using a deep learning model to obtain a pixel-level mask of the target object. Based on the deep convolutional network, the pixel-level mask of the target object is extracted to obtain the high-dimensional feature vector of the target object; Based on the angle information of the marker points and the high-dimensional feature vector, semantic positioning fusion processing is performed to obtain the three-dimensional spatial coordinates and direction information of the target object; The specific process of obtaining the two-dimensional coordinate information, angle information and identification code information of the marking point is as follows: Grayscale the scene image data containing the preset AR marker; The grayscaled scene image data is identified and matched with the preset AR marker library to obtain each AR marker point in each scene image data. The specific process is as follows: the corner point data in the grayscaled scene image data is identified based on the corner point detection algorithm. For each corner point, the corner point neighborhood within the set range is selected. For each corner point neighborhood, normalization is performed respectively. The descriptor of the normalized corner point neighborhood is generated based on the SIFT algorithm and stored in the form of a feature vector. Analyze each AR marker point based on the corner point data to obtain the two-dimensional coordinate information and angle information of each AR marker point; Identify and locate the pattern information in each AR marker based on image processing technology; Analyze the pattern information in each AR marker separately and convert the analysis results into binary data; Convert binary data into identification code information in character form; The specific process of obtaining the three-dimensional spatial coordinates and direction information of the target object is: Read the angle information of the marking point and the high-dimensional feature vector of the target object, and perform scalar processing; Based on the series strategy, the angle information of the marker points after scalar processing and the high-dimensional feature vector of the target object are fused; The fused feature vectors are further processed based on multiple fully connected layers in the preset neural network fusion model, and the three-dimensional spatial coordinates and direction information of the target object are output through the output layer in the neural network fusion model.

2. The target positioning method based on monocular vision according to claim 1, characterized in that: The specific steps to obtain the two-dimensional coordinate information and angle information of each AR marker point are as follows: For each AR marker point, identify the set number of corner points; The center coordinates are analyzed based on the coordinates of each corner point of the set number of corner points, and recorded as the two-dimensional coordinate information of the corresponding AR marker point; The angle of each marker point relative to the horizontal axis of the scene image is analyzed based on the first corner point and the second corner point. The calculation formula is as follows: ; in, For the The angle of the AR markers relative to the horizontal axis of the scene image, For the The coordinates of the first corner point of the AR marker, For the The coordinates of the second corner point of the AR marker, , is the number of AR markers.

3. The target positioning method based on monocular vision according to claim 2, characterized in that: The specific process of obtaining the pixel-level mask of the target object is as follows: Read the two-dimensional coordinate information and identification code information of each AR marker point, and determine the position of the target object in the scene image based on the marker relationship between the preset marker point and the target object; Establishing size information and position information of ROI based on identification code information; The image data of the target object covered by the ROI is input into the pre-trained U-Net model for segmentation processing to obtain the pixel-level mask of the target object.

4. The target positioning method based on monocular vision according to claim 3 is characterized in that: The U-Net model includes an encoder for reducing the spatial size of an image and a decoder for restoring the spatial size of an image, and in each step, the decoder performs feature fusion with the corresponding layer of the encoder through a jump connection.

5. The target positioning method based on monocular vision according to claim 3, characterized in that: The image data of the target object covered by the ROI is input into the pre-trained U-Net model for segmentation processing. Specifically, the image data of the target object covered by the ROI is forward propagated through the U-Net model. At this time, the model outputs a probability map of the same size as the ROI, and converts the probability map into a binary mask based on the set threshold. The binary threshold processing formula is as follows: ; in, is the generated binary mask, is the pixel coordinate, is the set threshold, Probability map generated for the U-Net model.

6. The target positioning method based on monocular vision according to claim 5, characterized in that: The specific process of obtaining the high-dimensional feature vector of the target object is as follows: Read the pixel-level mask of the target object for format conversion; The pixel-level mask of the target object after the format conversion is resized based on the image scaling algorithm; The pixel-level mask of the target object after mask size adjustment is input into the pre-trained ResNet-50 model for feature extraction to obtain a high-dimensional feature vector of the target object.

7. The target positioning method based on monocular vision according to claim 6, characterized in that: The pre-training steps of the ResNet-50 model are as follows: Obtain pixel-level mask data of several target objects and establish a pixel-level mask dataset; Divide the pixel-level mask dataset into a mask training dataset and a mask verification dataset; The ResNet-50 model is trained based on the mask training dataset, and the cross entropy loss function is calculated; The model performance was evaluated using the masked validation dataset and the model parameters were adjusted based on the validation results until they met the expected standards.

8. The target positioning method based on monocular vision according to claim 7, characterized in that: The calculation formula of the cross entropy loss function is as follows: ; in, is the cross entropy loss function, The target category is The real label, Predict the first The probability of the class, is the number of categories.

Citation Information

Patent Citations

  • Application of monocular three-dimensional target detection in grassland forest fire positioning

    CN117392366A

  • Meal delivery robot target positioning and obstacle avoidance method and system based on monocular vision

    CN118154687A