Screen button detection method and device based on 2D and 3D cameras and related medium thereof
By combining detection methods from 2D and 3D cameras, spatial mapping relationships and feature point matching are established, solving the problem of limited field of view of a single camera. This achieves higher accuracy and more stable screen button detection, suitable for scenarios such as electronic product manufacturing and warehouse logistics.
Patent Information
- Application Number
- CN202310624939.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-30
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-05-30
AI Technical Summary
In existing technologies, the limited field of view of a single camera leads to a decrease in the accuracy and stability of screen button detection. Especially in situations with a large detection range and short detection distance, traditional screen button detection methods cannot obtain complete three-dimensional data of the screen, resulting in a decrease in detection accuracy and stability.
A detection method based on 2D and 3D cameras is adopted to acquire 2D image data and 3D image data respectively, establish spatial mapping relationship, perform image preprocessing and model reconstruction, extract button area data of 3D screen model, and perform feature point matching between 2D button data and 3D button data to obtain the final button data.
It improves the accuracy and stability of screen button detection, avoids the risk of single camera failure, and reduces the possibility of screen damage by eliminating the need for contact detection.
Smart Images

Figure CN116645418B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data detection, and in particular relates to a screen button detection method and device based on 2D and 3D cameras and a related medium thereof. BACKGROUND
[0002] In the prior art, the screen button detection method is a widely used technology, especially in touch screen device applications. Traditional screen button detection methods usually rely on 2D image data, that is, the position and state of the button are detected by analyzing a planar image; however, this method may have some limitations, for example, for buttons with perspective effects or complex three-dimensional shapes, the accuracy and stability may be affected. In the working condition of large detection range and short detection distance, the existing screen button detection method cannot obtain complete three-dimensional data of the screen, and the detection accuracy and stability of the screen button are greatly reduced. SUMMARY
[0003] The embodiments of the present application provide a screen button detection method and device based on 2D and 3D cameras and a related medium thereof, aiming to solve the problem of reduced accuracy and stability of screen button detection due to the limited field of view of a single camera in the prior art.
[0004] In a first aspect, the embodiments of the present application provide a screen button detection method based on 2D and 3D cameras, comprising:
[0005] 2D image data and 3D image data are acquired respectively, and a spatial mapping relationship is established using the 2D image data and the 3D image data to obtain an image mapping result;
[0006] The image mapping result is preprocessed to obtain a preprocessing result; wherein the preprocessing result includes preprocessing results of the 2D image data and the 3D image data;
[0007] Image data matching and model reconstruction processing are performed according to the preprocessing result to obtain a 3D screen model;
[0008] Button region data of the 3D screen model is extracted according to a preset button parameter to obtain 3D button data;
[0009] 2D button data is collected, and feature point matching is performed between the 2D button data and the 3D button data to obtain final button data; wherein the button data is used for uploading to a detection system.
[0010] In a second aspect, the embodiments of the present application provide a screen button detection device based on 2D and 3D cameras, comprising:
[0011] An image mapping unit is configured to acquire 2D image data and 3D image data respectively, and establish a spatial mapping relationship by using the 2D image data and the 3D image data to obtain an image mapping result;
[0012] An image processing unit is configured to perform image preprocessing on the image mapping result to obtain a preprocessing result; wherein the preprocessing result includes preprocessing results of the 2D image data and the 3D image data;
[0013] An image reconstruction unit is configured to perform image data matching and model reconstruction processing according to the preprocessing result to obtain a 3D screen model;
[0014] An image extraction unit is configured to extract button area data of the 3D screen model according to a preset button parameter to obtain 3D button data;
[0015] An image matching unit is configured to acquire 2D button data, and perform feature point matching between the 2D button data and the 3D button data to obtain final button data; wherein the button data is used for uploading to a detection system.
[0016] In a third aspect, an embodiment of the present application provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the 2D and 3D camera-based screen button detection method of the first aspect when executing the computer program.
[0017] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executable on a processor to implement the 2D and 3D camera-based screen button detection method of the first aspect.
[0018] An embodiment of the present application provides a 2D and 3D camera-based screen button detection method, which comprises acquiring 2D image data and 3D image data respectively and establishing a spatial mapping relationship to obtain an image mapping result; performing image preprocessing on the image mapping result to obtain a preprocessing result; performing image data matching and model reconstruction processing according to the preprocessing result to obtain a 3D screen model; extracting button area data of the 3D screen model according to a preset button parameter to obtain 3D button data; acquiring 2D button data, and performing feature point matching between the 2D button data and the 3D button data to obtain final button data. The present application establishes a 3D screen model by using 2D image data and 3D image data, and obtains 3D button data from the 3D screen model, matches the 3D button data with 2D button data, and obtains final button data, thereby improving the accuracy and stability of screen button detection.
[0019] The embodiment of the present application also provides a screen button detection device based on 2D and 3D cameras, a computer device and a storage medium, which also have the beneficial effects described above. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0021] Figure 1 A flowchart of a screen button detection method based on 2D and 3D cameras provided by the embodiment of the present application is shown in the figure.
[0022] Figure 2 Another flowchart of a screen button detection method based on 2D and 3D cameras provided by the embodiment of the present application is shown in the figure.
[0023] Figure 3 A schematic block diagram of a screen button detection device based on 2D and 3D cameras provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0025] It should be understood that when used in the specification and the appended claims, the terms "comprise" and "include" indicate the presence of described features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0026] It should also be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, unless otherwise clearly indicated by the context, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0027] It should be further understood that the term "and / or" used in the present application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0028] Please refer to Figure 1 , Figure 1 A flowchart of a screen button detection method based on 2D and 3D cameras provided by an embodiment of the present application is shown, which specifically includes steps S101-S105.
[0029] S101, 2D image data and 3D image data are acquired respectively, and a spatial mapping relationship is established using the 2D image data and the 3D image data, to obtain an image mapping result;
[0030] S102, image preprocessing is performed on the image mapping result to obtain a preprocessing result; wherein the preprocessing result includes preprocessing results of the 2D image data and the 3D image data;
[0031] S103, image data matching and model reconstruction processing are performed according to the preprocessing result to obtain a 3D screen model;
[0032] S104, button region data of the 3D screen model is extracted according to a preset button parameter to obtain 3D button data;
[0033] S105, 2D button data is collected, and feature point matching is performed between the 2D button data and the 3D button data to obtain final button data; wherein the button data is used for uploading to a detection system.
[0034] In combination with Figure 2 , in step S101, 2D image data and 3D image data are collected using a camera, the 2D image data includes color and texture information of an image, and the 3D image data includes geometric shape and depth information of an object; the 2D image data and the 3D image data are imported into a mapping modeling module of a system, and the mapping modeling module establishes a spatial mapping relationship according to features of the 2D image data and the 3D image data to obtain an image mapping result.
[0035] Specifically, pixel points in the 2D image data are mapped to corresponding points in the 3D image data, and the mapping relationship can be established through internal parameters (such as focal length and principal point position) and external parameters (such as position and orientation of the camera) of the camera, and the establishment of the mapping relationship can convert pixel coordinates in the 2D image data into three-dimensional coordinates in the real world.
[0036] In an embodiment, before step S101, the following steps are included:
[0037] The calibration board of the selected camera is selected and uniformity optimization and background noise minimization of the calibration board are respectively performed to obtain calibration board positioning parameters; an angle point detection algorithm is used to detect the angle points of the calibration board positioning parameters to obtain calibration board angle point parameters; a nonlinear optimization algorithm is used to estimate the internal parameters of the camera to obtain internal parameter calibration values; a three-dimensional reconstruction algorithm is used to estimate the external parameters of the camera to obtain external parameter calibration values; and a polynomial distortion correction is used to respectively perform radial distortion and tangential distortion processing on the internal parameter calibration values and the external parameter calibration values to obtain distortion correction values; wherein the distortion correction values are used to obtain the 2D image data and the 3D image data.
[0038] In the embodiment, according to the calibration requirements, a suitable calibration board is selected, and the calibration board is accurately positioned in space using an accurate positioning device, so as to ensure that different positions and angles of the observed object can be completely observed, and the uniformity optimization of light and the minimization of background noise are considered during the arrangement process. For the uniformity optimization of light, the uniformity and stability of the light source can be ensured, the shadow and strong light change are avoided, the uniform light source such as the diffuse light source or the light box is used, and the position and intensity of the light source are adjusted as needed, and the stray light and reflection are reduced as much as possible. For the minimization of background noise, a plurality of image frames can be collected and averaged, which helps to reduce the image fluctuation caused by random noise and improve the image quality and stability. A precise angle point detection algorithm such as the Shi-Tomasi algorithm based on image gradient, the convolutional neural network based on machine learning, etc. is adopted to detect the angle points on the calibration board with high precision and stability to obtain the calibration board angle point parameters; a nonlinear optimization algorithm (such as Gauss-Newton or Levenberg-Marquardt) or a method based on Bayesian inference is used to estimate the internal parameters of the camera, and the internal parameter calibration values can be obtained; a three-dimensional reconstruction algorithm such as Bundle Adjustment based on multi-view geometric relationship, Structure-from-Motion, etc. is used to obtain the external parameter calibration values (such as the position, attitude and scale information of the camera); and an accurate distortion correction method such as the polynomial distortion correction, the optical model fitting, etc. is adopted to accurately eliminate the radial distortion and the tangential distortion in the image, and the distortion correction values can be obtained after the processing is completed.
[0039] In an embodiment, the step S101 comprises:
[0040] The coordinate values of the 2D image data and the 3D image data are captured by using a high dynamic range algorithm; the coordinate values of the 2D image data are normalized to obtain normalized coordinate values; and the normalized coordinate values are matched with the coordinate values of the 3D image data to obtain the image mapping result.
[0041] In this embodiment, the high dynamic range algorithm is used to capture the coordinate values of the 2D image data and the coordinate values of the 3D image data in the case of sufficient light and noise minimization. The high dynamic range (HDR) algorithm is a technique for synthesizing images with a wide range of brightness; it combines multiple images with different exposure levels into one comprehensive image to reveal a larger dynamic range than a standard image can capture, such as exposure fusion algorithm, ToneMapping algorithm, image alignment algorithm, etc. In addition, common normalization methods include dividing the coordinate values by the width and height of the image to range between 0 and 1; the normalized coordinate values of the 2D image data are matched with the coordinate values in the 3D image data, and the matching is achieved by selected matching algorithms such as nearest neighbor matching, least squares fitting, etc. The purpose of matching is to associate the coordinates in the 2D image data with the corresponding coordinates in the 3D image data to establish a spatial mapping relationship. This mapping relationship can be used to infer the 3D coordinates of the pixels in the 2D image data or to perform inverse projection, thereby realizing the association between 2D and 3D.
[0042] In step S102, the 2D image data in the image mapping result is preprocessed to extract and enhance useful information in the image; the 3D image data in the image mapping result is preprocessed to optimize the 3D image data for further analysis and application; the quality and accuracy of the preprocessing result directly affect the effect of subsequent image processing and related applications.
[0043] In an embodiment, the step S102 includes:
[0044] The image mapping result is subjected to median filtering and histogram equalization to obtain first intervention data; the first processing data is subjected to size adjustment and image registration to obtain second intervention data; the second intervention data is subjected to interference exclusion using a de-artifact algorithm to obtain the preprocessing result.
[0045] In this embodiment, median filtering can effectively reduce noise in the image while preserving edge information; histogram equalization can enhance the contrast and details of the image, making it clearer; size adjustment and image registration are performed on the first intervention data to obtain the second intervention data; size adjustment can adjust the image to the desired size, which can be adjusted according to actual needs; image registration can calibrate the position and angle between different images to align them in space, facilitating subsequent processing; the second intervention data is subjected to interference exclusion using a de-artifact algorithm to obtain the preprocessing result; the de-artifact algorithm can identify and remove interference signals in the image, such as artifacts, halos, etc. By excluding these interferences, a more accurate and clear preprocessing result can be obtained.
[0046] In step S103, the purpose of image data matching is to align the pre-processed 2D image data and 3D image data in space, to provide accurate data basis for subsequent model reconstruction. Based on the image data matching result, model reconstruction processing is performed, and using three-dimensional reconstruction algorithms and techniques, model reconstruction is performed according to 2D image data and corresponding 3D image data, including point cloud reconstruction, surface reconstruction, voxel reconstruction and other methods, and finally a 3D screen model is obtained. The 3D screen model represents the three-dimensional shape and structure of the screen, and can be used for subsequent visualization, analysis and application.
[0047] In an embodiment, the step S103 comprises:
[0048] The feature points of the pre-processing result are extracted using a feature extraction algorithm, and image data matching is performed to obtain a feature point matching result. Based on the feature point matching result, a fundamental matrix estimation and a triangulation processing are performed to obtain a plurality of three-dimensional point data. A dense point cloud is established using the three-dimensional point data, and the dense point cloud is filtered and optimized to obtain a point cloud filtering result. Surface reconstruction is performed on the point cloud filtering result to generate a surface network model, and the surface network model is refined to obtain the 3D screen model. The refinement processing includes smoothing, hole repair, texture mapping and subdivision.
[0049] In this embodiment, the feature points can be extracted using common feature extraction algorithms such as SIFT, SURF, ORB, etc. Then, the corresponding relationship between images is established by matching the extracted feature points to obtain a feature point matching result. Based on the feature point matching result, a fundamental matrix estimation and a triangulation processing are performed. The fundamental matrix estimation can infer the camera motion relationship between images, which is used for subsequent three-dimensional point cloud reconstruction. The triangulation processing can calculate the corresponding three-dimensional point data through the coordinates of the matched feature points. A dense point cloud is established using a plurality of three-dimensional point data, and the dense point cloud is filtered and optimized to remove noise and invalid points to obtain a point cloud filtering result. The point cloud filtering can use filtering algorithms such as Gaussian filtering, statistical filtering, etc. to improve the quality and accuracy of the point cloud. Through surface reconstruction, a surface mesh model representing an object or a scene can be generated. The surface reconstruction method includes geometry-based methods such as Poisson reconstruction, Marching Cubes, etc. The surface network model is refined, including smoothing, hole repair, texture mapping and subdivision, etc. to obtain a more detailed and realistic 3D screen model.
[0050] In step S104, parameters of the button are preset, including position, size, shape and other information of the button. These parameters can be obtained by manual annotation or automatic detection algorithm; according to the preset button parameters, button region data is extracted from the 3D screen model, according to the position and size of the button, a volume or region can be defined in the 3D screen model to represent the range of the button, and the button region data is extracted by geometric shape cutting of the 3D model, and the 3D button data is obtained by button region extraction.
[0051] In an embodiment, the step S104 comprises:
[0052] According to the button parameters, the 3D screen model is regionally filtered to obtain a regionally filtered result; the regionally filtered result is segmented and extracted from the 3D screen model by a region segmentation algorithm to obtain a regionally extracted result; and the regionally extracted result is resampled and interpolated to obtain the 3D button data.
[0053] In this embodiment, according to the position, size and shape of the button, a region or volume can be defined to represent the range of the button, and the regionally filtered result is obtained by regionally filtering the region that meets the button parameters; the regionally filtered result is segmented and extracted from the 3D screen model by a region segmentation algorithm, which can be based on geometric shape to separate the region that meets the button parameters from the model to obtain the regionally extracted result; the regionally extracted result is resampled and interpolated to obtain more detailed and smooth 3D button data, resampling can adjust the sampling rate or density of the data to meet user needs, and interpolation can fill in the gaps or missing data to make the data more complete and continuous. Through resampling and interpolation processing, the final 3D button data can be obtained.
[0054] In step S105, 2D button data is collected by a camera, the camera is aimed at the button on the screen, and an image of the button is taken. Multiple images under different angles and lighting conditions can be used to match the feature points in the 2D button data with the feature points in the 3D button data. Feature point matching algorithms such as FLANN, RANSAC, etc. can be used for matching; according to the matching results of the feature points, the button information corresponding to the 2D button data and the 3D button data is associated to obtain the final button data; the button data includes the position, size, shape, color and other information of the button, and the final button data can be used as input to the detection system for button detection and recognition.
[0055] In an embodiment, the step S105 comprises:
[0056] The feature points of the 2D button data and the 3D button data are extracted to generate feature descriptors; the scores of different feature descriptors with similar features are calculated, and the feature descriptors with adjacent scores are matched to obtain matching candidate data; similarity measurement and index evaluation are performed on the matching candidate data to obtain the final button data; wherein the final button data includes position information and depth information of the screen button.
[0057] In the embodiment, feature extraction algorithms such as SURF, PPF algorithm, etc. are used to extract feature points from 2D button data and 3D button data and generate feature descriptors; feature points represent the salient features of buttons, and descriptors describe the features of the region around the feature points; the similarity scores between different feature descriptors are calculated, and the feature descriptors with adjacent scores are matched as matching candidate data; similarity measurement methods such as Euclidean distance, cosine similarity, etc. can be used to calculate the similarity scores between feature descriptors, and the feature descriptors with adjacent scores represent that they have similar features; similarity measurement and index evaluation are performed on the matching candidate data to obtain the final button data; similarity measurement can be performed by calculating the similarity scores between matching candidate data, such as distance-based similarity measurement; index evaluation can use various evaluation indexes such as precision, recall rate, F1 score, etc. to evaluate the quality of matching candidate data; the final button data can include position information and depth information of the screen button, providing accurate position and distance information of the button.
[0058] In summary, the application can be applied to various screen detection scenarios, such as electronic product manufacturing, warehouse logistics, etc., and does not require contact detection of the screen, avoiding the possibility of screen damage and affecting the detection result. In addition, the application uses the fusion of multiple 2D cameras and 3D cameras to obtain more comprehensive and accurate screen images, thereby improving the detection accuracy and reducing the risk of failure of a single camera to a certain extent.
[0059] In combination Figure 3 As shown in the figure, Figure 3 A schematic block diagram of a screen button detection device based on 2D and 3D cameras provided by the embodiment of the application, the screen button detection device based on 2D and 3D cameras 300 includes:
[0060] An image mapping unit 301 is configured to obtain 2D image data and 3D image data respectively, and establish a spatial mapping relationship using the 2D image data and the 3D image data to obtain an image mapping result;
[0061] An image processing unit 302 is configured to perform image preprocessing on the image mapping result to obtain a preprocessing result; wherein the preprocessing result includes preprocessing results of the 2D image data and the 3D image data;
[0062] The image reconstruction unit 303 is configured to perform image data matching and model reconstruction processing according to the pre-processing result, to obtain a 3D screen model.
[0063] The image extraction unit 304 is configured to extract button region data of the 3D screen model according to a preset button parameter, to obtain 3D button data.
[0064] The image matching unit 305 is configured to collect 2D button data, and perform feature point matching between the 2D button data and the 3D button data, to obtain final button data; wherein the button data is used for uploading to a detection system.
[0065] In the embodiment, the image mapping unit 301 acquires 2D image data and 3D image data respectively, and establishes a spatial mapping relationship by using the 2D image data and the 3D image data, to obtain an image mapping result; the image processing unit 302 performs image pre-processing on the image mapping result, to obtain a pre-processing result; wherein the pre-processing result includes pre-processing results of the 2D image data and the 3D image data; the image reconstruction unit 303 performs image data matching and model reconstruction processing according to the pre-processing result, to obtain a 3D screen model; the image extraction unit 304 extracts button region data of the 3D screen model according to a preset button parameter, to obtain 3D button data; the image matching unit 305 collects 2D button data, and performs feature point matching between the 2D button data and the 3D button data, to obtain final button data; wherein the button data is used for uploading to a detection system.
[0066] In an embodiment, the image mapping unit 301 includes, before the image mapping unit 301:
[0067] The calibration unit is configured to select a calibration board of a camera, and perform light uniformity optimization and background noise minimization processing on the calibration board respectively, to obtain calibration board positioning parameters.
[0068] The detection unit is configured to perform corner point detection on the calibration board positioning parameters by using a corner point detection algorithm, to obtain calibration board corner point parameters.
[0069] The intrinsic parameter unit is configured to estimate intrinsic parameters of the camera by using a nonlinear optimization algorithm, to obtain intrinsic parameter calibration values.
[0070] The extrinsic parameter unit is configured to estimate extrinsic parameters of the camera by using a three-dimensional reconstruction algorithm, to obtain extrinsic parameter calibration values.
[0071] a distortion unit configured to perform radial distortion and tangential distortion on the intrinsic calibration values and the extrinsic calibration values respectively by using a polynomial distortion correction to obtain distortion correction values, wherein the distortion correction values are used to obtain the 2D image data and the 3D image data.
[0072] In an embodiment, the image mapping unit 301 comprises:
[0073] a dynamic unit configured to capture coordinate values of the 2D image data and coordinate values of the 3D image data respectively by using a high dynamic range algorithm;
[0074] a normalization unit configured to normalize the coordinate values of the 2D image data to obtain normalized coordinate values;
[0075] a mapping unit configured to match the normalized coordinate values with the coordinate values of the 3D image data to obtain the image mapping result.
[0076] In an embodiment, the image processing unit 302 comprises:
[0077] a histogram unit configured to perform median filtering and histogram equalization on the image mapping result to obtain first intervention data;
[0078] an intervention unit configured to perform size adjustment and image registration on the first intervention data to obtain second intervention data;
[0079] an artifact unit configured to perform interference exclusion on the second intervention data by using an artifact removal algorithm to obtain the pre-processing result.
[0080] In an embodiment, the image reconstruction unit 303 comprises:
[0081] a matching unit configured to perform image data matching on feature points of the pre-processing result by using a feature extraction algorithm to obtain a feature point matching result;
[0082] a triangulation unit configured to perform basis matrix estimation and triangulation processing according to the feature point matching result to obtain a plurality of three-dimensional point data;
[0083] a filtering unit configured to establish a dense point cloud by using the three-dimensional point data and perform filtering optimization on the dense point cloud to obtain a point cloud filtering result;
[0084] a repairing unit configured to perform surface reconstruction on the point cloud filtering result to generate a surface network model and perform refinement processing on the surface network model to obtain the 3D screen model, wherein the refinement processing comprises smoothing, hole repairing, texture mapping and subdivision.
[0085] In an embodiment, the image extraction unit 304 comprises:
[0086] The screening unit is configured to perform region screening on the 3D screen model according to the button parameter, and obtain a region screening result.
[0087] The extraction unit is configured to perform region segmentation and extraction on the region screening result from the 3D screen model by using a region segmentation algorithm, and obtain a region extraction result.
[0088] The interpolation unit is configured to perform resampling and interpolation processing on the region extraction result, and obtain the 3D button data.
[0089] In an embodiment, the image matching unit 305 comprises:
[0090] The feature unit is configured to extract feature points of the 2D button data and the 3D button data, and generate feature descriptors.
[0091] The candidate unit is configured to calculate scores of feature descriptors with similar features in different feature descriptors, and perform matching candidate on the feature descriptors with adjacent scores, and obtain matching candidate data.
[0092] The index unit is configured to perform similarity measurement and index evaluation on the matching candidate data, and obtain the final button data; wherein the final button data comprises position information and depth information of a screen button.
[0093] Since the embodiments of the device part correspond to the embodiments of the method part, the embodiments of the device part are described in the description of the embodiments of the method part, and are not described herein.
[0094] The embodiment of the present application further provides a computer readable storage medium, which has a computer program stored thereon, and the computer program can implement the steps provided by the above embodiment when executed. The storage medium can include a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various storage medium capable of storing program codes.
[0095] The embodiment of the present application further provides a computer device, which can include a memory and a processor, the memory has a computer program stored therein, and the processor can implement the steps provided by the above embodiment when calling the computer program in the memory. Of course, the computer device can further include various network interfaces, power supplies and other components.
[0096] The various embodiments described in the specification are presented for purposes of illustration and description. Each of the embodiments highlight a different aspect of the application. The embodiments are not mutually exclusive, and can be combined in various manners. The embodiments disclosed herein are not exhaustive of the ways in which the application can be practiced. Numerous modifications and adaptations will be apparent to those skilled in the art. The embodiments disclosed herein are merely exemplary in nature and are not intended to limit the scope of the application.
[0097] It should also be noted that the terms "first", "second", and the like, do not denote any order, quantity, combination, or importance, but rather are used to identify one element from another. Also, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
Claims
1. A 2D and 3D camera based screen button detection method, characterized in that, The method comprises the following steps: respectively acquiring 2D image data and 3D image data, and establishing a spatial mapping relationship by using the 2D image data and the 3D image data to obtain an image mapping result; performing image preprocessing on the image mapping result to obtain a preprocessing result; wherein the preprocessing result comprises preprocessing results of the 2D image data and the 3D image data; performing image data matching and model reconstruction processing according to the preprocessing result to obtain a 3D screen model; extracting button area data of the 3D screen model according to a preset button parameter to obtain 3D button data; collecting 2D button data and performing feature point matching on the 2D button data and the 3D button data to obtain final button data; wherein the button data is used for uploading to a detection system; Before the step of respectively acquiring 2D image data and 3D image data, and establishing a spatial mapping relationship by using the 2D image data and the 3D image data to obtain an image mapping result, the method comprises the following steps: selecting a calibration board of a camera and performing light uniformity optimization and background noise minimization processing on the calibration board respectively to obtain calibration board positioning parameters; performing corner point detection on the calibration board positioning parameters by using a corner point detection algorithm to obtain calibration board corner point parameters; estimating internal parameters of the camera by using a nonlinear optimization algorithm to obtain internal parameter calibration values; estimating external parameters of the camera by using a three-dimensional reconstruction algorithm to obtain external parameter calibration values; performing radial distortion and tangential distortion processing on the internal parameter calibration values and the external parameter calibration values respectively by using a polynomial distortion correction to obtain distortion correction values; wherein the distortion correction values are used for acquiring the 2D image data and the 3D image data; the corner point detection algorithm is a Shi-Tomasi algorithm based on image gradient; and the nonlinear optimization algorithm is a Gauss-Newton algorithm; The step of respectively acquiring 2D image data and 3D image data, and establishing a spatial mapping relationship by using the 2D image data and the 3D image data to obtain an image mapping result comprises the following steps: respectively capturing coordinate values of the 2D image data and coordinate values of the 3D image data by using a high dynamic range algorithm; performing normalization processing on the coordinate values of the 2D image data to obtain normalized coordinate values; and matching the normalized coordinate values with the coordinate values of the 3D image data to obtain the image mapping result.
2. The 2D and 3D camera based screen button detection method according to claim 1, wherein, The step of performing image preprocessing on the image mapping result to obtain a preprocessing result comprises the following steps: performing median filtering and histogram equalization on the image mapping result to obtain first intervention data; performing size adjustment and image registration on the first intervention data to obtain second intervention data; performing interference exclusion on the second intervention data by using an artifact removal algorithm to obtain the preprocessing result. 3.The 2D and 3D camera based screen button detection method of claim 1, wherein, The step of performing image data matching and model reconstruction processing according to the preprocessing result to obtain a 3D screen model comprises the following steps: extracting feature points of the preprocessing result by using a feature extraction algorithm and performing image data matching to obtain a feature point matching result; According to the feature point matching result, a fundamental matrix estimation and a triangulation process are performed to obtain a plurality of three-dimensional point data; A dense point cloud is established using the three-dimensional point data, and the dense point cloud is filtered and optimized to obtain a point cloud filtering result; A surface network model is generated by performing surface reconstruction on the point cloud filtering result, and the surface network model is refined to obtain the 3D screen model; wherein the refinement process includes smoothing, hole repair, texture mapping and subdivision.
4. The 2D and 3D camera based screen button detection method of claim 1, wherein, The 3D screen model is extracted according to the preset button parameters to obtain 3D button data, including: According to the button parameters, the 3D screen model is regionally screened to obtain a regionally screened result; The regionally screened result is segmented and extracted from the 3D screen model using a region segmentation algorithm to obtain a regionally extracted result; The regionally extracted result is resampled and interpolated to obtain the 3D button data.
5. The 2D and 3D camera based screen button detection method according to claim 1, wherein, The 2D button data is collected, and the 2D button data is matched with the 3D button data according to the feature points to obtain the final button data, including: The feature points of the 2D button data and the 3D button data are extracted to generate feature descriptors; The scores of different feature descriptors with similar features are calculated, and the feature descriptors with adjacent scores are matched as candidates to obtain matching candidate data; The similarity of the matching candidate data is measured and the index is evaluated to obtain the final button data; wherein the final button data includes position information and depth information of the screen button.
6. A 2D and 3D camera based screen button detection apparatus, characterized in that, Including: An image mapping unit is configured to obtain 2D image data and 3D image data respectively, and establish a spatial mapping relationship using the 2D image data and 3D image data to obtain an image mapping result; An image processing unit is configured to perform image preprocessing on the image mapping result to obtain a preprocessing result; wherein the preprocessing result includes the preprocessing result of the 2D image data and 3D image data; An image reconstruction unit is configured to perform image data matching and model reconstruction processing according to the preprocessing result to obtain a 3D screen model; An image extraction unit is configured to extract button region data of the 3D screen model according to preset button parameters to obtain 3D button data; An image matching unit is configured to collect 2D button data, and match the 2D button data with the 3D button data according to feature points to obtain final button data; wherein the button data is used for uploading to a detection system; The 2D and 3D camera-based screen button detection device is also used for selecting a calibration board of a camera, and performing light uniformity optimization and background noise minimization processing on the calibration board respectively to obtain calibration board positioning parameters; an angle point detection algorithm is used to detect the calibration board positioning parameters to obtain calibration board angle point parameters; a nonlinear optimization algorithm is used to estimate internal parameters of the camera to obtain internal parameter calibration values; a three-dimensional reconstruction algorithm is used to estimate external parameters of the camera to obtain external parameter calibration values; a polynomial distortion correction is used to perform radial distortion and tangential distortion processing on the internal parameter calibration values and the external parameter calibration values respectively to obtain distortion correction values; wherein the distortion correction values are used to obtain the 2D image data and the 3D image data; the angle point detection algorithm is a Shi-Tomasi algorithm based on image gradients; and the nonlinear optimization algorithm is a Gauss-Newton algorithm. The image mapping unit is specifically used for capturing coordinate values of the 2D image data and coordinate values of the 3D image data respectively by using a high dynamic range algorithm; performing normalization processing on the coordinate values of the 2D image data to obtain normalized coordinate values; and matching the normalized coordinate values with the coordinate values of the 3D image data to obtain the image mapping result.
7. A computer device, comprising: The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the 2D and 3D camera-based screen button detection method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the 2D and 3D camera-based screen button detection method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Data matching method and device, electronic equipment and storage medium
CN113705669A
Defect detection method, system, device, equipment, storage medium and product
CN115809983A