Robust Image Feature Matching Against Motion Blur
By using deep learning network to transform and compare the motion descriptor of the image sensor, the problem of low feature matching accuracy caused by motion blur of image sensors is solved, the positioning and mapping performance in AR/VR applications is improved, and the accuracy and stability of SLAM calculations are ensured.
Patent Information
- Application Number
- CN202180015780.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-19
- Filing Date
- 2021-02-08
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-02-08
AI Technical Summary
Existing image feature matching methods have low accuracy when facing image sensor motion blur, resulting in poor positioning and mapping performance in augmented and virtual reality applications.
Deep learning network is used as a key point descriptor motion fuzzy converter or comparison module, and the trained artificial neural network converts the image sensor motion descriptors of different time periods into converted descriptors for feature matching.
Improve feature matching accuracy in different image motion situations, enhance positioning and mapping performance in AR/VR applications, and ensure the accuracy and stability of SLAM calculations.
Smart Images

Figure CN115210758B_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] Augmented reality (AR) technology superimposes virtual content on a user's view of the real world. With the development of AR software development kits (SDKs), smartphone AR has become mainstream in the mobile phone industry. AR SDKs typically provide six degrees of freedom (6DoF) tracking capabilities. A user can use a smartphone's camera to scan the environment, and the smartphone performs visual inertial odometry (VIO) in real time. Once the pose of the camera is continuously tracked, virtual objects can be placed into the AR scene to create an illusion that real and virtual objects are merged together.
[0002] A key point (or "interest point") is a point in an image. Different from other points in the image, a key point has a well-defined spatial position or is otherwise located in the image and remains stable under local and global changes (e.g., scale change, illumination change, etc.). A key point descriptor can be defined as a multi-element vector. This multi-element vector (usually in the scale space) describes the neighborhood of the key point in the image. Examples of key point descriptor frameworks include scale-invariant feature transform (SIFT), speeded-up robust features (SURF), and binary robust invariant scalable keypoints (BRISK). Image features can be defined as key points and corresponding key point descriptors.
[0003] Feature matching between different images (e.g., between different frames of a video sequence) is an important part of many image processing applications. Images captured when an image sensor is moving may have significant motion blur or dynamic blur. Since the descriptions of most widely used image features are very sensitive to motion blur, feature matching is unlikely to succeed when features are extracted from images with different motion blurs. Therefore, there is a need in the art to improve methods for performing feature matching. SUMMARY OF THE INVENTION
[0004] The present disclosure generally relates to methods and systems related to image processing. More specifically, embodiments of the present disclosure provide methods and systems for performing feature matching in augmented reality applications. Embodiments of the present disclosure are applicable to various applications in augmented reality and computer-based display systems.
[0005] An image processing method according to a general configuration includes: calculating descriptors of key points in a first image, the first image being acquired by an image sensor during a first time period, wherein a first motion descriptor describes the motion of the image sensor during the first time period; calculating descriptors of key points in a second image, the second image being acquired by the image sensor during a second time period, the second time period being different from the first time period, wherein a second motion descriptor describes the motion of the image sensor during the second time period; converting, based on the first motion descriptor and the second motion descriptor, the calculated descriptors for the key points in the first image into converted descriptors using a trained artificial neural network; and comparing the converted descriptors with the calculated descriptors for the key points in the second image.
[0006] A computer system according to another general configuration includes one or more memories configured to store a first image and a second image. The first image is acquired by an image sensor during a first time period. The second image is acquired by the image sensor during a second time period. The second time period is different from the first time period. The computer system further includes one or more processors. The one or more memories are further configured to store computer-readable instructions that, when executed by the one or more processors, configure the computer system to: calculate descriptors for key points in the first image, wherein a first motion descriptor describes the motion of the image sensor during the first time period; calculate descriptors for key points in the second image, wherein a second motion descriptor describes the motion of the image sensor during the second time period; convert, based on the first motion descriptor and the second motion descriptor, the calculated descriptors for the key points in the first image into converted descriptors using a trained artificial neural network; and compare the converted descriptors with the calculated descriptors for the key points in the second image.
[0007] One or more non-transitory computer storage media according to yet another general configuration store instructions that, when executed on a computer system, cause the computer system to perform the following operations: select key points in a first image, the first image being acquired by an image sensor during a first time period, wherein a first motion descriptor describes the motion of the image sensor during the first time period; calculate corresponding descriptors for the selected key points in the first image; select key points in a second image, the second image being acquired by an image sensor during a second time period, the second time period being different from the first time period, wherein a second motion descriptor describes the motion of the image sensor during the second time period; calculate corresponding descriptors for the selected key points in the second image; convert, based on the first motion descriptor and the second motion descriptor, the calculated descriptors for the selected key points in the first image into converted descriptors using a trained artificial neural network; and compare the converted descriptors with the calculated descriptors for the selected key points in the second image.
[0008] Compared with traditional techniques, many benefits can be obtained through the approach of the present disclosure. For example, the methods and systems according to embodiments of the present disclosure utilize a deep learning network as a key point descriptor motion blur converter or as a key point descriptor comparison module to address the challenge of image feature matching when there is different motion blur in camera frames. For example, embodiments of the present disclosure can be used to improve the positioning and mapping performance in AR / VR applications. In addition, when extracting features from camera frames captured under different image motion conditions, embodiments of the present disclosure can improve the feature matching accuracy. Therefore, when the device is in rapid motion, the SLAM calculation can be more accurate and stable. This is because one or more blurred camera frames that were previously wasted can now be effectively used for positioning and mapping estimation. Thus, embodiments of the present disclosure avoid the disadvantages of existing methods, such as the disadvantage of lower matching accuracy when there is slight motion or the image is blurred by similar motion. These and other embodiments of the present disclosure, along with many of their advantages and features, will be described in more detail in conjunction with the following text and drawings. Description of the Drawings
[0009] Figure 1 A simplified schematic diagram showing a key point descriptor converter according to an embodiment of the present disclosure.
[0010] Figure 2 A simplified flowchart showing an image processing method according to an embodiment of the present disclosure.
[0011] Figure 3A An example of an image according to an embodiment of the present disclosure.
[0012] Figure 3B Showing according to an embodiment of the present disclosure in Figure 3A Examples of key points of the image shown.
[0013] Figure 4 A simplified flowchart showing a method of performing image processing according to an embodiment of the present disclosure.
[0014] Figure 5 An example showing six degrees of freedom (6DOF).
[0015] Figure 6 A simplified flowchart showing a method of generating training data for network training according to an embodiment of the present disclosure.
[0016] Figure 7 A simplified flowchart showing a method of training an ANN according to an embodiment of the present disclosure.
[0017] Figure 8 A simplified block diagram showing a device according to an embodiment of the present disclosure.
[0018] Figure 9 A simplified flowchart according to an embodiment of the present disclosure is shown, which shows a method of performing feature matching using a key point descriptor comparison module.
[0019] Figure 10 A block diagram of a computer system according to an embodiment of the present disclosure is shown. Detailed Description of the Invention
[0020] Hereinafter, various embodiments will be described. For ease of explanation, specific configurations and details are listed to provide a comprehensive understanding of the embodiments. However, it will be apparent to those skilled in the art that these embodiments can be practiced without these specific details. In addition, well-known features may be omitted or simplified so as not to obscure the embodiments being described.
[0021] Many applications rely heavily on the performance of image feature matching. Such applications may include image alignment (e.g., image stitching, image registration, panoramic stitching), three-dimensional (3D) reconstruction (e.g., stereo vision), indexing and content retrieval, motion tracking, object recognition, and many other applications.
[0022] A basic requirement for many augmented reality or virtual reality (AR / VR) applications is to determine the position and orientation of the device in three-dimensional space. Such applications can use simultaneous localization and mapping (SLAM) algorithms to determine the real-time position and orientation of the device and infer the structure of the environment (or "scene") in which the device operates. In an example of a SLAM application, video sequence frames from a device camera are input into a module that executes the SLAM algorithm. Features are extracted from the frames and matched between different frames, and the SLAM algorithm searches for matching features corresponding to the same locations (spots) in the captured scene. By tracking the positions of the features in different frames, the SLAM module can determine the movement of the image sensor within the scene and infer the main structure of the scene.
[0023] In traditional feature matching, key point detection is performed on each of multiple images (e.g., on each frame of a video sequence), and a corresponding key point descriptor is calculated for each detected key point from its neighborhood (usually in the scale space). The number of key points detected in each image is typically at least several dozen and can be as many as five hundred or more. The neighborhood from which the key point descriptor is calculated usually has a radius of approximately fifteen pixels around the key point. As Figure 1As shown, the key point descriptors f0 from the first image I0 (e.g., a frame in a video sequence) and the key point descriptors f1 from the second image I1 (e.g., a different frame in the video sequence, e.g., consecutive frames in the sequence) are used to calculate a matching score. The score metric is typically the distance d(f0, f1) between the key point descriptors f0 and f1 in the descriptor space, e.g., the distance according to any of the following exemplary distance metrics. This score calculation is repeated for different pairs of key point descriptors from the two images, and the resulting scores are thresholded to identify matching features: e.g., to determine whether the current pair of key point descriptors matches and thus whether the corresponding features from the two images match. In a typical SLAM application, a pair of matching features corresponds to a point in the physical environment, and this correspondence represents a mathematical constraint. Subsequent stages of the SLAM calculation can derive the camera motion and the environment model as the best solution that satisfies multiple constraints. These multiple constraints include the constraints generated by the matching feature pairs.
[0024] Examples of distance metrics that can be used for matching score calculation include Euclidean distance, city block distance, chi-square distance, cosine distance, and Minkowski distance. Assume that f0 and f1 are n-dimensional vectors such that f0 = x 0,1 , x 0,2 , x 0,3 ,..., x 0,n , while f1 = x 1,1 , x 1,2 , x 1,3 ,..., x 1,n ,, then the distance d(f0, f1) between f0 and f1 can be determined according to these distance metrics in the following manner:
[0025] Euclidean distance
[0026] City block distance
[0027] Cosine distance Where
[0028] Chi-square distance (assuming all elements of f0 and f1 have values greater than zero)
[0029] Minkowski distance (also known as generalized Euclidean distance)
[0030] The resulting scores are thresholded to determine whether corresponding feature matching can be performed according to steps such as those described below:
[0031] If (d(f0, f1) < T) {
[0032] f0 and f1 match
[0033] } else {
[0034] f0 and f1 do not match
[0035] },
[0036] where T represents a threshold value.
[0037] Examples of AR / VR devices include mobile phones and head-mounted devices (e.g., AR or "smart" glasses). Given the nature of AR / VR devices, many video frames are captured while the image sensor (e.g., a video camera) is moving. Consequently, the captured frames may have significant motion blur. If the motion of the image sensor during the capture of image I0 is the same as its motion during the capture of image I1, and each descriptor f0 and f1 corresponds to the same key point in the two images, then the values of descriptors f0 and f1 are often similar, resulting in a small calculated distance between them. However, in practical AR / VR applications, the image sensor typically undergoes different motions when capturing each image, so descriptors f0 and f1 may be distorted by different amounts of motion blur. Unfortunately, almost all widely used image features (such as SIFT, SURF, BRISK) are very sensitive to motion blur, such that any significant motion blur (e.g., blur of five or more pixels) may cause distortion of the key point descriptors. Thus, even if descriptors f0 and f1 correspond to the same key point, the values of descriptors f0 and f1 may differ significantly. When extracting such features from images with different amounts of motion blur and using the above-mentioned score metric, feature matching is likely to fail.
[0038] For many applications where the image sensor may be in motion (e.g., AR / VR applications), it is possible to quantify the image motion during each capture interval. The direction and magnitude of the motion can be estimated from the input of one or more motion sensors of the device, and / or can be calculated, for example, from two temporally adjacent frames in a video sequence. Motion sensors (which may include one or more gyroscopes, accelerometers, and / or magnetometers) can indicate the displacement and / or change in orientation of the device and can be implemented within an inertial measurement unit (IMU).
[0039] Examples of techniques for dealing with motion blur can include the following:
[0040] 1) Do not perform matching: Since the matching of image features in blurred images is so unreliable, a potential solution is to completely abandon image feature matching, at least between image pairs with significant and different amounts of motion blur.
[0041] 2) First, perform deblurring on the image: Before the image is used for feature extraction, deblurring is performed to remove motion blur in the image.
[0042] 3) Compensate for motion blur when calculating keypoint descriptors: When calculating keypoint descriptors, the estimated motion of the image is used to compensate for the effect of motion blur. The difference between this method and "First, perform deblurring on the image" is that the removal or compensation of motion blur is performed on the neighborhood of the keypoints rather than on the entire image.
[0043] 4) Extract blur-invariant features: When calculating keypoint descriptors from the neighborhood of the keypoints, only the components of the blur invariants are used, while those components that are sensitive to motion blur are ignored. Therefore, even in the case of motion-blurred images, the keypoint descriptors will remain roughly the same.
[0044] The deficiencies of the above methods may include the following points:
[0045] 1) Mismatch: To prevent the quality of the SLAM estimate from degrading due to incorrect matching of features in blurred images, when significant image motion is detected, it may be chosen not to perform feature matching at all. For example, it may be chosen to perform SLAM using only the output of the motion sensor at these times. However, this method may cause the output of the image sensor to be completely wasted at these moments, while the SLAM calculation may become less accurate and less stable.
[0046] 2) First, perform deblurring on the image: Image deblurring usually involves a large amount of computation. Since SLAM calculations are usually performed on mobile platforms, the additional computation required for deblurring may not always be available or affordable. In addition, the deblurring operation is prone to adding new artifacts to the original image, which in turn is prone to negatively affecting the accuracy of image feature matching.
[0047] 3) Compensate for motion blur when calculating keypoint descriptors: Since the calculation only involves the neighborhood of the keypoints, compensating for motion blur when calculating keypoint descriptors usually requires less additional computation than deblurring the entire image. However, the drawback of introducing new artifacts in the image still exists.
[0048] 4) Extract blur-invariant features: Since this method ignores the components that are sensitive to motion blur, there is less information available for performing feature matching. In other words, this method can improve the matching accuracy and stability in cases of significant camera motion, but at the cost of reducing the matching performance in other cases (e.g., cases with little motion).
[0049] It may be desirable to improve the robustness of the keypoint descriptor framework to motion blur. Thus, embodiments described herein, implemented using suitable systems, methods, apparatuses, devices, etc. disclosed herein, can support improving the accuracy of feature matching operations in applications prone to motion blur. The embodiments described herein can be implemented in various applications using feature matching, including image alignment (e.g., image stitching, image registration, panoramic stitching), three-dimensional reconstruction (e.g., stereo vision), indexing and content retrieval, endoscopic imaging, motion tracking, object tracking, object recognition, autonomous navigation, SLAM, etc.
[0050] According to an embodiment of the present disclosure, a deep learning network is used as a keypoint descriptor motion blur converter or as a keypoint descriptor comparison module to address the challenges of image feature matching in the presence of different motion blurs in camera frames. For example, embodiments of the present disclosure can be used to improve the localization and mapping performance in AR / VR applications. Additionally, when extracting features from camera frames captured under different image motion conditions, embodiments of the present disclosure can improve the feature matching accuracy. Thus, when the device is in rapid motion, SLAM calculations can be more accurate and stable. This is because one or more blurred camera frames that were previously wasted can now be effectively used for localization and mapping estimation. Therefore, embodiments of the present disclosure avoid the drawbacks of existing methods, such as lower matching accuracy during small motions or when the image is blurred by similar motion.
[0051] Figure 1 A simplified schematic diagram of a keypoint descriptor converter 100 according to an embodiment of the present disclosure is shown. As Figure 1 shown, the conversion is performed on the keypoint descriptor f0 extracted from the image I0. The descriptor converter accepts three inputs: the keypoint descriptor f0; the motion descriptor M0 of the image motion at the time of acquiring the image I0; and the motion descriptor M1 of the image motion at the time of acquiring the image I1. The output of the converter is the transformed keypoint descriptor f1'. The design goal is that if the descriptors f0 and f1 correspond to the same keypoint, the transformed keypoint descriptor f1' is similar to the keypoint descriptor f1 in the image I1; and if the descriptors f0 and f1 refer to different keypoints, the transformed descriptor f1' and the descriptor f1 are significantly different.
[0052] Figure 2 A simplified flowchart of an image processing method according to an embodiment of the present disclosure is shown. Figure 1The image processing method 100 shown includes tasks 210, 220, 230, and 240. In task 210, descriptors of key points in a first image are calculated. The first image is acquired by an image sensor during a first time period. The first motion descriptor describes the motion of the image sensor during the first time period. In task 220, descriptors of key points in a second image are calculated. The second image is acquired by the image sensor during a second time period different from the first time period. The second motion descriptor describes the motion of the image sensor during the second time period. In task 230, using a trained artificial neural network (ANN), based on the first motion descriptor and the second motion descriptor, the calculated descriptors for the key points in the first image are converted into converted descriptors. In task 240, the converted descriptors are compared with the calculated descriptors for the key points in the second image. For example, task 240 may include calculating the distance (e.g., according to distance metrics disclosed herein, such as Euclidean distance, chi-square distance, etc.) between the converted descriptors and the calculated descriptors in the descriptor space, and comparing the calculated distance with a threshold (e.g., according to the procedure described above).
[0053] It should be understood that the specific steps shown in Figure 2 provide a specific method for performing image processing according to an embodiment of the present disclosure. As described above, other step sequences may also be performed according to alternative embodiments. For example, alternative embodiments of the present disclosure may perform the above steps in a different order. In addition, Figure 2 a single step shown in may include multiple sub-steps, and these sub-steps may be performed in different orders according to the situation of the single step. In addition, according to specific applications, additional steps may be added or steps may be deleted. Those skilled in the art will recognize many variations, modifications, and alternatives.
[0054] A key point is a point in an image. Different from other points in the image, a key point has a well-defined spatial position, or is otherwise located in the image and remains stable under local and global changes (e.g., scale change, illumination change, etc.).
[0055] Figure 3A An example of an image (e.g., a frame in a video sequence) according to an embodiment of the present disclosure is shown. Figure 3B An example of the key points of the image shown in according to an embodiment of the present disclosure is shown. Refer to Figure 3A in, Figure 3B Figure 3B The circles shown in indicate the positions of several examples of key points 310 - 322 in the image. In a typical feature matching application, the number of key points detected in each image is at least twelve, twenty-five, or fifty, and can reach one hundred, two hundred, five hundred, or more.
[0056] In a typical application of feature matching, for multiple different pairs of key-point descriptors (f0, f1), on paired consecutive frames in a video sequence (e.g., on each consecutive pair of frames), the method 200 shown in Figure 2 is repeated. In some embodiments, for each of the multiple key points detected in the first image (each key point having a corresponding location in the first image), the method 200 is repeated for each pair of descriptors. Each pair of descriptors includes a descriptor of a key point in the first image and a descriptor of a key point within a threshold distance (e.g., 20 pixels) of the same location in the second image.
[0057] The method 200 is not limited to any specific size or format of the first and second images. Examples from typical feature matching applications are now provided. Most current AR / VR devices are configured to acquire video in VGA format (i.e., having a frame size of 640 x 480 pixels), with each pixel having red, green, and blue components. The largest frame format common in such devices is 1280 x 720 pixels, such that in a typical application, the maximum size of each of the first and second images is approximately one thousand by two thousand pixels. In a typical application, the minimum size of each of the first and second images is approximately one-quarter VGA (i.e., 320 x 240 pixels), as smaller image sizes are likely to be insufficient to support algorithms such as SLAM.
[0058] The key-point descriptor calculation tasks 210 and 220 can be implemented using existing key-point descriptor frameworks. Existing key-point descriptor frameworks are, for example, Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), Binary Robust Invariant Scalable Keypoints (BRISK), etc. Such tasks can include calculating the orientation of key points. This can include determining how or in what direction the pixel neighborhood (also referred to as an "image patch") around the key point is oriented. The calculation of the orientation of key points is typically performed on image patches at different scales in the scale space. The calculation of the orientation of key points can include detecting the most dominant orientation of the gradient angles in the image patch. For example, the SIFT framework assigns a 128-dimensional feature vector to each key point based on the gradient directions of the pixels in 16 local neighborhoods of the key point.
[0059] Some keypoint descriptor frameworks (e.g., SIFT and SURF) include both keypoint detection and keypoint descriptor calculation. Other keypoint descriptor frameworks (e.g., Binary Robust Independent Elementary Features (BRIEF)) include keypoint descriptor calculation but do not include keypoint detection. For cases where the latter type of framework is used to perform Task 210 and Task 220, it may be desirable to perform a keypoint detection operation on the first image to select corresponding keypoints before performing Task 210, and to perform a keypoint detection operation on the second image to select corresponding keypoints before performing Task 220. In either case, in response to either the first motion descriptor or the second motion descriptor having a value that exceeds a threshold (e.g., blur of twenty pixels), it may be desirable to implement the methods herein to skip that pair of images. This is because keypoint detection is not likely to succeed for images with such extensive blur.
[0060] Figure 4 FIG. shows a simplified flowchart of a method for performing image processing according to another embodiment of the present disclosure. As Figure 4 shown, the elements used in Method 200 are also used in Figure 4 Method 400 shown in, and Method 400 further includes additional elements described below. Thus, the description related to Figure 2 also applies to Figure 4 . Those skilled in the art will recognize many variations, modifications, and alternatives.
[0061] As Figure 4 shown, in Task 410, keypoints in the first image are selected, and in Task 420, keypoints in the second image are selected. Examples of keypoints that can be used to implement Task 410 and Task 420 include corner detectors (e.g., Harris corner detector, Features from Accelerated Segment Test (FAST)) and blob detectors (e.g., Laplacian of Gaussian (LoG), Difference of Gaussians (DoG), Determinant of Hessian (DoH)). Such keypoint detectors are typically configured to blur the image with different blur widths and resample the image with different sampling rates to create a scale space, and to detect corners and / or blobs at different scales.
[0062] The first motion descriptor describes the motion of the image sensor when acquiring the first image, and the second motion descriptor describes the motion of the image sensor during the acquisition of the second image. The first motion descriptor and the second motion descriptor can be output from a motion sensor (e.g., an IMU) and / or calculated from adjacent frames in a video sequence (e.g., from frames acquired just before or after the frame for which the motion descriptor is being calculated). The acquisition period for each frame in a video sequence is typically the reciprocal of the frame rate, but the acquisition period may be shorter. The frame rate of a typical video sequence (e.g., taken by an Android phone) is 30 frames per second (fps). The frame rate of an iPhone or a head-mounted device can be as high as 120 fps.
[0063] Each motion descriptor can be implemented to describe a trajectory or path in a coordinate space of one, two, or three spatial dimensions. Such a trajectory can be described as a sequence of one or more positions of the image sensor sampled at uniform intervals during the corresponding acquisition period. Each sampled position can be represented as a motion vector relative to the previous sampled position (e.g., taking the position at the start of the acquisition period as the origin of the coordinate space). In one such example, the first motion descriptor describes a first trajectory having three dimensions in space, and the second motion descriptor describes a second trajectory having three dimensions in space, the second trajectory being different from the first trajectory.
[0064] Each motion descriptor can be further implemented to describe a six-degree-of-freedom (6DOF) motion. In addition to the three spatial dimensions, such a motion can also include rotations about one or more axes around these dimensions. As Figure 5 shown, Figure 5 a diagram showing six degrees of freedom (6DOF). These rotations can be labeled as roll, pitch, and yaw. For example, for each sampled position of the image sensor, the motion descriptor can include the corresponding orientation (e.g., the viewing direction) of the reference orientation of the image sensor relative to the orientation of the previous sampled position (e.g., taking the reference orientation as the orientation at the start of the acquisition period).
[0065] Referring again to Figure 4, in Task 230, using the trained artificial neural network (ANN), based on the first motion descriptor and the second motion descriptor, the computed descriptor for the key points in the first image is converted into a converted descriptor. Examples of ANNs that can be trained to perform such complex conversions between multi-element vectors include convolutional neural networks (CNNs) and autoencoders. It is desirable for the ANN to be implemented to be rather small and fast: for example, the ANN includes fewer than ten thousand parameters, or fewer than five thousand parameters, and / or the trained ANN occupies less than five megabytes of storage space. In some embodiments, the ANN is implemented such that both the input layer and the output layer are arrays of size 32×32. In a typical production environment, a copy of the trained ANN (e.g., during manufacturing and / or supply) is stored in each of a string of devices having the same model of video camera (and, possibly also having the same model of IMU). Before inputting the motion descriptor into the trained ANN, it is desirable to normalize the values of the motion descriptor so that they occupy the same range as the values of the computed key point descriptors.
[0066] It should be understood that the specific steps shown in Figure 4 provide a specific method for performing image processing according to another embodiment of the present disclosure. As described above, other step sequences can also be performed according to alternative embodiments. For example, alternative embodiments of the present disclosure can perform the above steps in a different order. In addition, Figure 4 the individual steps shown in
[0067] may include multiple sub-steps, and these sub-steps can be performed in different orders depending on the situation of the individual step. In addition, depending on the specific application, additional steps can be added or steps can be deleted. Those skilled in the art will recognize many variations, modifications, and alternatives.
[0068] Figure 6 shows a simplified flowchart according to an embodiment of the present disclosure, which shows a method for generating training data for network training. As Figure 6 shown, for example, for each image in the set of training images 620, some key points 622 can be detected. Then some different motion blurs (each described by a corresponding one in the set of motion descriptors 610a - 610n) can be applied to the image to generate a corresponding number of blurred images B1 - Bn. Key point descriptors for the detected key points (and corresponding motion descriptors) can be computed from each blurred image.
[0069] Figure 7A simplified flowchart according to an embodiment of the present disclosure is shown, which shows a method for training an ANN. As Figure 7 shown, during the training of the ANN, the corresponding calculated key point descriptors provide reference values or groundtruth for the loss function. It is desirable to enhance the synthetic training data from descriptors calculated from a relatively small number of images. The relatively small number of images have real motion blur and are manually annotated by human annotators.
[0070] Performing inference through a deep learning network typically involves more computations than calculating the distance between two key point descriptors. Therefore, to save some computations, it is desirable to implement Method 200 or Method 400 to compare the value of a first motion descriptor with the value of a second motion descriptor, and for cases where the comparison result indicates that the first motion descriptor and the second motion descriptor have similar values, avoid using the trained network (e.g., use a traditional scoring metric instead). Method 200 or Method 400 can be implemented, for example, to include a task. The task calculates the distance between the first motion descriptor and the second motion descriptor and compares the distance with a threshold. This implementation of Method 200 or Method 400 can be configured to, in response to the motion descriptor comparison task indicating that the motion blur of a first image is similar to the motion blur of a second image, use a scoring metric (e.g., the distance as described above) instead of using the trained network to determine whether the key point descriptor from the first image matches the key point descriptor from the second image.
[0071] Figure 8 A simplified block diagram of a device according to an embodiment of the present disclosure is shown. As an example, according to a general configuration including a key point descriptor calculator 810, a key point descriptor converter 820, and a key point descriptor comparator 830, the Figure 8 device 800 shown in can be used for image processing in a mobile device (e.g., a cellular phone such as a smart phone, or a head-mounted device).
[0072] The key point descriptor calculator 810 is configured to calculate the descriptors of the key points in the first image and calculate the descriptors of the key points in the second image (e.g., as described herein with reference to tasks 210 and 220, respectively). The first image is acquired by the image sensor during a first time period. The second image is acquired by the image sensor during a second time period different from the first time period. The key point descriptor converter 820 is configured to convert the calculated descriptor for the key points in the first image into a converted descriptor using a trained ANN based on the first motion descriptor and the second motion descriptor (e.g., as described herein with reference to task 230). The first motion descriptor describes the motion of the image sensor during the first time period. The second motion descriptor describes the motion of the image sensor during the second time period. The key point descriptor comparator 830 is configured to compare the converted descriptor with the calculated descriptor for the key points in the second image (e.g., as described herein with reference to task 240).
[0073] In some embodiments, the apparatus 800 is implemented in a device such as a mobile phone. The device typically has a video camera. The video camera is configured to generate a sequence of frames including the first image and the second image. The device may also include one or more motion sensors. The one or more motion sensors may be configured to determine the 6DOF motion of the device in space. In some embodiments, the apparatus 800 is implemented in a head-mounted device such as a set of AR glasses. The head-mounted device may also have motion sensors and one or more cameras.
[0074] Figure 9 A simplified flowchart according to an embodiment of the present disclosure is shown, which shows a method of performing feature matching using the key point descriptor comparison module 900. As shown in Figure 9 As shown, the key point descriptor comparison module 900 is used to determine whether two key point descriptors match, even if these features belong to two images with different image motions. Using the key point descriptor comparison module 900 is a more comprehensive method to solve the feature matching failure caused by motion blur. As Figure 9 shown, the key point descriptor comparison module 900 receives four inputs. The four inputs include two key point descriptors f0 and f1 and motion descriptors M0 and M1. The motion descriptors M0 and M1 correspond to the motion blur of the source images of f0 and f1. The output is a binary decision 910 (i.e., whether f0 and f1 match) and a value P 912 representing the confidence of the output binary decision.
[0075] In a manner similar to the reasons described above, a deep learning network is trained to produce the match indication and confidence value output of the keypoint descriptor comparison module 900. In this case, the network can be implemented as a classifier network known in the field of deep learning. Given sufficient training data, the classifier generally produces a good output. The training of such a network (e.g., a CNN) can be performed using the training data obtained as described above (e.g., referring to Figure 6 ).
[0076] Compared with the keypoint descriptor transformer 820, the keypoint descriptor comparison module 900 encapsulates score metrics and often has higher matching accuracy. On the other hand, the comparison module 900 requires more inputs than the transformer 820 and typically includes a larger network. Therefore, this solution tends to occupy more memory and consume more computing resources. As described above, when the keypoint descriptors f0 and f1 are from images with similar image motion, it may be desirable to use traditional score metrics instead to save some computational effort.
[0077] The embodiments discussed herein can be implemented in various fields. These fields may include feature matching, such as image alignment (e.g., panoramic stitching), 3D reconstruction (e.g., stereo vision), indexing, and content retrieval, etc. The first image and the second image are not limited to images generated by visible light cameras (e.g., in the RGB or other color spaces). For example, the first image and the second image can be images generated by cameras sensitive to non-visible light (e.g., infrared (IR) images, ultraviolet (UV) images), images generated by structured light cameras, and / or images generated by image sensors other than cameras (e.g., imaging using radar, lidar, sonar, etc.). In addition, the embodiments described herein can also be extended beyond motion blur to cover other factors that can distort keypoint descriptors, such as illumination changes, etc.
[0078] Figure 10 A block diagram of a computer system 1000 according to an embodiment of the present disclosure is shown. As described herein, the computer system 1000 and its components can be configured to perform the implementation of the methods described herein. Although these components are shown as belonging to the same computer system 1000 (e.g., a smartphone or a head-mounted device), the computer system 1000 can also be implemented such that these components are distributed (e.g., distributed among different servers, distributed between a smartphone and one or more network entities, etc.).
[0079] Computer system 1000 includes at least a processor 1002, a memory 1004, a storage device 1006, input / output peripherals or I / O 1008, a communication peripheral 1010, and an interface bus 1012. The interface bus 1012 is configured to communicate, transfer, and transmit data, control, and commands between the various components of computer system 1000. The memory 1004 and / or the storage device 1006 may be configured to store a first image and a second image (e.g., store video sequence frames), and may include computer-readable storage media such as RAM, ROM, electrically erasable programmable read-only memory (EEPROM), hard disk, CD-ROM, optical storage device, magnetic storage device, electronic non-volatile computer storage such as memory, and other tangible storage media. Any such computer-readable storage media may be configured to store instructions or program code embodying various aspects of the present disclosure. The memory 1004 and the storage device 1006 also include computer-readable signal media. The computer-readable signal media includes a propagated data signal embodying computer-readable program code. Such a propagated signal takes any of a variety of forms, including but not limited to electromagnetic, optical, or any combination thereof. The computer-readable signal media includes any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transmit a program for use in conjunction with computer system 1000.
[0080] In addition, the memory 1004 includes an operating system, programs, and applications. The processor 1002 is configured to execute the stored instructions, and the processor 1002 includes, for example, a logic processing unit, a microprocessor, a digital signal processor, and other processors. The memory 1004 and / or the processor 1002 may be virtual and may be hosted within another computer system such as a cloud network or a data center. The I / O peripherals 1008 include a user interface, other input / output devices (e.g., an image sensor configured to acquire an image to be indexed), and computing components. The user interface is such as a keyboard, a screen (e.g., a touch screen), a microphone, and a speaker. The computing components are such as a graphics processing unit, a serial port, a parallel port, a universal serial bus, and other input / output peripherals. The I / O peripherals 1008 are connected to the processor 1002 through any port coupled to the interface bus 1012. The communication peripheral 1010 is configured to facilitate communication between computer system 1000 and other computing devices (e.g., a cloud computing entity configured to execute a part of the indexing and / or query search methods described herein) through a communication network, and includes, for example, a network interface controller, a modem, wireless and wired interface cards, an antenna, and other communication peripherals.
[0081] Although the subject matter has been described in detail with respect to its specific embodiments, it should be understood that those skilled in the art can readily modify, vary, and make equivalents of these embodiments after understanding the above. Therefore, it should be understood that this disclosure is presented by way of example and not limitation, and does not exclude modifications, variations, and / or additions to the subject matter that are readily apparent to those skilled in the art. In fact, the methods and systems described herein can be embodied in various other forms. Additionally, various omissions, substitutions, and changes can be made to the forms of the methods and systems described herein without departing from the spirit of this disclosure. The appended claims and their equivalents are intended to cover such forms or modifications that fall within the scope and spirit of this disclosure.
[0082] Unless otherwise specifically stated, it should be understood that in this specification, discussions using terms such as "processing", "computing", "determining", "identifying", or similar terms refer to activities or processes of a computing device, such as one or more computers or similar electronic computing devices that manipulate or transform data represented as physical electronic or magnetic quantities in memory, registers, or other information storage devices, transmission devices, or display devices of a computing platform.
[0083] One or more of the systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device can include any suitable component configuration that provides results conditioned on one or more inputs. Suitable computing devices include computer systems based on general-purpose microprocessors. The computer system accesses the stored software. The software programs or configures the computer system from a general-purpose computing device into a specialized computing device for implementing one or more embodiments of the subject matter. Any suitable programming, scripting, or other type of language or combination of languages can be used to implement the teachings contained herein in the software for programming or configuring the computing device.
[0084] Embodiments of the methods disclosed herein can be executed in the operation of such computing devices. The order of the blocks presented in the above examples can be varied - for example, the blocks can be reordered, combined, and / or divided into sub-blocks. Certain blocks or processes can be executed in parallel.
[0085] The terms "comprising", "including", "having", etc. are synonyms and are used in an open-ended inclusive manner and do not exclude other elements, features, acts, operations, etc. Additionally, the term "or" is used in its inclusive sense (rather than in its exclusive sense). Thus, when "or" is used, for example, to connect a series of elements, the term "or" means one, some, or all of the elements in that series. The use of "adapted to" or "configured to" herein is meant as open and inclusive language and does not exclude devices that are adapted to or configured to perform additional tasks or steps. The headings, lists, and numbers included herein are for ease of explanation only and are not meant to be restrictive.
[0086] The various features and processes described above can be used independently of each other or in various combinations. All possible combinations and sub - combinations are intended to fall within the scope of this disclosure. Additionally, in some embodiments, certain method or process blocks may be omitted. The methods and processes described herein are also not limited to any particular order, and the associated blocks or states can be executed in other appropriate orders. For example, the described blocks or states can be executed in an order not specifically disclosed, or multiple blocks or states can be combined into a single block or state. These example blocks or states can be performed serially, in parallel, or otherwise. Blocks or states can be added to or removed from the disclosed examples. Similarly, the configurations of the example systems and components described herein can also be different. For example, elements can be added, removed, or rearranged compared to the disclosed examples.
[0087] The various elements of an embodiment of the apparatus or system disclosed herein (e.g., apparatus 800) can be embodied in any combination of hardware and software and / or firmware deemed suitable for the intended application. For example, these elements can be fabricated as electronic and / or optical devices, e.g., electronic and / or optical devices residing on the same chip or between two or more chips in a chipset. An example of such an element is a fixed or programmable logic element array, such as transistors or logic gates. Any one of these elements can be implemented as one or more such arrays. Any two or more, or even all, of these elements can be implemented in the same or multiple arrays. One or more such arrays can be implemented within one or more chips (e.g., within a chipset including two or more chips). Such a device can also be implemented to include a memory configured to store a first image and a second image.
[0088] The processors or other devices disclosed herein can be fabricated as one or more electronic and / or optical devices, e.g., one or more electronic and / or optical devices residing on the same chip or between two or more chips in a chipset. An example of such an element is an array of fixed or programmable logic elements, such as transistors or logic gates. Any of these elements can be implemented as one or more such arrays. One or more such arrays can be implemented within one or more chips (e.g., within a chipset including two or more chips). Examples of such arrays include arrays of fixed or programmable logic elements such as microprocessors, embedded processors, IP cores, DSPs (Digital Signal Processors), FPGAs (Field Programmable Gate Arrays), ASSPs (Application Specific Standard Products), and ASICs (Application Specific Integrated Circuits). The processors or other processing means disclosed herein can also be embodied as one or more computers (e.g., machines including one or more arrays programmed to execute one or more instruction sets or sequences of instructions) or other processors. The processors described herein can be used to perform tasks or execute other instruction sets. These tasks or instruction sets have no direct relation to the implementation procedures of method M100 or 400 (or other methods disclosed with reference to the operation of the devices or systems described herein). These tasks or instruction sets are, for example, tasks related to another operation of the device or system in which the processor is embedded (e.g., a voice communication device such as a smartphone or a smart speaker). It is also possible for a part of the methods disclosed herein to be executed under the control of one or more other processors.
[0089] Each task of the methods disclosed herein (e.g., method 200 and / or method 400) can be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. In typical applications of the implementation of the methods disclosed herein, an array of logic elements (e.g., logic gates) is configured to perform one, more than one, or even all of the various tasks of the method. One or more tasks (possibly all tasks) can also be implemented as code (e.g., one or more instruction sets). This code is embodied in a computer program product (e.g., one or more data storage media, such as disks, flash memory, or other non-volatile memory cards, semiconductor memory chips, etc.). This code can be read and / or executed by a machine (e.g., a computer) including an array of logic elements (e.g., a processor, a microprocessor, a microcontroller, or other finite state machine). The implementation tasks of the methods disclosed herein can also be performed by more than one such array or machine. In these or other embodiments, these tasks can be performed in a wireless communication device such as a cellular phone or other device having such communication capabilities. Such a device can be configured to communicate with a circuit-switched network and / or a packet-switched network (e.g., using one or more protocols such as VoIP). For example, such a device can include an RF circuit configured to receive and / or transmit encoded frames.
[0090] In one or more example embodiments, the operations described herein can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the operations can be stored on or transmitted via a computer-readable medium as one or more instructions or code. The term "computer-readable medium" includes both computer-readable storage media and communication (e.g., transmission) media. By way of example and not limitation, computer-readable storage media may include an array of storage elements such as semiconductor memories (which may include, but are not limited to, dynamic or static RAM, ROM, EEPROM, and / or flash memory), or ferroelectric memory, magnetoresistive memory, ovonic memory, polymeric memory, or phase change memory; CD-ROM or other optical disk storage; and / or magnetic disk storage or other magnetic storage devices. Such storage media may store information in the form of instructions or data structures that are accessible by a computer. Communication media may include any medium that is capable of carrying the desired program code in the form of instructions or data structures and that is accessible by a computer, including any medium that facilitates the transfer of a computer program from one place to another. Additionally, any connection is suitable to be termed a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and / or microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and / or microwave are included in the definition of the medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray DiscTM (Universal City, California, Blu-ray Disc Association). Disks typically reproduce data magnetically, while discs reproduce data optically by laser. Combinations of the above should also be included within the scope of computer-readable media.
[0091] In some embodiments, the non-transitory computer-readable storage medium includes code. When executed by at least one processor, the code causes the at least one processor to perform the image processing methods described herein (e.g., method 200 or method 400). Other examples of such storage media include media that also include code. When executed by at least one processor, the code causes the at least one processor to perform the image processing methods described herein.
[0092] Unless explicitly limited by the context, the term "signal" is used herein to denote any of its ordinary meanings, including the state of a memory location (or set of memory locations) expressed on a wire, bus, or other transmission medium. Unless explicitly limited by the context, the term "generate" is used herein to denote any of its ordinary meanings, such as calculate or otherwise produce. Unless explicitly limited by the context, the term "calculate" is used herein to denote any of its ordinary meanings, such as compute, evaluate, estimate, and / or select from a plurality of values. Unless explicitly limited by the context, the term "obtain" is used to indicate any of its ordinary meanings, such as calculate, derive, receive (e.g., from an external device), and / or retrieve (e.g., from an array of storage elements). Unless explicitly limited by the context, the term "select" is used to denote any of its ordinary meanings, such as identify, indicate, apply, and / or use at least one, but less than all, of a group formed of two or more. Unless explicitly limited by the context, the term "determine" is used to denote any of its ordinary meanings, such as decide, ascertain, conclude, calculate, select, and / or evaluate. When the term "comprising" is used in the specification and claims, other elements or operations are not excluded. The term "based on" (e.g., "A is based on B") is used to denote any of its ordinary meanings, including the following cases: case (i) "derived from" (e.g., "B is a precursor of A"); (ii) "at least based on" (e.g., "A is at least based on B") and (iii) "equal to" (e.g., "A is equal to B") if appropriate in a particular context. Similarly, the term "responsive to" is used to denote any of its ordinary meanings, including "at least responsive to". Unless otherwise specified, the terms "at least one of A, B, and C", "one or more of A, B, and C", and "one or more among A, B, and C" mean "A and / or B and / or C". Unless otherwise specified, "each of A, B, and C" and "each among A, B, and C" mean "A and B and C".
[0093] Unless otherwise specified, any disclosure of the operation of a device with specific features is also clearly intended to disclose a method with similar features, and vice versa, and any disclosure of the operation of a device according to a specific configuration is also clearly intended to disclose a method according to a similar configuration, and vice versa. The term "configuration" can be used to refer to a method, a device, and / or a system, as specifically indicated according to its particular context. Unless specifically stated otherwise in the context, the terms "method", "process", "procedure", and "technique" are general and can be used interchangeably. A "task" with multiple subtasks is also a method. Unless specifically stated otherwise in the context, the terms "device" and "equipment" are also general and interchangeable. The terms "element" and "module" are typically used to denote a part of a larger configuration. Unless explicitly restricted by the context, the term "system" is used herein to denote any ordinary meaning thereof, including "a set of elements that interact to achieve a common purpose".
[0094] Unless initially introduced by the definite article, ordinal numbers (such as "first", "second", "third", etc.) used to modify claim elements do not themselves indicate any priority or order of the claim element relative to another element, but merely distinguish the claim element from another claim element with the same name (but using an ordinal number). Unless explicitly restricted by the context, each of the terms "plurality" and "set" herein is used to denote an integer greater than 1.
[0095] The foregoing description is provided to enable a person skilled in the art to make or use the disclosed embodiments. Various modifications to these embodiments will be readily apparent to those skilled in the art. The principles defined herein can be applied to other embodiments without departing from the scope of the disclosure. Therefore, the disclosure is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features defined by the following claims.
Claims
1. An image processing method, comprising: Calculating descriptors of key points in a first image, the first image being acquired by an image sensor during a first time period, wherein a first motion descriptor describes the motion of the image sensor during the first time period; Calculating descriptors of key points in a second image, the second image being acquired by the image sensor during a second time period different from the first time period, wherein a second motion descriptor describes the motion of the image sensor during the second time period; Based on the first motion descriptor and the second motion descriptor, using a trained artificial neural network, converting the calculated descriptors for the key points in the first image into converted descriptors; and Comparing the converted descriptors with the calculated descriptors for the key points in the second image.
2. The method according to claim 1, wherein The first motion descriptor describes a first trajectory having three dimensions in space; and The second motion descriptor describes a second trajectory having three dimensions in space, the second trajectory being different from the first trajectory.
3. The method according to claim 1, characterized in that, Each of the first motion descriptor and the second motion descriptor describes a motion of six degrees of freedom.
4. The method according to claim 1, wherein The calculated descriptors for the key points in the first image describe neighborhoods of the key points in the first image; And The calculated descriptors for the key points in the second image describe neighborhoods of the key points in the second image.
5. The method according to claim 4, wherein The motion described by the first motion descriptor is based on information of an image acquired by the image sensor, the image not being the first image.
6. The method according to claim 4, characterized in that, Further comprising: Determining that a distance between the first motion descriptor and the second motion descriptor is not less than a threshold, wherein, depending on the determination, the trained artificial neural network is used.
7. The method according to claim 1, characterized in that Further comprising: Selecting key points in the first image, and selecting key points in the second image.
8. A computer system, comprising: One or more memories configured to store: A first image acquired by an image sensor during a first time period, and A second image acquired by the image sensor during a second time period different from the first time period; And One or more processors; wherein, The one or more memories are further configured to store computer-readable instructions that, when executed by the one or more processors, configure the computer system to: Calculate descriptors for key points in the first image, wherein a first motion descriptor describes the motion of the image sensor during the first time period; Calculate descriptors for key points in the second image, wherein a second motion descriptor describes the motion of the image sensor during the second time period; Based on the first motion descriptor and the second motion descriptor, using a trained artificial neural network, converting the calculated descriptors for the key points in the first image into converted descriptors; and Compare the transformed descriptor with the computed descriptor for the key points in the second image.
9. The computer system according to claim 8, wherein the first motion descriptor describes a first trajectory having three dimensions in space; and the second motion descriptor describes a second trajectory having three dimensions in space, the second trajectory being different from the first trajectory.
10. The computer system according to claim 8, wherein Each of the first motion descriptor and the second motion descriptor describes a motion of six degrees of freedom.
11. The computer system according to claim 8, wherein the computed descriptor for the key points in the first image describes the neighborhood of the key points in the first image; and the computed descriptor for the key points in the second image describes the neighborhood of the key points in the second image.
12. The computer system according to claim 8, wherein The motion described by the first motion descriptor is based on the information of the image acquired by the image sensor, and the image is not the first image.
13. The computer system according to claim 8, wherein The computer-readable instructions are further operable to configure the computer system to: determine that the distance between the first motion descriptor and the second motion descriptor is not less than a threshold, and depending on the determination, use the trained artificial neural network.
14. The computer system according to claim 8, characterized in that, Further comprising: selecting key points in the first image and selecting key points in the second image.
15. One or more non-transitory computer storage media storing instructions that, when executed on a computer system, cause the computer system to perform the following operations: Select key points in a first image, the first image being acquired by an image sensor during a first time period, wherein, The first motion descriptor describes the motion of the image sensor during the first time period; compute the corresponding descriptors for the selected key points in the first image; select key points in a second image, the second image being acquired by the image sensor during a second time period different from the first time period, wherein the second motion descriptor describes the motion of the image sensor during the second time period; compute the corresponding descriptors for the selected key points in the second image; based on the first motion descriptor and the second motion descriptor, use the trained artificial neural network to transform the computed descriptor for the selected key points in the first image into a transformed descriptor; and compare the transformed descriptor with the computed descriptor for the selected key points in the second image.
16. The one or more non-transitory computer storage media storing instructions according to claim 15, wherein: the first motion descriptor describes a first trajectory having three dimensions in space; and the second motion descriptor describes a second trajectory having three dimensions in space, the second trajectory being different from the first trajectory.
17. The one or more non-transitory computer storage media storing instructions according to claim 15, wherein, Each of the first motion descriptor and the second motion descriptor describes a motion of six degrees of freedom.
18. The one or more non-transitory computer storage media storing instructions according to claim 15, wherein: the computed descriptor for the key points in the first image describes the neighborhood of the key points in the first image; and The computed descriptors for the key points in the second image describe the neighborhoods of the key points in the second image.
19. The one or more non-transitory computer storage media storing instructions according to claim 18, wherein, The motion described by the first motion descriptor is based on information from an image acquired by the image sensor, which is not the first image.
20. The one or more non-transitory computer storage media storing instructions according to claim 18, wherein, The instructions further cause the computer system to perform an operation including determining that a distance between the first motion descriptor and the second motion descriptor is not less than a threshold, and, depending on the determination, using the trained artificial neural network.
Citation Information
Patent Citations
Device and method for detecting moving objects
CN103733607A
Image processing device and method of corrected image
CN110049230A