Image data processing method, image processing device and readable storage medium

By judging and eliminating dynamic objects in the SLAM algorithm, using the contour coordinates and pixel center of mass of the camera picture frame, combined with inertial measurement unit data, the positioning accuracy problem of the SLAM algorithm in the existence of dynamic objects is solved, and higher positioning accuracy and calculation efficiency are achieved.

CN115272417BActive Publication Date: 2025-08-19GEER TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211000997.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-19
Publication Date
2025-08-19
Estimated Expiration
2042-08-19

AI Technical Summary

Technical Problem

The existing SLAM algorithm is difficult to accurately process when there are dynamic objects, resulting in a reduced positioning accuracy in VIO tracking mode.

Method used

By obtaining the picture frames collected by the camera, determining the outline coordinates and center of mass of the object, combining the data of the inertial measurement unit, dynamic objects are judged and eliminated, improving positioning accuracy and reducing calculation amount.

Benefits of technology

Effectively eliminates dynamic object interference, improves the positioning accuracy of the SLAM algorithm and reduces the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115272417B_ABST
    Figure CN115272417B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of digital image processing technology, and in particular to a method for processing image data, an image processing device, and a readable storage medium, wherein the method comprises: obtaining a first frame and a second frame captured by a camera; determining the contour coordinates of each object in the first frame, and determining the pixel centroid of each object between the first frame and the second frame; and determining a moving object among the objects based on the contour coordinates and / or the pixel centroid. By determining the contour coordinates and pixel centroid of the object in the image, and judging whether the object is a dynamic object based on the contour coordinates and / or the pixel centroid, the interference of the dynamic object in the image on the subsequent algorithm is eliminated, the positioning accuracy is improved, and the computational complexity of the subsequent algorithm is reduced, thus solving the problem of how to judge and eliminate dynamic objects in static images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital image processing, and in particular to an image data processing method, an image processing device, and a readable storage medium. Background Art

[0002] In current SLAM (Simultaneous Localization and Mapping) algorithms for VR (Virtual Reality) and AR (Augmented Reality), the VIO (Visual-Inertial Odometry) tracking model, which combines a camera and an IMU (Inertial Measurement Unit), is a common implementation. The pose positioning accuracy and speed of the VIO tracking model are key competitive advantages for AR and VR products.

[0003] Traditional SLAM algorithms in related technical solutions are based on the assumption that the entire scene is static and there are no dynamic objects. However, when there are significant dynamic objects in the scene, traditional SLAM algorithms have difficulty processing dynamic objects, resulting in reduced positioning accuracy in VIO tracking mode.

[0004] The above content is only used to assist in understanding the technical solution of the present invention and does not constitute an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of the present invention is to provide a method for processing image data, aiming to solve the problem of how to determine and eliminate dynamic objects in static images.

[0006] To achieve the above-mentioned object, the present invention provides a method for processing image data, the method comprising:

[0007] Obtain the first picture frame and the second picture frame captured by the camera;

[0008] Determining the outline coordinates of each object in the first picture frame, and determining the pixel centroid of each object between the first picture frame and the second picture frame;

[0009] A moving object among the objects is determined according to the contour coordinates and / or the pixel centroids.

[0010] Optionally, the step of determining a moving object among the objects according to the contour coordinates and / or the pixel centroids includes:

[0011] determining a deflection vector of the camera within a corresponding time difference between the first picture frame and the second picture frame;

[0012] determining a field of view angle difference of the camera between the first picture frame and the second picture frame according to the deflection vector;

[0013] Determining a displacement of the pixel centroid of each object between the first picture frame and the second picture frame based on the field of view angle difference;

[0014] determining the moving object among the objects according to the mass center displacement, wherein when the mass center displacement is greater than a displacement threshold, determining that the object is the moving object;

[0015] Alternatively, determining a contour model of the object according to the contour coordinates;

[0016] The moving object among the objects is determined based on the matching degree between the contour model and a reference contour model in a preset dynamic training set, wherein when the matching degree is greater than a matching threshold, the object is determined to be the moving object.

[0017] Optionally, before the step of determining the moving object among the objects based on the contour coordinates and / or the pixel centroids, the method further comprises:

[0018] determining a picture complexity between the first picture frame and the second picture frame;

[0019] When the picture complexity is greater than or equal to a complexity threshold, performing the step of determining the deflection vector of the camera within the corresponding time difference between the first picture frame and the second picture frame;

[0020] Otherwise, the step of determining the contour model of the target object according to the contour coordinates is performed.

[0021] Optionally, the step of determining the deflection vector of the camera within the corresponding time difference between the first picture frame and the second picture frame includes:

[0022] Obtain inertial data from the inertial measurement unit;

[0023] The yaw vector is determined based on the inertial data.

[0024] Optionally, before the step of determining the difference in field of view angle of the camera between the first picture frame and the second picture frame according to the deflection vector, the step includes:

[0025] Acquire a visual coordinate system corresponding to the camera and a picture coordinate system corresponding to the first picture frame and the second picture frame;

[0026] Determining a coordinate mapping between the camera and the first picture frame and the second picture frame according to the visual coordinate system and the picture coordinate system;

[0027] The step of determining the difference in field of view angle of the camera between the first picture frame and the second picture frame according to the deflection vector includes:

[0028] determining, based on the coordinate mapping, a picture translation amount and a picture deflection amount between the first picture frame and the second picture frame according to the deflection vector;

[0029] The field of view angle difference is determined according to the picture translation amount and the picture deflection amount.

[0030] Optionally, the step of determining a moving object among the objects according to the contour coordinates and / or the pixel centroids includes:

[0031] Determining a contour model of each object according to the contour coordinates;

[0032] determining a first moving object among the objects according to a first matching degree between the contour model and a reference contour model in a preset dynamic training set;

[0033] determining a first deflection vector of the camera within a corresponding time difference between the first picture frame and the second picture frame;

[0034] determining a first field of view angle difference of the camera between the first image frame and the second image frame according to the first deflection vector;

[0035] determining, based on the first field of view angle difference, first centroid displacements of first pixel centroids of the objects other than the first moving object among the objects between the first picture frame and the second picture frame;

[0036] determining a second moving object among the other objects according to the first mass center displacement;

[0037] The moving object is determined according to the first moving object and the second moving object.

[0038] Optionally, the step of determining whether the object is a moving object based on the contour coordinates and / or the pixel centroid includes:

[0039] determining a second deflection vector of the camera within a corresponding time difference between the first picture frame and the second picture frame;

[0040] determining a second field of view angle difference of the camera between the first picture frame and the second picture frame according to the second deflection vector;

[0041] determining, based on the second field of view angle difference, a second centroid displacement of a second pixel centroid of each object between the first picture frame and the second picture frame;

[0042] determining a third moving object among the objects according to the second mass center displacement;

[0043] determining, according to the contour coordinates of the third moving object, contour models of other objects among the objects except the third moving object;

[0044] determining a fourth moving object among the other objects based on a second matching degree between the contour model of the other object and a reference contour model in a preset dynamic training set;

[0045] The moving object is determined according to the third moving object and the fourth moving object.

[0046] Optionally, after the step of determining whether the object is a moving object based on the contour coordinates and / or the pixel centroid, the following steps are included:

[0047] When it is determined that the object is the dynamic object, determining a coordinate area associated with the moving object in the second picture frame as the dynamic object coordinate area;

[0048] using the areas other than the dynamic object coordinate area in the first picture frame and the second picture frame as static object coordinate areas;

[0049] Determining static feature points in the static object coordinate area;

[0050] The static feature points are input into a target algorithm to obtain an image in which the dynamic objects are removed.

[0051] In addition, to achieve the above-mentioned purpose, the present invention also provides an image processing device, which includes a memory, a processor, and an image data processing program stored in the memory and runnable on the processor. When the image data processing is executed by the processor, the steps of the image data processing method described above are implemented.

[0052] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a processing program for image data is stored. When the processing program for image data is executed by a processor, the steps of the image data processing method described above are implemented.

[0053] An embodiment of the present invention provides an image data processing method, an image processing device, and a readable storage medium, wherein the method includes: obtaining a first frame and a second frame captured by a camera; determining the contour coordinates of each object in the first frame, and determining the pixel centroid of each object between the first frame and the second frame; and determining a moving object among the objects based on the contour coordinates and / or the pixel centroid. By determining the contour coordinates and pixel centroid of the object in the image, and judging whether the object is a dynamic object based on the contour coordinates and / or the pixel centroid, the interference of the dynamic object in the image on the algorithm is eliminated in the subsequent algorithm, thereby improving the positioning accuracy while reducing the computational complexity of the subsequent algorithm, and solving the problem of how to judge and eliminate dynamic objects in static images. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 Schematic diagram of the hardware architecture of an image processing device involved in an embodiment of the present invention;

[0055] Figure 2 1 is a flow chart of a first embodiment of a method for processing image data according to the present invention;

[0056] Figure 3 is a schematic diagram of a marking object in a first embodiment of the method for processing image data of the present invention;

[0057] Figure 4 is a flow chart of a second embodiment of the method for processing image data of the present invention;

[0058] Figure 5 This is a schematic diagram of a picture captured by a camera in the second embodiment of the image data processing method of the present invention;

[0059] Figure 6 Schematic diagram of pixel value distribution within a 3*3 sliding frame in a second embodiment of the image data processing method of the present invention;

[0060] Figure 7 Schematic diagram of feature point marking in the second embodiment of the method for processing image data of the present invention;

[0061] Figure 8 A schematic diagram of a dynamic object coordinate area marking in a second embodiment of the image data processing method of the present invention;

[0062] Figure 9 This is a schematic diagram of extraction results in the second embodiment of the method for processing image data of the present invention;

[0063] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0064] The present application relates to a method of identifying and capturing moving objects in vision, and reducing the impact of dynamic images on feature point matching errors by cutting out images, thereby increasing the pose data accuracy of the SLAM algorithm. The recognition process can be AI recognition of action posture, the relative motion between the movement of the IMU and the Camera image, etc.; cutting out images mainly involves the characterization of the contours of moving objects, and the deletion of feature points through sparse pixel values; the input of the SLAM algorithm is a stationary object, which reduces the algorithm logic of mismatching and elimination of feature points, thereby improving the positioning rate and accuracy. In addition, due to the addition of the inertial measurement unit (IMU), it can ensure that the camera can accurately capture dynamic objects in the image while moving.

[0065] To better understand the above technical solutions, exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0066] As an implementation solution, the image processing device can be as follows Figure 1 shown.

[0067] The embodiment of the present invention relates to an image processing device, which includes a processor 101, such as a CPU, a memory 102, and a communication bus 103. The communication bus 103 is used to achieve connection and communication between these components.

[0068] The memory 102 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Figure 1 As shown, the memory 102 as a computer-readable storage medium may include an image data processing program; and the processor 101 may be used to call the image data processing program stored in the memory 102 and perform the following operations:

[0069] Obtain the first picture frame and the second picture frame captured by the camera;

[0070] Determining the outline coordinates of each object in the first picture frame, and determining the pixel centroid of each object between the first picture frame and the second picture frame;

[0071] A moving object among the objects is determined according to the contour coordinates and / or the pixel centroids.

[0072] In one embodiment, the processor 101 may be configured to call an image data processing program stored in the memory 102 and perform the following operations:

[0073] determining a deflection vector of the camera within a corresponding time difference between the first picture frame and the second picture frame;

[0074] determining a field of view angle difference of the camera between the first picture frame and the second picture frame according to the deflection vector;

[0075] Determining a displacement of the pixel centroid of each object between the first picture frame and the second picture frame based on the field of view angle difference;

[0076] determining the moving object among the objects according to the mass center displacement, wherein when the mass center displacement is greater than a displacement threshold, determining that the object is the moving object;

[0077] Alternatively, determining a contour model of the object according to the contour coordinates;

[0078] The moving object among the objects is determined based on the matching degree between the contour model and a reference contour model in a preset dynamic training set, wherein when the matching degree is greater than a matching threshold, the object is determined to be the moving object.

[0079] In one embodiment, the processor 101 may be configured to call an image data processing program stored in the memory 102 and perform the following operations:

[0080] determining a picture complexity between the first picture frame and the second picture frame;

[0081] When the picture complexity is greater than or equal to a complexity threshold, performing the step of determining the deflection vector of the camera within the corresponding time difference between the first picture frame and the second picture frame;

[0082] Otherwise, the step of determining the contour model of the target object according to the contour coordinates is performed.

[0083] In one embodiment, the processor 101 may be configured to call an image data processing program stored in the memory 102 and perform the following operations:

[0084] Obtain inertial data from the inertial measurement unit;

[0085] The yaw vector is determined based on the inertial data.

[0086] In one embodiment, the processor 101 may be configured to call an image data processing program stored in the memory 102 and perform the following operations:

[0087] Acquire a visual coordinate system corresponding to the camera and a picture coordinate system corresponding to the first picture frame and the second picture frame;

[0088] Determining a coordinate mapping between the camera and the first picture frame and the second picture frame according to the visual coordinate system and the picture coordinate system;

[0089] The step of determining the difference in field of view angle of the camera between the first picture frame and the second picture frame according to the deflection vector includes:

[0090] determining, based on the coordinate mapping, a picture translation amount and a picture deflection amount between the first picture frame and the second picture frame according to the deflection vector;

[0091] The field of view angle difference is determined according to the picture translation amount and the picture deflection amount.

[0092] In one embodiment, the processor 101 may be configured to call an image data processing program stored in the memory 102 and perform the following operations:

[0093] Determining a contour model of each object according to the contour coordinates;

[0094] determining a first moving object among the objects according to a first matching degree between the contour model and a reference contour model in a preset dynamic training set;

[0095] determining a first deflection vector of the camera within a corresponding time difference between the first picture frame and the second picture frame;

[0096] determining a first field of view angle difference of the camera between the first image frame and the second image frame according to the first deflection vector;

[0097] determining, based on the first field of view angle difference, first centroid displacements of first pixel centroids of the objects other than the first moving object among the objects between the first picture frame and the second picture frame;

[0098] determining a second moving object among the other objects according to the first mass center displacement;

[0099] The moving object is determined according to the first moving object and the second moving object.

[0100] In one embodiment, the processor 101 may be configured to call an image data processing program stored in the memory 102 and perform the following operations:

[0101] determining a second deflection vector of the camera within a corresponding time difference between the first picture frame and the second picture frame;

[0102] determining a second field of view angle difference of the camera between the first picture frame and the second picture frame according to the second deflection vector;

[0103] determining, based on the second field of view angle difference, a second centroid displacement of a second pixel centroid of each object between the first picture frame and the second picture frame;

[0104] determining a third moving object among the objects according to the second mass center displacement;

[0105] determining, according to the contour coordinates of the third moving object, contour models of other objects among the objects except the third moving object;

[0106] determining a fourth moving object among the other objects based on a second matching degree between the contour model of the other object and a reference contour model in a preset dynamic training set;

[0107] The moving object is determined according to the third moving object and the fourth moving object.

[0108] In one embodiment, the processor 101 may be configured to call an image data processing program stored in the memory 102 and perform the following operations:

[0109] When it is determined that the object is the dynamic object, determining a coordinate area associated with the moving object in the second picture frame as the dynamic object coordinate area;

[0110] using the areas other than the dynamic object coordinate area in the first picture frame and the second picture frame as static object coordinate areas;

[0111] Determining static feature points in the static object coordinate area;

[0112] The static feature points are input into a target algorithm to obtain an image in which the dynamic objects are removed.

[0113] Based on the hardware architecture of the above-mentioned image processing device based on digital image processing technology, an embodiment of the method for processing image data of the present invention is proposed.

[0114] In the current AR and VR SLAM algorithms, VIO (binocular + IMU) tracking mode is a relatively common implementation method, and the accuracy and speed of POSE (pose) are the competitive points of current AR, VR and other products. Among them, the camera is the dominant logic in the algorithm, and improving its accuracy and speed is the key to improving the entire algorithm; and the image processing as the original input generally extracts feature points in the image, and triangulates the current depth information through the common view relationship between the two eyes or the matching relationship between the previous and next frames. Therefore, feature points are the key input, and the accuracy of the input determines the complexity of subsequent algorithm calculations.

[0115] In binocular scenes, the main input is the feature point extraction of static objects. Dynamic objects are noise in this process. Therefore, these dynamically moving objects need to be removed from the frame image to ensure that the retained information is static and stable.

[0116] Reference Figure 2 In a first embodiment, the method for processing image data includes the following steps:

[0117] Step S10, obtaining a first picture frame and a second picture frame captured by a camera;

[0118] In this embodiment, a first picture frame and a second picture frame captured by a camera are obtained;

[0119] In this embodiment, the first picture frame and the second picture frame captured by the camera are firstly acquired.

[0120] Optionally, the camera may be a binocular camera or a monocular camera. If the camera is a binocular camera, depth of field information can be determined based on images captured by the binocular camera. If the camera is a monocular camera, depth of field information of images captured by the monocular camera can also be determined based on other devices or methods. This allows the distance of imaged objects in the image to be distinguished based on the depth of field information.

[0121] Optionally, in some embodiments, taking a binocular camera as an example, the principle of capturing picture frames is as follows: first, the binocular camera is calibrated to obtain the internal and external parameters and homography matrix of the left / right cameras; the original image is corrected according to the calibration results, and the two corrected images captured at the same time are located in the same plane and parallel to each other; then binocular matching is performed to match pixels of the two corrected images; the depth of each pixel is calculated based on the matching results, thereby obtaining a picture frame containing the imaging distance of each object in the image.

[0122] For a single frame, it's impossible to determine whether each object within the image is dynamic or has a dynamic trend. Therefore, it's necessary to compare objects between two adjacent frames to make this determination. It's important to note that the first frame is the previous frame of the second frame, and the sampling interval between the first and second frames is set by the developer based on actual needs.

[0123] Step S20, determining the outline coordinates of each object in the first picture frame, and determining the pixel centroid of each object between the first picture frame and the second picture frame;

[0124] In this embodiment, each object in the picture is identified based on the difference in pixel mean values between various regions in the picture frame, and the outline of each object in the picture is marked.

[0125] Optionally, the outline marking method can be polygonal outline marking. Figure 3 , Figure 3 The following is a schematic diagram of marking a target object identified in an image frame with a parallelogram in a specific embodiment. After marking, the pixel information of the polygon corner points is recorded. In subsequent processing, the target object can be eliminated based on the recorded pixel information.

[0126] It should be noted that each identified object is present in both the first and second frames. In some embodiments, when a meaningful object appears only in the previous frame and disappears in the next frame (for example, a bird flying by quickly), the object is directly determined to be a dynamic object and no further determination is made. Of course, in most implementation scenarios, the time interval between the capture of the two frames is usually short, and in most cases the target object is present in both the first and second frames.

[0127] Specifically, the outline coordinates of each object marked in the first frame are determined. The outline coordinates are a corresponding coordinate group of the circumscribed polygon outlined based on the object model. The pattern shape formed based on the coordinate group can reflect the appearance model of the object.

[0128] Specifically, the pixel centroid of the objects that appear between the first picture frame and the second picture frame is determined. The pixel centroid is the pixel point centroid of the image area corresponding to the target object, which is characterized by the fact that the sum of pixels in the horizontal and vertical directions of the image at this point is equal. The pixel centroid is used as an object identifier in subsequent calculations to determine whether the image area (i.e., the object) corresponding to the pixel centroid has moved.

[0129] Step S30: determining a moving object among the objects according to the contour coordinates and / or the pixel centroids.

[0130] After determining the outline coordinates of each object in the first frame image and the pixel centroids of each object between adjacent image frames, the moving object among the objects is determined based on the outline coordinates and / or the pixel centroids.

[0131] Optionally, in some embodiments, the more mobile objects with a motion trend among the various objects can be determined based on the contour coordinates. The contour coordinates are input into a trained AI dynamic training model. The AI dynamic training model is based on machine deep learning technology. Models of more mobile objects that may move in daily scenes, such as models of cats, dogs, birds, vehicles, people and other objects, are input as dynamic training sets to train a model that can determine whether the object is an object that may move based on the contour coordinates. The moving objects among the various objects are determined based on the degree of match between the contour model and the reference contour model in the preset dynamic training set. Optionally, a matching threshold is set. When the matching degree is greater than the matching threshold, the object is judged to be a moving object. This recognition method is fast and effective in scenes with relatively simple images.

[0132] Optionally, in some embodiments, moving objects can be identified from the pixel centroids. First, the camera's change between the first and second frames is determined: the camera's deflection vector is determined within the corresponding time difference between the first and second frames. The deflection vector includes the offset distance and the deflection angle. Based on the deflection vector, the angular difference between the field of view angles of the two frames is obtained. This recognition method is more accurate in scenes with complex images.

[0133] In some specific embodiments, an inertial measurement unit (IMU) can be used to determine the object's sheet vector. An IMU is characterized by its relatively accurate measurement of the wearer's position change over a short period of time. However, over longer periods of measurement, errors accumulate over time. That is, the longer the measurement interval, the lower the accuracy of the measurement. In this embodiment, the time difference between the first and second frames is relatively short, so an IMU is introduced to measure the change in the camera's field of view angle, resulting in a more accurate final detection result. Inertial data from the inertial measurement unit is obtained, and the change in the camera's field of view angle during the corresponding time difference between the first and second frames is calculated based on the inertial data. In some specific embodiments, before determining the field of view angle difference, it is also necessary to obtain the camera's corresponding visual coordinate system and the corresponding image coordinate system between the first and second frames, and determine a coordinate mapping between the two coordinate systems. The purpose of determining the coordinate mapping is to determine the relative positional relationship between the object's actual movement distance under the camera lens and the object's movement distance between the first and second frames. For example, an object 3 meters away from the user actually moves 1 meter in space and 0.5 centimeters in the image. Therefore, the coordinate mapping parameter associated with the 3-meter distance is determined to be 200. Based on this coordinate mapping, the camera's deflection vector is converted into the image translation and deflection between the first and second frames. Based on the image translation and deflection, the image change of the second frame relative to the first frame, i.e., the field of view angle difference, can be determined.

[0134] Optionally, in the above two embodiments, the method to be used for judgment can be determined based on the picture complexity between the first picture frame and the second picture frame. Picture complexity is characterized by a quantitative parameter reflecting the complexity of the picture content, which is jointly defined by the structure of the object model in the picture, the number of objects, and the difference in picture pixel values. Among them, the more polygons the object model has, the more objects there are, and the greater the difference between adjacent pixel values in the picture, the higher the corresponding picture complexity. When the picture complexity is greater than or equal to the preset complexity threshold, it is judged that the scene in the picture is relatively complex, and the solution of determining the moving objects among the various objects based on the pixel centroid is executed; otherwise, the solution of determining the moving objects among the various objects based on the pixel centroid is executed.

[0135] Furthermore, if there is an angle difference between the field of view angles of the two picture frames, the field of view angle of the latter frame is converted to the field of view angle corresponding to the previous frame based on the field of view angle difference, that is, the adjacent image frames obtained after the conversion have the same field of view angle. After the field of view angles of the two frames are made consistent, the movement of the pixel center of mass of the target object between the two frames, that is, the center of mass movement, is compared, and whether the target object is a dynamic object is determined based on the center of mass movement. Optionally, in some embodiments, a movement threshold can be set, and when the center of mass movement is greater than the movement threshold, the object is determined to be a dynamic object. Optionally, in other embodiments, when the coordinates of the pixel center of mass of the same object between the first picture frame and the second picture frame are inconsistent, the object is determined to be a dynamic object.

[0136] Alternatively, in some embodiments, the object can be determined to be dynamic based on both the contour coordinates and the pixel centroid. It should be noted that the two different determination methods have different execution orders, resulting in different technical effects. The following describes the different execution orders:

[0137] 1. First, identify the object model in the first frame according to the contour coordinates to obtain the contour model of each object, and then determine the first matching degree between the contour model and the reference contour model. When the first matching degree is greater than the threshold, the object is determined to be the first moving object, that is, the AI model is used to quickly identify the moving object in the picture that may move in the subsequent picture. Then, based on the pixel centroid, other objects other than the first moving object between the first picture frame and the second picture frame are judged to obtain the second moving object, that is, the objects in the first picture frame and the second picture frame are further judged as dynamic objects through pixel centroid judgment, and the first moving object and the second moving object are both regarded as moving objects. This method is suitable for eliminating objects that are not recognized by the AI model when they move in the picture.

[0138] For example, when the device recognizes that there is a piece of paper in the picture captured by the camera, for AI model recognition, the detection result of the model recognition corresponding to the outline coordinates of the paper should be stationary. However, when the paper moves in the next frame due to interference from external factors (such as being blown by the wind), the paper should also be removed as a dynamic object that interferes with the positioning of the subsequent algorithm, and the AI model recognition solution cannot recognize that the paper is a dynamic object. Therefore, in further pixel centroid recognition, by detecting that the pixel centroid corresponding to the paper has changed, the paper can be identified as a dynamic object through pixel centroid recognition.

[0139] Second, first, based on the pixel centroid, identify all objects other than the first moving object between the first and second frames to obtain a third moving object. Then, using the AI model, identify a moving object among the other objects, i.e., a fourth moving object. Both the third and fourth moving objects are considered moving objects. This method is suitable for removing objects that did not move within the corresponding time difference between the first and second frames, but may subsequently move.

[0140] For example, when the device recognizes that there is a person moving at a slower speed in the picture captured by the camera, due to the slow moving speed, the center of mass of the person in the picture captured within the corresponding time difference between the first picture frame and the second picture frame does not change obviously, and the device cannot detect that the pixel center of mass of the object corresponding to the person has changed, but the object is a dynamic object that should be eliminated. Therefore, in further AI model recognition, the walking person model corresponding to the contour coordinates of the object is identified, and the walking person is identified as a dynamic object through the AI model.

[0141] In the solution provided in this embodiment, the first picture frame and the second picture frame captured by the camera are obtained, and the contour coordinates of each object in the first picture frame and the pixel center of mass of each object between the first picture frame and the second picture frame are determined. According to the contour coordinates and / or pixel center of mass, the moving objects among the objects are judged, so that the interference of dynamic objects in the image on the algorithm is eliminated in the subsequent algorithm, thereby improving the positioning accuracy while reducing the computational complexity of the subsequent algorithm, and solving the problem of how to judge and eliminate dynamic objects in static images.

[0142] Reference Figure 4 In the second embodiment, based on the first embodiment, after step S30, the following steps are included:

[0143] Step S40, when it is determined that the object is the dynamic object, determining a coordinate region associated with the moving object in the second picture frame as a dynamic object coordinate region;

[0144] Step S50, using the area in the first picture frame and the second picture frame except the dynamic object coordinate area as a static object coordinate area;

[0145] Step S60, determining static feature points in the static object coordinate area;

[0146] Step S70: input the static feature points into a target algorithm to obtain an image with the dynamic objects removed.

[0147] In this embodiment, the coordinate area associated with the dynamic object in the second picture frame is used as the dynamic object coordinate area, and the area other than the dynamic object coordinate area is used as the static object coordinate area. Then, the feature points corresponding to the static object coordinate area are determined as static feature points. The feature points are represented as points in the image that have a large difference in pixel value from the surrounding pixels. The static feature points are input into the target algorithm as independent variables to obtain an image without the dynamic object.

[0148] For example, in some specific implementations, the target algorithm is a SLAM algorithm. The following is an example to illustrate the extraction process of the SLAM algorithm:

[0149] Assume that the picture taken by the camera is as follows Figure 5 As shown in the feature point recognition algorithm, a 3*3 sliding frame is used to slide horizontally and vertically from the origin (0, 0) inside the original image. The pixel value distribution within the 3*3 sliding frame is as follows: Figure 6 shown.

[0150] If C3 minus A1 or A3 minus C1 is greater than a certain value, point B2 is considered to be a feature point. Figure 7 shown.

[0151] Furthermore, if the sphere is identified as a dynamic object, the dynamic object coordinate area marking diagram is as follows: Figure 8 As shown in FIG, the multi-deformation angle coordinates of the dynamic object are recorded during the algorithm process, that is, the minimum circumscribed rectangle of the sphere is the coordinate area of the dynamic object.

[0152] When inputting the feature points of the image into the SLAM algorithm, the marked object will be ignored, that is, the four feature points at the sphere will be eliminated, so that the feature points recorded in this image are all extracted from static objects. The final extracted result is as follows Figure 9 shown.

[0153] In the technical solution provided in this embodiment, by identifying, marking, and eliminating dynamic objects in the image, the feature point extraction results are made more effective, and image preprocessing can be performed in the front-end module, such as ISP, DSP, etc., so as to achieve more accurate and high-frequency Pose output.

[0154] Furthermore, those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a computer-readable storage medium. The program instructions are executed by at least one processor in the image processing device to implement the steps of the above-described method embodiment.

[0155] Therefore, the present invention also provides a computer-readable storage medium, which stores an image data processing program. When the image data processing program is executed by a processor, the steps of the image data processing method described in the above embodiment are implemented.

[0156] The computer-readable storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0157] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0158] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a computer-readable storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0159] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for processing image data, characterized in that: The method comprises: Obtain the first picture frame and the second picture frame captured by the camera; Determining the outline coordinates of each object in the first picture frame, and determining the pixel centroid of each object between the first picture frame and the second picture frame; Determining a moving object among the objects according to the contour coordinates and the pixel centroids includes: Determining a contour model of each object according to the contour coordinates; determining a first moving object among the objects according to a first matching degree between the contour model and a reference contour model in a preset dynamic training set; determining a first deflection vector of the camera within a corresponding time difference between the first picture frame and the second picture frame; determining a first field of view angle difference of the camera between the first image frame and the second image frame according to the first deflection vector; determining, based on the first field of view angle difference, first centroid displacements of first pixel centroids of the objects other than the first moving object among the objects between the first picture frame and the second picture frame; determining a second moving object among the other objects according to the first mass center displacement; The moving object is determined according to the first moving object and the second moving object.

2. A method for processing image data, characterized in that: The method comprises: Obtain the first picture frame and the second picture frame captured by the camera; Determining the outline coordinates of each object in the first picture frame, and determining the pixel centroid of each object between the first picture frame and the second picture frame; Determining a moving object among the objects according to the contour coordinates or the pixel centroids includes: determining a deflection vector of the camera within a corresponding time difference between the first picture frame and the second picture frame; determining a field of view angle difference of the camera between the first picture frame and the second picture frame according to the deflection vector; Determining a displacement of the pixel centroid of each object between the first picture frame and the second picture frame based on the field of view angle difference; determining the moving object among the objects according to the mass center displacement, wherein when the mass center displacement is greater than a displacement threshold, determining that the object is the moving object; Alternatively, determining a contour model of the object according to the contour coordinates; The moving object among the objects is determined based on the matching degree between the contour model and a reference contour model in a preset dynamic training set, wherein when the matching degree is greater than a matching threshold, the object is determined to be the moving object.

3. The method for processing image data according to claim 2, wherein: Before the step of determining the moving object among the objects according to the contour coordinates or the pixel centroid, the method includes: determining a picture complexity between the first picture frame and the second picture frame; When the picture complexity is greater than or equal to a complexity threshold, performing the step of determining the deflection vector of the camera within the corresponding time difference between the first picture frame and the second picture frame; Otherwise, the step of determining the contour model of the object according to the contour coordinates is performed.

4. The method for processing image data according to claim 1, wherein: The step of determining a first deflection vector of the camera within a corresponding time difference between the first picture frame and the second picture frame comprises: Obtain inertial data from the inertial measurement unit; The first yaw vector is determined based on the inertial data.

5. The method for processing image data according to claim 1, wherein: Before the step of determining a first field of view angle difference of the camera between the first picture frame and the second picture frame according to the first deflection vector, the method includes: Acquire a visual coordinate system corresponding to the camera and a picture coordinate system corresponding to the first picture frame and the second picture frame; Determining a coordinate mapping between the camera and the first picture frame and the second picture frame according to the visual coordinate system and the picture coordinate system; The step of determining a first field of view angle difference of the camera between the first picture frame and the second picture frame according to the first deflection vector includes: determining, based on the coordinate mapping, a picture translation amount and a picture deflection amount between the first picture frame and the second picture frame according to the deflection vector; The first field of view angle difference is determined according to the picture translation amount and the picture deflection amount.

6. A method for processing image data, characterized in that: The method comprises: Obtain the first picture frame and the second picture frame captured by the camera; Determining the outline coordinates of each object in the first picture frame, and determining the pixel centroid of each object between the first picture frame and the second picture frame; Determining a moving object among the objects according to the contour coordinates and the pixel centroids includes: determining a second deflection vector of the camera within a corresponding time difference between the first picture frame and the second picture frame; determining a second field of view angle difference of the camera between the first picture frame and the second picture frame according to the second deflection vector; determining, based on the second field of view angle difference, a second centroid displacement of a second pixel centroid of each object between the first picture frame and the second picture frame; determining a third moving object among the objects according to the second mass center displacement; determining, according to the contour coordinates of the third moving object, contour models of other objects among the objects except the third moving object; determining a fourth moving object among the other objects based on a second matching degree between the contour model of the other object and a reference contour model in a preset dynamic training set; The moving object is determined according to the third moving object and the fourth moving object.

7. The method for processing image data according to claim 1, wherein: After the step of determining whether the object is a moving object based on the contour coordinates and the pixel centroid, the method further includes: When it is determined that the object is the moving object, determining a coordinate area associated with the moving object in the second picture frame as a dynamic object coordinate area; using the areas other than the dynamic object coordinate area in the first picture frame and the second picture frame as static object coordinate areas; Determining static feature points in the static object coordinate area; The static feature points are input into a target algorithm to obtain an image in which the dynamic objects are removed.

8. An image processing device, characterized in that The image processing device includes: a memory, a processor, and an image data processing program stored in the memory and executable on the processor. When the image data processing program is executed by the processor, the steps of the image data processing method according to any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an image data processing program, which, when executed by a processor, implements the steps of the image data processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Parking lot CCTV monitoring system and method capable of intelligently tracking target

    CN110536114A

  • Image processing method and device before positioning, storage medium and computer equipment

    CN112767409A

  • Computer device, state variable estimation method and state variable estimation program

    JP2020144651A