Depth calculation method, system, readable storage medium, and depth image processing device
By distinguishing the input images and calculating the parallax, the problems of high complexity and high power consumption of the deep vision imaging system are solved, and deep image processing with low power consumption and high accuracy are achieved.
Patent Information
- Application Number
- CN202110145518.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-02
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-02-02
AI Technical Summary
In the prior art, the algorithms of deep vision imaging systems are complex and have high power consumption, especially in power-sensitive devices such as mobile phones and VR glasses.
By distinguishing the input image, the stationary area and the motion area are distinguished. The stationary area is searched in a small range from the previous frame, the motion area and the reference frame are searched in a full range, the parallax of pixels in each area is calculated, and the parallax is accumulated to calculate the depth map.
When the stationary area is large, the low power consumption of the structured light system is ensured, the accuracy of the stationary area judgment is improved, the accumulated error is eliminated, and the system robustness is improved.
Smart Images

Figure CN114842061B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and relates to a calculation method and system, in particular to a depth calculation method, system, readable storage medium and depth image processing device. Background Art
[0002] Depth vision imaging systems include three technical routes: 3D structured light, binocular imaging, and time-of-flight. 3D structured light requires an active light source emitter (infrared LED or VCSEL), and the binocular imaging scheme does not require a light source emitter. Both 3D structured light and binocular imaging are based on the principle of triangulation ranging and feature matching to generate depth image data. The algorithm complexity of 3D structured light and binocular imaging is high, and the power consumption of the entire system path is very obvious. For power-sensitive devices (such as mobile phones, VR glasses, etc.), attention needs to be paid to the power consumption of the entire system path.
[0003] Therefore, how to provide a depth calculation method, system, readable storage medium and depth image processing device to solve the defects of high algorithm complexity and large power consumption in the prior art has actually become a technical problem that those skilled in the art urgently need to solve. Summary of the Invention
[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a depth calculation method, system, readable storage medium and depth image processing device, which are used to solve the problems of high algorithm complexity and large power consumption in the prior art.
[0005] To achieve the above object and other related objects, on the one hand, the present invention provides a depth calculation method, including: differentiating an input image to distinguish a static region and a moving region; performing a small-range search on the static region with the previous frame to calculate the disparity between the current frame and the previous frame of the corresponding pixel points in the static region, and accumulating the disparity between the previous frame and the reference frame of the corresponding pixel points to obtain the disparity between the corresponding pixel points in the static region and the reference frame; performing a full-range search on the moving region with the reference frame to obtain the disparity between the corresponding pixel points in the moving region and the reference frame; calculating the depth map of the input image according to the disparity between the current frame and the reference frame of all pixel points.
[0006] In an embodiment of the present invention, the step of accumulating the disparity between the previous frame and the reference frame of the corresponding pixel points includes: after obtaining the disparity between the current frame and the previous frame of the corresponding pixel points in the static region, determining whether the disparity between the previous frame and the reference frame of the corresponding pixel points is an invalid value; if so, assigning the disparity between the current frame and the reference frame of the corresponding pixel points as an invalid value; if not, determining whether the disparity between the current frame and the previous frame of the corresponding pixel points is an invalid value; when the disparity between the current frame and the previous frame of the corresponding pixel points is an invalid value, the accumulated disparity between the current frame and the reference frame is equal to the disparity between the previous frame and the reference frame of the corresponding pixel points or is assigned as an invalid value; when the disparity between the current frame and the previous frame of the corresponding pixel points is a valid value, the disparity between the previous frame and the reference frame of the corresponding pixel points is accumulated or the disparity between the current frame and the previous frame of the corresponding pixel points is multiplied by a filtering coefficient and then the disparity between the previous frame and the reference frame of the corresponding pixel points is accumulated to obtain the disparity between the current frame and the reference frame of the corresponding pixel points.
[0007] In an embodiment of the present invention, the step of differentiating the input image to distinguish the static region and the moving region includes: subtracting the corresponding pixel points of the current frame and the previous frame of the input image and taking the absolute value of the difference. When the absolute value of the difference is greater than a preset threshold, the motion and stillness flag of the pixel point is set to 1; when the absolute value of the difference is less than or equal to the preset threshold, the motion and stillness flag of the pixel point is set to 0, generating a motion and stillness flag map based on the pixel points; starting from the m*n sized rectangular region in the upper left corner of the motion and stillness flag map, row by row, with a predetermined row interval as the step size, respectively, counting the number of pixel points with a motion and stillness flag of 1 in all m*n sized rectangular regions with a predetermined column interval as the step size, and marking the rectangular regions where the number of pixel points with a motion and stillness flag of 1 exceeds a predetermined number threshold as the moving region; the region not marked as the moving region is the static region.
[0008] In an embodiment of the present invention, the step of differentiating the input image to distinguish the static region and the moving region includes: extracting the speckle information of the input image, marking the pixel points with speckles as 1; marking the pixel points without speckles as 0; performing a per-pixel exclusive OR operation on the speckle map after extracting the speckle information of the current frame and the speckle map after extracting the speckle information of the previous frame to generate a motion and stillness flag map based on the pixel points; starting from the m*n sized rectangular region in the upper left corner of the motion and stillness flag map, row by row, with a predetermined row interval as the step size, respectively, counting the number of pixel points with a motion and stillness flag of 1 in all m*n sized rectangular regions with a predetermined column interval as the step size, and marking the rectangular regions where the number of pixel points with a motion and stillness flag of 1 exceeds a predetermined number threshold as the moving region; the region not marked as the moving region is the static region.
[0009] In an embodiment of the present invention, the step of differentiating the input image to distinguish a static region and a motion region includes: subtracting the corresponding pixel points of the previous frame and the frame before the previous frame of the disparity map or depth map and taking the absolute value of the difference. When the absolute value of the difference is greater than a preset threshold, the static / dynamic flag of the pixel point is set to 1; when the absolute value of the difference is less than or equal to the preset threshold, the static / dynamic flag of the pixel point is set to 0, generating a static / dynamic flag map based on pixel points. Starting from the m*n-sized rectangular region in the upper left corner of the static / dynamic flag map, row by row with a predetermined row interval as the step length, respectively count the number of pixel points with a static / dynamic flag of 1 in all m*n-sized rectangular regions with a predetermined column interval as the step length. Expand the rectangular regions where the number of pixel points with a static / dynamic flag of 1 exceeds a predetermined number threshold and mark them as motion regions; the regions not marked as motion regions are static regions; the distinguished static regions and motion regions are used for the next frame calculation of the depth calculation method described in claim 1; for expanding the rectangle of the motion region, according to the maximum number of moving pixels k between two preset frames, expand the m*n-sized rectangular region marked as the motion region into a (m + 2k)*(n + 2k)-sized rectangular region.
[0010] In an embodiment of the present invention, after calculating the disparity between the current frame and the previous frame of the pixel points corresponding to the static region, the depth calculation method further includes updating the motion region according to the result of a small-range search between the static region and the previous frame. The step of updating the motion region according to the result of the small-range search between the static region and the previous frame includes: marking the points where the disparity between the current frame and the previous frame of the static region is an invalid value as 1, and other points as 0, thereby obtaining a static / dynamic flag map based on points; starting from the m*n-sized rectangular region in the upper left corner of the static / dynamic flag map, row by row with a predetermined row interval as the step length, respectively count the number of pixel points with a static / dynamic flag of 1 in all m*n-sized rectangular regions with a predetermined column interval as the step length. Mark the rectangular regions where the number of pixel points with a static / dynamic flag of 1 exceeds a predetermined number threshold as motion regions; the regions not marked as motion regions remain unchanged with their original values.
[0011] In an embodiment of the present invention, any combination of the differentiation methods for differentiating the input image to distinguish a static region and a motion region is performed. If the combined differentiation methods all mark the rectangular region as a static region, then mark the rectangular region as a static region, and the regions not marked as static regions are motion regions.
[0012] In an embodiment of the present invention, the depth calculation method further includes: adding a static region counter to each region. When the region is a motion region in a certain frame, the counter is assigned a value of 0. When the region is a static region, the counter is incremented by 1. When the counter is greater than or equal to a preset threshold, the region is set as a motion region, and at the same time, the counter is assigned a value of 0.
[0013] On the other hand, the present invention provides a depth calculation system, including: a discrimination module for discriminating an input image to distinguish a static area and a moving area; a static area parallax calculation module for performing a small-range search on the static area and the previous frame to calculate the parallax between the current frame and the previous frame of corresponding pixel points in the static area, and accumulating the parallax between the previous frame and the reference frame of corresponding pixel points to obtain the parallax between the corresponding pixel points in the static area and the reference frame; a moving area parallax calculation module for performing a full-range search on the moving area and the reference frame to obtain the parallax between the corresponding pixel points in the moving area and the reference frame; and a depth calculation module for calculating the depth map of the input image according to the parallax between the current frame and the reference frame of all pixel points.
[0014] On yet another aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the depth calculation method is implemented.
[0015] Finally, the present invention provides a depth image processing device, including: a processor and a memory; the memory is used for storing a computer program, and the processor is used for executing the computer program stored in the memory so that the depth image processing device executes the depth calculation method.
[0016] As described above, the depth calculation method, system, readable storage medium and depth image processing device of the present invention have the following beneficial effects:
[0017] First, when the static area is relatively large, the present invention can ensure the low power consumption of the structured light system.
[0018] Second, the present invention improves the accuracy of static area judgment by combining multiple dynamic and static detection algorithms.
[0019] Third, the present invention improves the system robustness by eliminating the cumulative error. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It shows a schematic diagram of an embodiment of the depth calculation method of the present invention.
[0021] Figure 2 It shows a schematic diagram of rectangular division and statistical order of the present invention.
[0022] Figure 3 It shows a schematic diagram of another embodiment of the depth calculation method of the present invention.
[0023] Figure 4 It shows a schematic diagram of the process of the depth calculation system of the present invention in an embodiment.
[0024] DESCRIPTION OF REFERENCE NUMERALS
[0025] 4 Depth calculation system 41 Differentiation module 42 Stationary area parallax calculation module 43 Moving area parallax calculation module 44 Depth calculation module Specific Embodiments
[0026] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0027] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex. Embodiment 1
[0028] This embodiment provides a depth calculation method, including:
[0029] Differentiate the input image to distinguish the static region and the moving region;
[0030] Perform a small-range search on the static region with the previous frame to calculate the disparity between the current frame and the previous frame of the corresponding pixel points in the static region, and accumulate the disparity between the previous frame and the reference frame of the corresponding pixel points to obtain the disparity between the corresponding pixel points in the static region and the reference frame;
[0031] Perform a full-range search on the moving region with the reference frame to obtain the disparity between the corresponding pixel points in the moving region and the reference frame;
[0032] Calculate the depth map of the input image according to the disparity between the current frame and the reference frame of all pixel points, that is, the disparity map.
[0033] The following will describe in detail the depth calculation method provided in this embodiment with reference to the diagrams. Please refer to Figure 1 , which shows a schematic diagram of an embodiment of the depth calculation method. As Figure 1 shown, the depth calculation method specifically includes the following steps:
[0034] S11, Differentiate the input image to distinguish the static region and the moving region.
[0035] In this embodiment, one way to differentiate the input image to distinguish the static region and the moving region includes:
[0036] Subtract the corresponding pixel points of the current frame of the input image from those of the previous frame and take the absolute value of the difference. When the absolute value of the difference is greater than a preset threshold, set the motion flag of the pixel point to 1; when the absolute value of the difference is less than or equal to the preset threshold, set the motion flag of the pixel point to 0, and generate a motion flag map based on pixel points.
[0037] Starting from the m*n-sized rectangular area in the upper left corner of the motion flag map, count the number of pixel points with a motion flag of 1 in all m*n-sized rectangular areas with a predetermined column interval sx as the step size row by row with a predetermined row interval sy as the step size. Mark the rectangular areas where the number of pixel points with a motion flag of 1 exceeds a predetermined number threshold as motion areas; the areas not marked as motion areas are static areas. In this embodiment, add a static area counter to each area. When the area is a motion area in a certain frame, assign 0 to the counter, and when it is a static area, increment the counter by 1. When the counter is greater than or equal to a preset threshold, set the area as a motion area and assign 0 to the counter at the same time.
[0038] For example, if the resolution of the input image is 640*480, take m = 128, n = 96, sx = 64, and sy = 48. After obtaining the motion flag map based on points (some simple filtering such as Gaussian filtering can be performed on the image before subtracting the corresponding pixels of the two frames), the map can be divided into 100 rectangular blocks of 64*48. First, count the number of 1s in these rectangular blocks respectively to obtain a 10*10 two-dimensional array or image. Just add the number of 1s in adjacent 4 rectangular blocks to obtain the number of 1s in the required rectangular block of 128*96 size, which greatly reduces the calculation amount. The rectangular division and statistical order are as shown in the figure of Figure 2 as shown.
[0039] Or another way to distinguish the input image to distinguish static areas and motion areas includes:
[0040] Extract the speckle information of the input image, mark the pixel points with speckles as 1; mark the pixel points without speckles as 0; perform a pixel-by-pixel exclusive OR operation on the speckle map after extracting the speckle information of the current frame and the speckle map after extracting the speckle information of the previous frame to generate a motion flag map based on pixel points; starting from the m*n-sized rectangular area in the upper left corner of the motion flag map, count the number of pixel points with a motion flag of 1 in all m*n-sized rectangular areas with a predetermined column interval as the step size row by row with a predetermined row interval as the step size. Mark the rectangular areas where the number of pixel points with a motion flag of 1 exceeds a predetermined number threshold as motion areas; the areas not marked as motion areas are static areas.
[0041] Or yet another way to distinguish the input image to distinguish static areas and motion areas includes:
[0042] Subtract the corresponding pixel points of the previous frame of the disparity map or depth map from those of the frame before the previous frame and take the absolute value of the difference. When the absolute value of the difference is greater than a preset threshold, set the motion and stillness flag of this pixel point to 1; when the absolute value of the difference is less than or equal to the preset threshold, set the motion and stillness flag of the pixel point to 0, and generate a motion and stillness flag map based on the pixel points.
[0043] Starting from the m*n-sized rectangular area in the upper left corner of the motion and stillness flag map, count the number of pixel points with a motion and stillness flag of 1 in all m*n-sized rectangular areas with a predetermined column interval as the step size row by row with a predetermined row interval as the step size. Expand the rectangles with the number of pixel points with a motion and stillness flag of 1 exceeding a predetermined number threshold and mark them as motion areas; the areas not marked as motion areas are still areas; the distinguished still areas and motion areas are used for the next frame calculation of the depth calculation method described in claim 1.
[0044] Regarding the expansion of the rectangles of the motion areas, according to the maximum number k of moving pixels between two frames preset,
[0045] Expand the m*n-sized rectangular area marked as a motion area into a (m + 2k)*(n + 2k)-sized rectangular area.
[0046] In this embodiment, combine the above three methods of distinguishing the input image to distinguish still areas and motion areas. If all three distinguishing methods mark this rectangular area as a still area, then mark this rectangular area as a still area, and the areas not marked as still areas are motion areas.
[0047] S12. Conduct a small-range search on the still area and the previous frame to calculate the disparity between the current frame and the previous frame of the corresponding pixel points in the still area, update the motion area according to the result of the small-range search on the still area and the previous frame, and accumulate the disparity between the previous frame of the corresponding pixel points in the updated still area and the reference frame to obtain the disparity between the corresponding pixel points in the updated still area and the reference frame.
[0048] Specifically, in the small-range search process, a rectangular window rec1 is obtained centered on any pixel point p1 in the static area of the current frame. The corresponding small-range search area area1 of the rectangular window rec1 in the previous frame is obtained, and a matching search is performed within the corresponding small-range search area area1 to obtain the matching pixel point p2 of the pixel point p1. The matching cost can adopt normalized cross-correlation. According to the positions and matching costs of the pixel point p1 and its most matching pixel point p2, the disparity between the pixel point p1 and the previous frame is obtained. Wherein, in this embodiment, the moving direction of the structured light pattern when the distance becomes farther is taken as the positive direction, and the absolute value of the above-mentioned disparity value refers to the distance value in the disparity direction between the position of the pixel point p1 in the image to be processed and the position of its matching pixel point p2 in the previous frame image. When the matching cost is less than a preset threshold, the disparity value of the pixel point p1 is set to an invalid value. It should be noted that in this embodiment, other directions can also be selected as the positive direction, and the above-mentioned scheme can be adjusted accordingly according to the actually selected positive direction in specific applications.
[0049] In this embodiment, the step of accumulating the disparity between the previous frame and the reference frame of the corresponding pixel point in S12 includes:
[0050] After obtaining the disparity between the current frame and the previous frame of the corresponding pixel point in the updated static area, it is judged whether the disparity between the previous frame and the reference frame of the corresponding pixel point is an invalid value; if so, the disparity between the current frame and the reference frame of the corresponding pixel point is assigned an invalid value; if not, it is judged whether the disparity between the current frame and the previous frame of the corresponding pixel point is an invalid value;
[0051] If the disparity between the current frame and the previous frame of the corresponding pixel point is an invalid value, then the disparity between the accumulated current frame and the reference frame is equal to or assigned an invalid value to the disparity between the previous frame and the reference frame of the corresponding pixel point;
[0052] If the disparity between the current frame and the previous frame of the corresponding pixel point is a valid value, then the disparity between the previous frame and the reference frame of the corresponding pixel point is accumulated, or the disparity between the current frame and the previous frame of the corresponding pixel point is multiplied by a filtering coefficient and then the disparity between the previous frame and the reference frame of the corresponding pixel point is accumulated to obtain the disparity between the current frame and the reference frame of the corresponding pixel point.
[0053] In this embodiment, if the accumulated disparity exceeds the search range, the accumulated disparity can be limited so that it still remains within the search range; actually, the range of the accumulated disparity can also be a little larger than the actual search range, as long as it does not contain invalid values.
[0054] To reduce the impact of the cumulative error caused by adjacent frame matching in the static area, a static area counter can be added to each area. When the area is a moving area in a certain frame, the counter is assigned a value of 0. When the area is a static area, the counter is incremented by 1. When the counter is greater than or equal to a threshold value, the area does not perform a small-range search with the previous frame to calculate the disparity, but directly performs a search within a certain search range with the reference frame to calculate the disparity, and at the same time, the counter is assigned a value of 0. This search range can be expanded based on the range between the maximum and minimum disparities of all pixel points in this area and the reference frame, which can reduce the computational amount.
[0055] The step of updating the moving area according to the result of the small-range search between the static area and the previous frame includes:
[0056] Mark the points with an invalid disparity between the current frame and the previous frame in the static area as 1, and other points as 0, so as to obtain a point-based static and dynamic flag map;
[0057] Starting from the m*n-sized rectangular area in the upper left corner of the static and dynamic flag map, count the number of static and dynamic flags as 1 in all m*n-sized rectangular areas with a predetermined column interval as the step size row by row with a predetermined row interval as the step size, and mark the rectangular areas where the number of static and dynamic flags as 1 exceeds a predetermined number threshold as moving areas; the areas not marked as moving areas remain unchanged with their original values.
[0058] S13, perform a full-range search on the moving area and the reference frame to obtain the disparity between the corresponding pixel points in the moving area and the reference frame. In this embodiment, the process of calculating the disparity is similar to the process of calculating the disparity in the previous small-range search, except that the previous frame is replaced with the reference frame and the search range is increased, which will not be elaborated here.
[0059] S14, calculate the depth map of the input image according to the disparity between the current frame and the reference frame of all pixel points; for all pixel points with valid disparity, calculate the depth according to the formula for calculating depth from disparity, and for all pixel points with invalid disparity, assign the depth value as an invalid value.
[0060] The formula for calculating depth from disparity,
[0061] where Z is the distance from the target point to be measured in the scene to the camera lens plane, that is, the depth; R is the distance from the reference plane to the lens plane (determined by the calibration environment), d is the calculated disparity (the direction of the structured light pattern movement is the positive direction when the distance becomes farther), f is the focal length of the receiver camera (which can be obtained through camera internal parameter calibration), and b is the length of the line connecting the perspective centers of the projector and the receiver (that is, the baseline distance, which can be controlled within a range through structural design).
[0062] Please refer to Figure 3, shown as a schematic diagram of another embodiment of the depth calculation method. As Figure 3 shown, the depth calculation method specifically includes the following steps:
[0063] S31, starting from the second frame of the input image, processing the entire frame image in the manner of a static region, performing a small-range search with the previous frame to calculate the disparity between the current frame and the previous frame for each pixel point, and updating the motion region according to the result of the small-range search of the entire frame image and the previous frame.
[0064] Among them, the step of updating the motion region according to the result of the small-range search of the entire frame image and the previous frame includes:
[0065] Mark the points with invalid disparity between the current frame and the previous frame in the static region as 1, and other points as 0, so as to obtain a point-based static and dynamic flag map;
[0066] Starting from the m*n-sized rectangular region in the upper left corner of the static and dynamic flag map, statistically count the number of static and dynamic flags as 1 in all m*n-sized rectangular regions with a predetermined column interval as the step size row by row with a predetermined row interval as the step size, and mark the rectangular regions with the number of static and dynamic flags as 1 exceeding the predetermined number threshold as the motion region; the regions not marked as the motion region remain unchanged with their original values, that is, the static region.
[0067] S32, accumulate the disparity between the previous frame and the reference frame corresponding to the pixel points in the updated static region to obtain the disparity between each pixel point and the reference frame. The specific implementation method is similar to S12 and will not be elaborated here.
[0068] S33, perform a full-range search on the updated motion region and the reference frame to obtain the disparity between the pixel points corresponding to the motion region and the reference frame. In this embodiment, the specific process is similar to S13 and will not be elaborated here.
[0069] S34, calculate the depth map of the input image according to the disparity between the current frame and the reference frame of all pixel points; the specific process is similar to S14 and will not be elaborated here.
[0070] The depth calculation method described in this embodiment has the following beneficial effects:
[0071] First, the depth calculation method described in this embodiment can ensure the low power consumption of the structured light system, especially when the static region is relatively large;
[0072] Second, the depth calculation method described in this embodiment improves the accuracy of static region judgment through the combined use of multiple static and dynamic detection algorithms.
[0073] Third, the depth calculation method described in this embodiment improves the system robustness by eliminating the cumulative error method.
[0074] This embodiment also provides a storage medium (also known as a computer-readable storage medium) on which a computer program is stored, and when the program is executed by a processor, the depth calculation method is implemented.
[0075] Those of ordinary skill in the art can understand that for a computer-readable storage medium: all or part of the steps of implementing the above method embodiments can be completed by hardware related to the computer program. The foregoing computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes. Embodiment 2
[0076] This embodiment provides a depth calculation system, including:
[0077] A discrimination module that discriminates an input image to distinguish a static area and a moving area;
[0078] A static area parallax calculation module for performing a small-range search on the static area and the previous frame to calculate the parallax between the current frame and the previous frame of the corresponding pixel points in the static area, and accumulating the parallax between the previous frame and the reference frame of the corresponding pixel points to obtain the parallax between the corresponding pixel points in the static area and the reference frame;
[0079] A moving area parallax calculation module for performing a full-range search on the moving area and the reference frame to obtain the parallax between the corresponding pixel points in the moving area and the reference frame;
[0080] A depth calculation module for calculating the depth map of the input image according to the parallax between the current frame and the reference frame of all pixel points.
[0081] The depth calculation system provided in this embodiment will be described in detail below with reference to the drawings. Please refer to Figure 4 , which shows a schematic diagram of the principle structure of the depth calculation system in an embodiment. As Figure 4 shown, the depth calculation system 4 includes a discrimination module 41, a static area parallax calculation module 42, a moving area parallax calculation module 43, and a depth calculation module 44.
[0082] The discrimination module 41 is used to discriminate an input image to distinguish a static area and a moving area.
[0083] Specifically, in one implementation of the discrimination module 41, the current frame of the input image is subtracted from the corresponding pixel points of the previous frame and the absolute value of the difference is taken. When the absolute value of the difference is greater than a preset threshold, the movement and stillness flag of the pixel point is set to 1; when the absolute value of the difference is less than or equal to the preset threshold, the movement and stillness flag of the pixel point is set to 0, generating a movement and stillness flag map based on pixel points. Starting from the m*n-sized rectangular area in the upper left corner of the movement and stillness flag map, the number of pixel points with a movement and stillness flag of 1 in all m*n-sized rectangular areas with a predetermined column interval as the step length is statistically calculated row by row with a predetermined row interval as the step length. The rectangular areas where the number of pixel points with a movement and stillness flag of 1 exceeds a predetermined number threshold are marked as moving areas; the areas not marked as moving areas are still areas. In this embodiment, a still area counter is added to each area. When the area is a moving area in a certain frame, the counter is assigned a value of 0. When the area is a still area, the counter is incremented by 1. When the counter is greater than or equal to a preset threshold, the area is set as a moving area, and at the same time, the counter is assigned a value of 0.
[0084] Specifically, in another implementation of the discrimination module 41, speckle information of the input image is extracted. Pixel points with speckles are marked as 1; pixel points without speckles are marked as 0. The speckle map after extracting speckle information of the current frame is subjected to an exclusive OR operation with the speckle map after extracting speckle information of the previous frame pixel by pixel, generating a movement and stillness flag map based on pixel points. Starting from the m*n-sized rectangular area in the upper left corner of the movement and stillness flag map, the number of pixel points with a movement and stillness flag of 1 in all m*n-sized rectangular areas with a predetermined column interval as the step length is statistically calculated row by row with a predetermined row interval as the step length. The rectangular areas where the number of pixel points with a movement and stillness flag of 1 exceeds a predetermined number threshold are marked as moving areas; the areas not marked as moving areas are still areas.
[0085] Specifically, in yet another implementation of the discrimination module 41, the absolute value of the difference between the previous frame and the frame before the previous frame of the disparity map or depth map is taken for the corresponding pixel points. When the absolute value of the difference is greater than a preset threshold, the movement and stillness flag of the pixel point is set to 1; when the absolute value of the difference is less than or equal to the preset threshold, the movement and stillness flag of the pixel point is set to 0, generating a movement and stillness flag map based on pixel points. Starting from the m*n-sized rectangular area in the upper left corner of the movement and stillness flag map, the number of pixel points with a movement and stillness flag of 1 in all m*n-sized rectangular areas with a predetermined column interval as the step length is statistically calculated row by row with a predetermined row interval as the step length. The rectangular areas where the number of pixel points with a movement and stillness flag of 1 exceeds a predetermined number threshold are enlarged and then marked as moving areas; the areas not marked as moving areas are still areas. The discriminated still areas and moving areas are used for the next frame calculation of the depth calculation method described in claim 1. For the enlargement of the rectangle of the moving area, according to the maximum number of moving pixels k between two frames preset, the m*n-sized rectangular area marked as a moving area is enlarged to a (m + 2k)*(n + 2k)-sized rectangular area.
[0086] In this embodiment, a static region counter is added to each region. When a certain frame shows that the region is a moving region, the counter is assigned a value of 0. When the region is a static region, the counter is incremented by 1. When the counter is greater than or equal to a preset threshold, the region is set as a moving region, and at the same time, the counter is assigned a value of 0.
[0087] The static region disparity calculation module 42 is configured to perform a small-range search on the static region and the previous frame to calculate the disparity between the current frame and the previous frame of the corresponding pixel points in the static region, and accumulate the disparity between the previous frame and the reference frame of the corresponding pixel points to obtain the disparity between the corresponding pixel points in the static region and the reference frame.
[0088] Specifically, in the small-range search process, a rectangular window rec1 is obtained with any pixel point p1 in the static region of the current frame as the center. The corresponding small-range search area area1 of the rectangular window rec1 in the previous frame is obtained, and a matching search is performed within the corresponding small-range search area area1 to obtain the matching pixel point p2 of the pixel point p1. The matching cost can adopt normalized cross-correlation. According to the positions and matching cost of the pixel point p1 and its most matching pixel point p2, the disparity between the pixel point p1 and the previous frame is obtained. Here, in this embodiment, the moving direction of the structured light pattern when the distance becomes farther is taken as the positive direction, and the absolute value of the above disparity value refers to the distance value in the disparity direction between the position of the pixel point p1 in the image to be processed and the position of its matching pixel point p2 in the previous frame image. When the matching cost is less than a preset threshold, the disparity value of the pixel point p1 is set to an invalid value. It should be noted that in this embodiment, other directions can also be selected as the positive direction, and the above scheme can be adjusted accordingly according to the actually selected positive direction in specific applications.
[0089] In this embodiment, after the static region disparity calculation module 42 obtains the disparity between the current frame and the previous frame of the corresponding pixel points in the static region, it is determined whether the disparity between the previous frame and the reference frame of the corresponding pixel points is an invalid value; if so, the disparity between the current frame and the reference frame of the corresponding pixel points is assigned an invalid value; if not, it is determined whether the disparity between the current frame and the previous frame of the corresponding pixel points is an invalid value; if the disparity between the current frame and the previous frame of the corresponding pixel points is an invalid value, then the accumulated disparity between the current frame and the reference frame is equal to the disparity between the previous frame and the reference frame of the corresponding pixel points or is assigned an invalid value; if the disparity between the current frame and the previous frame of the corresponding pixel points is a valid value, then the disparity between the previous frame and the reference frame of the corresponding pixel points is accumulated or the disparity between the current frame and the previous frame of the corresponding pixel points is multiplied by a filtering coefficient and then the disparity between the previous frame and the reference frame of the corresponding pixel points is accumulated to obtain the disparity between the current frame and the reference frame of the corresponding pixel points.
[0090] In this embodiment, if the disparity after accumulation by the static area disparity calculation module 42 exceeds the search range, the accumulated disparity can be limited to keep it within the search range; actually, the range of the accumulated disparity can also be a bit larger than the actual search range, as long as it does not contain invalid values.
[0091] The moving area disparity calculation module 43 is used to perform a full-range search on the moving area and the reference frame to obtain the disparity between the corresponding pixel points of the moving area and the reference frame.
[0092] The depth calculation module 44 is used to calculate the depth map of the input image according to the disparity between the current frame and the reference frame of all pixel points.
[0093] It should be noted that it should be understood that the division of each module of the above system is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements, or all in the form of hardware, or some modules in the form of software called by processing elements and some modules in the form of hardware. For example: The x module can be a separately established processing element, or can be integrated in a certain chip of the above system. In addition, the x module can also be stored in the memory of the above system in the form of program code, and the function of the above x module is called and executed by a certain processing element of the above system. The implementation of other modules is similar. These modules can be fully or partially integrated together, or can be independently implemented. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software. The above modules can be one or more integrated circuits configured to implement the above method, for example: one or more application specific integrated circuits (ASICs), one or more digital signal processors (DSPs), one or more field programmable gate arrays (FPGAs), etc. When a certain module above is implemented in the form of a program code called by a processing element, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. These modules can be integrated together and implemented in the form of a system-on-a-chip (SOC). Embodiment III
[0094] This embodiment provides a depth image processing device, which includes: a processor, a memory, a transceiver, a communication interface, or / and a system bus; the memory and the communication interface are connected to the processor and the transceiver through the system bus to complete communication with each other. The memory is used to store computer programs, the communication interface is used to communicate with other devices, and the processor and the transceiver are used to run the computer programs, so that the depth image processing device executes each step of the depth calculation method described in the above Embodiment 1.
[0095] The above-mentioned system bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to realize the communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0096] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0097] The protection scope of the depth calculation method described in the present invention is not limited to the execution order of the steps listed in this embodiment. Any solution achieved by adding or subtracting steps of the prior art and replacing steps according to the principle of the present invention is included in the protection scope of the present invention.
[0098] The present invention also provides a depth calculation system, which can implement the depth calculation access method of the present invention. However, the implementation device of the depth calculation method of the present invention includes, but is not limited to, the structure of the access system for a large amount of data listed in this embodiment. Any structural deformation and replacement of the prior art made according to the principles of the present invention are included in the protection scope of the present invention.
[0099] In summary, the depth calculation method, system, storage medium, and depth image processing device of the present invention have the following beneficial effects:
[0100] First, when the static area is relatively large, the present invention can ensure the low power consumption of the structured light system.
[0101] Second, the present invention improves the accuracy of static area judgment by combining multiple dynamic and static detection algorithms.
[0102] Third, the present invention improves the system robustness by eliminating the cumulative error. The present invention effectively overcomes various disadvantages in the prior art and has high industrial utilization value.
[0103] The above embodiments are only illustrative of the principles and effects of the present invention, and are not used to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A depth calculation method, characterized in that, comprising: Differentiating the input image to distinguish a static region and a moving region; Performing a small-range search on the static region with the previous frame to calculate the disparity between the current frame and the previous frame of the corresponding pixel points in the static region, and accumulating the disparity between the previous frame and the reference frame of the corresponding pixel points to obtain the disparity between the corresponding pixel points in the static region and the reference frame; Performing a full-range search on the moving region with the reference frame to obtain the disparity between the corresponding pixel points in the moving region and the reference frame; and Calculating the depth map of the input image according to the disparity between the current frame and the reference frame of all pixel points, wherein the differentiating method for the static region and the moving region includes: Determining the static / dynamic flag setting of pixel points based on the comparison between the absolute value of the difference obtained by subtracting the corresponding pixel points of the current frame and the previous frame of the input image and a preset threshold, generating a static / dynamic flag map based on the pixel points, and differentiating the moving region and the static region based on the static / dynamic flags in the static / dynamic flag map; Or, extracting the speckle information of the input image, determining the pixel point flag based on the speckle information, performing a per-pixel exclusive OR operation on the speckle map after extracting the speckle information of the current frame and the speckle map after extracting the speckle information of the previous frame, generating a static / dynamic flag map based on the pixel points, and differentiating the moving region and the static region based on the static / dynamic flags in the static / dynamic flag map; Or, determining the static / dynamic flag of pixel points based on the comparison between the absolute value of the difference obtained by subtracting the corresponding pixel points of the previous frame and the frame before the previous frame of the disparity map or the depth map and a preset threshold, generating a static / dynamic flag map based on the pixel points, and differentiating the moving region and the static region based on the static / dynamic flags in the static / dynamic flag map, wherein differentiating the input image to distinguish a static region and a moving region includes: combining the differentiating methods for the static region and the moving region. If all three differentiating methods mark a rectangular region as a static region, then mark this rectangular region as a static region, and mark the region not marked as a static region as a moving region.
2. The depth calculation method according to claim 1, characterized in that, the step of accumulating the disparity between the previous frame and the reference frame of the corresponding pixel points includes: After obtaining the disparity between the current frame and the previous frame of the corresponding pixel points in the static region, determining whether the disparity between the previous frame and the reference frame of the corresponding pixel points is an invalid value; if so, assigning the disparity between the current frame and the reference frame of the corresponding pixel points as an invalid value; if not, determining whether the disparity between the current frame and the previous frame of the corresponding pixel points is an invalid value; If the disparity between the current frame and the previous frame of the corresponding pixel points is an invalid value, then the disparity between the current frame and the reference frame after accumulation is equal to the disparity between the previous frame and the reference frame of the corresponding pixel points or is assigned as an invalid value; If the disparity between the current frame and the previous frame of the corresponding pixel points is a valid value, then accumulating the disparity between the previous frame and the reference frame of the corresponding pixel points or multiplying the disparity between the current frame and the previous frame of the corresponding pixel points by a filtering coefficient and then accumulating the disparity between the previous frame and the reference frame of the corresponding pixel points to obtain the disparity between the current frame and the reference frame of the corresponding pixel points.
3. The depth calculation method according to claim 1, wherein, determining the dynamic and static flag setting of pixel points based on the comparison between the absolute value of the difference obtained by subtracting the corresponding pixel points of the current frame and the previous frame of the input image and a preset threshold, generating a dynamic and static flag map based on pixel points, and distinguishing the moving area and the static area based on the dynamic and static flag of the dynamic and static flag map includes: subtracting the corresponding pixel points of the current frame and the previous frame of the input image and taking the absolute value of the difference. When the absolute value of the difference is greater than a preset threshold, set the dynamic and static flag of the pixel point to 1; when the absolute value of the difference is less than or equal to the preset threshold, set the dynamic and static flag of the pixel point to 0, generating a dynamic and static flag map based on pixel points; starting from the m*n-sized rectangular area in the upper left corner of the dynamic and static flag map, respectively counting the number of dynamic and static flags with a value of 1 in all m*n-sized rectangular areas with a predetermined column interval as the step size row by row with a predetermined row interval as the step size, and marking the rectangle with the number of dynamic and static flags with a value of 1 exceeding a predetermined number threshold as the moving area; the area not marked as the moving area is the static area.
4. The depth calculation method according to claim 3, wherein, extracting the speckle information of the input image, determining the pixel point flag based on the speckle information, performing a per-pixel exclusive OR operation on the speckle map after extracting the speckle information of the current frame and the speckle map after extracting the speckle information of the previous frame, generating a dynamic and static flag map based on pixel points, and distinguishing the moving area and the static area based on the dynamic and static flag of the dynamic and static flag map includes: extracting the speckle information of the input image, marking the pixel points with speckles as 1; marking the pixel points without speckles as 0; performing a per-pixel exclusive OR operation on the speckle map after extracting the speckle information of the current frame and the speckle map after extracting the speckle information of the previous frame, generating a dynamic and static flag map based on pixel points; starting from the m*n-sized rectangular area in the upper left corner of the dynamic and static flag map, respectively counting the number of dynamic and static flags with a value of 1 in all m*n-sized rectangular areas with a predetermined column interval as the step size row by row with a predetermined row interval as the step size, and marking the rectangle with the number of dynamic and static flags with a value of 1 exceeding a predetermined number threshold as the moving area; the area not marked as the moving area is the static area.
5. The depth calculation method according to claim 4, wherein, determining the dynamic and static flag of pixel points based on the comparison between the absolute value of the difference obtained by subtracting the corresponding pixel points of the previous frame and the frame before the previous frame of the disparity map or depth map and a preset threshold, generating a dynamic and static flag map based on pixel points, and distinguishing the moving area and the static area based on the dynamic and static flag of the dynamic and static flag map includes: subtracting the corresponding pixel points of the previous frame and the frame before the previous frame of the disparity map or depth map and taking the absolute value of the difference. When the absolute value of the difference is greater than a preset threshold, set the dynamic and static flag of the pixel point to 1; when the absolute value of the difference is less than or equal to the preset threshold, set the dynamic and static flag of the pixel point to 0, generating a dynamic and static flag map based on pixel points; Starting from the m*n-sized rectangular area in the upper left corner of the static and dynamic flag map, row by row with a predetermined row interval as the step, respectively count the number of static and dynamic flags being 1 in all m*n-sized rectangular areas with a predetermined column interval as the step. Expand the rectangles whose number of static and dynamic flags being 1 exceeds a predetermined number threshold and mark them as motion areas; the areas not marked as motion areas are static areas; the distinguished static areas and motion areas are used for the next-frame calculation of the depth calculation method described in claim 1; Expand the rectangles in the motion area. According to the maximum number of motion pixels k between two preset frames, Expand the m*n-sized rectangular area marked as a motion area into a (m + 2k)*(n + 2k)-sized rectangular area.
6. According to the depth calculation method described in claim 1, characterized in that, After calculating the disparity between the current frame and the previous frame of the pixel points corresponding to the static area, the depth calculation method further includes updating the motion area according to the result of a small-range search between the static area and the previous frame. The steps of updating the motion area according to the result of the small-range search between the static area and the previous frame include: Mark the points where the disparity between the current frame and the previous frame of the static area is an invalid value as 1, and other points as 0, thus obtaining a point-based static and dynamic flag map; Starting from the m*n-sized rectangular area in the upper left corner of the static and dynamic flag map, row by row with a predetermined row interval as the step, respectively count the number of static and dynamic flags being 1 in all m*n-sized rectangular areas with a predetermined column interval as the step. Mark the rectangles whose number of static and dynamic flags being 1 exceeds a predetermined number threshold as motion areas; the areas not marked as motion areas remain their original values unchanged.
7. According to the depth calculation method described in claim 2, characterized in that, the depth calculation method further includes: Add a static area counter to each area. When the area is a motion area in a certain frame, assign the counter a value of 0. When the area is a static area, increment the counter by 1. When the counter is greater than or equal to a preset threshold, directly perform a search calculation of the disparity within a certain search range between the area and the reference frame, and assign the counter a value of 0.
8. A depth calculation system, characterized in that, comprising: A distinguishing module that distinguishes the input image to distinguish static areas and motion areas; A static area disparity calculation module for performing a small-range search between the static area and the previous frame to calculate the disparity between the current frame and the previous frame of the pixel points corresponding to the static area, and accumulating the disparity between the previous frame and the reference frame of the corresponding pixel points to obtain the disparity between the pixel points corresponding to the static area and the reference frame; A motion area disparity calculation module for performing a full-range search between the motion area and the reference frame to obtain the disparity between the pixel points corresponding to the motion area and the reference frame; and A depth calculation module for calculating the depth map of the input image according to the disparity between the current frame and the reference frame of all pixel points, where the distinguishing method of the static area and the motion area includes: Determine the dynamic and static flag settings of pixel points based on the comparison between the absolute value of the difference obtained by subtracting the corresponding pixel points of the current frame and the previous frame of the input image and a preset threshold, generate a dynamic and static flag map based on the pixel points, and distinguish the moving area and the static area based on the dynamic and static flags of the dynamic and static flag map; Alternatively, extract the speckle information of the input image, determine the pixel point flags based on the speckle information, perform a pixel-by-pixel exclusive OR operation on the speckle map after extracting the speckle information of the current frame and the speckle map after extracting the speckle information of the previous frame, generate a dynamic and static flag map based on the pixel points, and distinguish the moving area and the static area based on the dynamic and static flags of the dynamic and static flag map; Alternatively, determine the dynamic and static flags of pixel points based on the comparison between the absolute value of the difference obtained by subtracting the corresponding pixel points of the previous frame and the frame before the previous frame of the disparity map or depth map and a preset threshold, generate a dynamic and static flag map based on the pixel points, and distinguish the moving area and the static area based on the dynamic and static flags of the dynamic and static flag map, wherein the distinguishing module is used to combine the distinguishing methods of the static area and the moving area. If all three distinguishing methods mark the rectangular area as the static area, then mark the rectangular area as the static area, and mark the area not marked as the static area as the moving area.
9. A readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, it implements the depth calculation method described in any one of claims 1 to 7.
10. A depth image processing device, characterized in that, comprising: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the depth image processing device executes the depth calculation method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Parallax calculation method and device
CN101790103A
Stereoscopic video depth map generation method and device
CN102111637A
Disparity estimation method and device for auto convergence of region of interest using adaptive search range prediction in video
KR101340086B1