Oral cavity image processing method, device and equipment
By acquiring oral video streams and gyroscope data, keyframe images are extracted and grid correction and four-way template matching are performed, corner point information is corrected, and image fusion is performed based on the homography matrix, which solves the problem of difficult stitching of the panoramic image of the handheld device, and the reconstruction of the panoramic image of the teeth of the handheld device is achieved.
Patent Information
- Application Number
- CN202311621215.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
Existing handheld devices are prone to difficulty in stitching images due to shaking and misalignment when shooting panoramic images of teeth, and there are problems with radiation risks and high cost when shooting dental films.
Oral image processing method is adopted, by obtaining oral video streams and gyroscope data, keyframe images are extracted, grid correction and four-way template matching are performed, corner point information is corrected, and image fusion is performed based on homography matrix to achieve accurate splicing of tooth panoramic images.
It realizes clear shooting of the panoramic image of the teeth independently on the handheld device, solves the difficulty of image stitching caused by shaking and dislocation, reduces radiation risks and costs, and provides accurate three-dimensional tooth observation capabilities.
Smart Images

Figure CN120070722A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of oral cavity detection, and particularly to a method, device, and equipment for processing oral cavity images. Background Art
[0002] When people need to clearly view the arrangement of teeth on the inner wall of their oral cavity or problems on the tooth surface, they often use the method of looking in the mirror. However, due to the field of view and the distance from the mirror, it is impossible to obtain a full view of the teeth by looking in the mirror. Therefore, people often consider using digital imaging to view.
[0003] Currently, digital images of the oral cavity usually need to be taken with a dental X-ray machine in the hospital. By using X-rays, the arrangement of teeth and the condition of tooth roots can be clearly seen. However, taking pictures with a dental X-ray machine is relatively expensive and may have certain radiation effects. Therefore, people cannot view the internal condition of the oral cavity at any time and place. So, handheld mobile devices that can facilitate people to detect the oral cavity have emerged, such as endoscopes and dental irrigators, with cameras installed on them for detection.
[0004] Since the dental X-ray machine is in a fixed state, it is relatively easy to splice after scanning the teeth at close range. However, due to human operation, handheld devices will have problems of shaking and dislocation during the process of taking pictures of teeth at close range. It is difficult for people to view the panoramic picture of teeth when they finish scanning the teeth. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention provides a method, device, and equipment for processing oral cavity images, which can meet the needs of users to independently complete the shooting of oral cavity images and obtain clear panoramic images of teeth at close range.
[0006] The present invention adopts the following technical solutions:
[0007] On the one hand, a method for processing oral cavity images includes:
[0008] Obtain the collected oral cavity video stream and the corresponding gyroscope data;
[0009] Perform frame splitting on the oral cavity video stream to obtain key frame images;
[0010] Perform grid correction on the key frame images to obtain corrected images;
[0011] Use the four-direction template matching method to sequentially match the corrected images to obtain the current frame image and the next frame image that meet the corner point information matching conditions, and perform alignment;
[0012] Correct the corner point information of the next frame image based on the gyroscope data;
[0013] Based on the corner information of the current frame image and the corrected corner information of the next frame image, obtain the homography matrix between the current frame image and the next frame image;
[0014] Based on the homography matrix, fuse the current frame image and the next frame image to obtain a fused image.
[0015] Preferably, perform frame splitting on the oral video stream to obtain key frame images, specifically including:
[0016] S1021, disassemble each frame from the oral video stream;
[0017] S1022, calculate and compare the similarity between the current frame and the next frame;
[0018] S1023, if the similarity is less than the key frame minimum threshold, it means the key frame is lost, and subsequent frames are not judged;
[0019] S1024, if the similarity is greater than the key frame maximum threshold, take the frame after the next frame as the next frame, and return to S1022;
[0020] S1025, if the similarity is less than or equal to the key frame maximum threshold and greater than or equal to the key frame minimum threshold, then take the next frame as a key frame image and add it to the key frame list;
[0021] S1026, judge whether there are still frames not compared. If so, take the next frame as the current frame, take the frame after the next frame as the next frame, and return to S1022.
[0022] Preferably, the calculation method of the similarity between the current frame and the next frame includes:
[0023] Perform single-channel similarity calculation on the current frame and the next frame;
[0024] Average the similarities of the three single channels to obtain the similarity between the current frame and the next frame.
[0025] Preferably, after performing frame splitting on the oral video stream to obtain key frame images, it further includes: performing tooth instance segmentation based on the key frame images, and extracting tooth contour information with the gingival edge area as a feature point; specifically as follows:
[0026] Use the TeethInsReg model to perform tooth instance segmentation and extract tooth contour information with the gingival edge area as a feature point; the TeethInsReg model includes an encoding module encode, a feature pyramid network module FPN, a decoding module decode, a classification module category, and a mask module mask connected in sequence;
[0027] The input of the TeethInsReg model is a key-frame image, and the output is the category and the mask of the corresponding segmentation instance;
[0028] Binarize the mask to obtain the contour information of each tooth.
[0029] Preferably, use the four-direction template matching method to match the corrected images in sequence to obtain the current frame image and the next frame image that meet the corner point information matching conditions, specifically including:
[0030] Divide the current frame image into four template blocks, and perform template matching on each of the four template blocks with the next frame image to obtain the block with the highest matching rate with the next frame;
[0031] If the highest matching rate is greater than the set matching threshold, it means that there is an overlapping area between the current frame and the next frame, and subsequent corner point matching can be performed; otherwise, match the current frame with the frame after the next frame.
[0032] Preferably, align the current frame and the next frame image, specifically including:
[0033] Match the corner point information of the current frame with the corner point information of the next frame;
[0034] If the number of matching corner feature points between the current frame and the next frame exceeds the preset number, align the current frame and the next frame image.
[0035] Preferably, correct the corner point information of the next frame image based on the gyroscope data, specifically including:
[0036] Extract the first corner point information based on the contour information of the next frame image, calculate the second corner point information of the next frame image based on the gyroscope data, compare the feature points of the first corner point information and the second corner point information in sequence, and replace the feature points of the first corner point information with a deviation greater than the preset pixel with the feature points of the second corner point information to obtain the corrected corner point information of the next frame image.
[0037] Preferably, calculate the second corner point information of the next frame image based on the gyroscope data, specifically including:
[0038] According to the attitude information of the gyroscope, calculate the actual rotation angle and speed between the current frame and the next frame image; the attitude information includes the roll angle, yaw angle and pitch angle of the current frame and the next frame image;
[0039] Based on the roll angle, yaw angle and pitch angle of the current frame and the next frame image, calculate the rotation angle and rotation speed from the current frame image to the next frame image;
[0040] Calculate the corner point information of the next frame image based on the rotation angle and rotation speed.
[0041] Preferably, the corner information of the current frame image is extracted based on the contour information of the current frame image.
[0042] Preferably, based on the homography matrix, the current frame image and the next frame image are fused to obtain a fused image, specifically including:
[0043] Obtain the image information of the current frame and the image information of the next frame, and based on the homography matrix, obtain the overlapping position of the two images;
[0044] Based on the overlapping position, obtain the stitching size;
[0045] Correspond the initial position of the current frame with the initial position of the stitching size, and correspond the end position of the next frame with the end position of the next frame, and fuse the current frame image and the next frame image by linear superposition.
[0046] Preferably, after fusing the current frame image and the next frame image based on the homography matrix to obtain a fused image, it further includes:
[0047] Perform three-dimensional reconstruction based on the fused image.
[0048] On the other hand, an oral image processing device includes:
[0049] A data acquisition module configured to acquire the acquired oral video stream and the corresponding gyroscope data;
[0050] A key frame image acquisition module configured to perform frame splitting on the oral video stream to obtain key frame images;
[0051] An image correction module configured to perform grid correction on the key frame images to obtain corrected images;
[0052] An image matching and alignment module configured to sequentially match the corrected images using the four-direction template matching method to obtain the current frame image and the next frame image that meet the corner information matching conditions, and perform alignment;
[0053] A corner information correction module configured to correct the corner information of the next frame image based on the gyroscope data;
[0054] A homography matrix acquisition module configured to acquire the homography matrix between the current frame image and the next frame image based on the corner information of the current frame image and the corrected corner information of the next frame image;
[0055] A fused image module configured to fuse the current frame image and the next frame image based on the homography matrix to obtain a fused image.
[0056] On the other hand, an oral image processing device includes:
[0057] An imaging device for collecting an oral video stream;
[0058] A gyroscope for collecting attitude information during movement;
[0059] A handheld part for moving the device to drive the imaging device and the gyroscope to move;
[0060] A control module connected to the imaging device and the gyroscope, configured to execute the oral image processing method described above.
[0061] The present invention has the following beneficial effects:
[0062] (1) The present invention can reconstruct an accurate and clear three-dimensional image through key frame image extraction, tooth instance segmentation, image correction, image alignment, corner information correction, image stitching, and image fusion, meeting the needs of users to independently observe teeth;
[0063] (2) The key frame image extraction method of the present invention can solve the problems that when the user pauses or moves too slowly during movement, multiple consecutive frames are repeated, increasing the duration of image stitching; and when the user moves too fast and loses key frames, the overlapping area between the front and rear frames is too small to extract key information for stitching;
[0064] (3) The oral video stream collected by the present invention is obtained by shooting deep into the oral cavity. The shooting angle of the teeth is very small during intraoral shooting, and there are reflective noise points. The instance segmentation method of the present invention can accurately segment each tooth to divide the ownership, extract the tooth contour information, and perform subsequent matching work based on the corner points of the contour;
[0065] (4) The present invention uses a four-direction template alignment method to align image frames, and alignment is only performed when the corner point information of the front and rear frames meets the matching conditions, ensuring the alignment quality;
[0066] (5) To prevent the corner point information extracted due to jitter during the user's movement from having too large a deviation, the present invention corrects the corner point information of the next frame based on gyroscope data, ensuring more accurate image stitching.
[0067] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the following described accompanying drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts. Brief Description of the Drawings
[0068] Figure 1 Flow chart of the method for processing oral images according to an embodiment of the present invention;
[0069] Figure 2 Flow chart of the key frame image extraction method according to an embodiment of the present invention;
[0070] Figure 3 Schematic diagram of stitching error caused by instance segmentation in the prior art;
[0071] Figure 4 Schematic diagram of the TeethInsReg model structure for instance segmentation according to an embodiment of the present invention;
[0072] Figure 5 Schematic diagram of the instance segmentation effect according to an embodiment of the present invention;
[0073] Figure 6 Schematic diagram of misjudgment of gums in instance segmentation in the prior art;
[0074] Figure 7 Schematic diagram of accurate gum segmentation according to an embodiment of the present invention;
[0075] Figure 8 Schematic diagram of the template cropped from the current frame in the four-direction matching method according to an embodiment of the present invention;
[0076] Figure 9 Schematic diagram of the position of the best template displayed in the four-direction matching method according to an embodiment of the present invention;
[0077] Figure 10 Schematic diagram of the process of obtaining the corner point position according to an embodiment of the present invention; wherein, (a) represents the camera shooting positions of the front and rear frames; (b) represents the positions of the front and rear frames after rotation;
[0078] Figure 11 Schematic diagram of the fusion of the direct superposition method in the prior art;
[0079] Figure 12 Schematic diagram of the fusion of the linear superposition method according to an embodiment of the present invention;
[0080] Figure 13 Structure block diagram of the device for processing oral images according to an embodiment of the present invention;
[0081] Figure 14 Schematic diagram of the structure of the three-dimensional reconstruction device according to an embodiment of the present invention. Detailed implementation manners
[0082] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention; obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0083] In the description of the present invention, it should be noted that the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, the element defined by the phrase "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0084] In the description of the present invention, it should be noted that the terms "upper", "lower", "inner", "outer", "top / bottom end", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0085] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "install", "be provided with", "sheath / connect", "connect", etc. should be understood in a broad sense. For example, "connect" can be a fixed connection, a detachable connection, or an integral connection, it can be a mechanical connection, an electrical connection, a direct connection, or an indirect connection through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0086] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the step identifiers S101, S102, S103, etc. are only for convenient expression and do not represent the execution order, and the corresponding execution order can be adjusted.
[0087] See Figure 1 As shown, a method for processing oral images of the present invention includes:
[0088] S101, obtaining the collected oral video stream and the corresponding gyroscope data;
[0089] S102, Split the oral video stream into frames to obtain key frame images;
[0090] S103, Based on the key frame images, perform dental instance segmentation to extract the dental contour information with the gingival margin area as the feature points;
[0091] S104, Perform mesh correction on the key frame images to obtain the corrected images;
[0092] S105, Use the four - azimuth template matching method to sequentially match the corrected images, obtain the current frame image and the next frame image that meet the corner point information matching conditions, and align them;
[0093] S106, Correct the corner point information of the next frame image based on the gyroscope data;
[0094] S107, Based on the corner point information of the current frame image and the corrected corner point information of the next frame image, obtain the homography matrix between the current frame image and the next frame image;
[0095] S108, Based on the homography matrix, fuse the current frame image and the next frame image to obtain the fused image;
[0096] S109, Perform 3D reconstruction based on the fused image.
[0097] In this embodiment, the execution subject of a method for processing oral images includes devices such as an MCU controller, which is not limited in this embodiment as long as it can execute the above - mentioned method.
[0098] In this embodiment, the collected oral video stream and the corresponding gyroscope data are the oral video stream and the collected gyroscope data taken in a specified order. Specifically, it can be to take pictures of the upper row of teeth on the outer side of the oral cavity from right to left, and then take pictures of the upper row of teeth on the inner side of the oral cavity from right to left to complete the complete oral image of the upper row of teeth. Then, complete the images of the lower row of teeth from the outer and inner sides of the lower row of the oral cavity in turn.
[0099] Traditional video stream image stitching uses frame extraction. For example, general video images are 25 - 30 frames per second, and one frame is extracted every 5 frames as a key frame. However, such frame extraction for subsequent stitching is relatively crude. The user may pause slightly during the movement, resulting in multiple consecutive frames being the same picture. In this way, multiple images will be repeated, increasing the duration of subsequent image stitching. Another situation is that the user moves too fast during the shooting, losing key frames, resulting in too small an overlapping area between the front and back frames and being unable to extract key information for stitching. To solve such problems, a high - quality image acquisition algorithm for key frames is proposed to achieve splitting the oral video stream into frames and obtaining key frame images. See Figure 2 As shown, the specific steps are as follows.
[0100] (1) First, disassemble each frame from the video stream, compare the similarity between two consecutive frames, and obtain the similarity evaluation index Similarity. Split the RGB channels of the two images respectively, and statistically calculate their histograms. Then calculate the similarity according to their respective channels. The similarity calculation method for a single channel is as follows:
[0101]
[0102] (2) First, disassemble each frame from the video stream, compare the similarity between two consecutive frames, and obtain the similarity evaluation index Similarity. Split the RGB channels of the two images respectively, and statistically calculate their histograms. Then calculate the similarity according to their respective channels. The similarity calculation method for a single channel is as follows:
[0103]
[0104] Among them, single is the similarity of the current channel, src i is the y-axis value corresponding to the x-axis subscript i of the histogram of the previous frame. dst i is the y-axis value corresponding to the x-axis subscript i of the histogram of the next frame.
[0105] (3) Calculate the average of the similarities of the three channels respectively to obtain the similarity Similarity of the whole image.
[0106] For example, when the image similarity Similarity > 0.9, it can be judged that the image similarities are too close, so the next frame is a non-key frame, and then the similarity of the frame after the next frame is judged. If 0.9 <= Similarity <= 0.75, it is judged that the next frame is a key frame, and the similarity between the current key frame and the next frame is judged. If Similarity < 0.75, it means that the key frame is lost, and subsequent frames are not judged.
[0107] In this embodiment, the obtained oral video stream is a video stream taken deep into the oral cavity, resulting in a very small shooting angle of the teeth and having reflective noise points. Since the white reflective noise points will be fixed on the tooth surface as the teeth change, the traditional stitching method (panoramic stitching) based on corner features will be stitched incorrectly, and the corner features will be concentrated at the reflective points. For details, see Figure 3 as shown.
[0108] To avoid this noise, the gum edge is used as a feature point for image stitching. Use the TeethInsReg algorithm to perform instance segmentation on the oral cavity with the gum edge as the feature point. The effect diagram is shown in Figure 4 as shown.
[0109] The TeethInsReg model structure of this embodiment is shown in Figure 5 as follows.
[0110] The TeethInsReg model includes an encoding module encode, a feature pyramid network module FPN, a decoding module decode, a classification module category, and a mask module mask that are connected in sequence.
[0111] In the encode part, resnet18 is used as the backbone, and features of different scales are extracted through 4 times of downsampling to prevent missed detections due to too small or too large features.
[0112] The FPN feature pyramid is used to fuse and enhance the features, and shared feature maps are used to reduce the computational amount and improve the computational accuracy.
[0113] In the decode part, different-scale upsampling is performed according to the features of different scales introduced by the FPN to restore them to the size of the original image. Different from the decode part of Unet that upsamples step by step from the bottom layer and connects the original features of different scales, different-scale feature layers are upsampled in different proportions to restore them to the size of the original image, and layer connection is directly performed in a layer connection manner. This further reduces the computational amount and speeds up the calculation speed.
[0114] Category and Mask calculate the category (tooth or gum) and the corresponding segmentation instance mask respectively.
[0115] The mask is binarized to extract each mask to obtain the precise contour information of each tooth.
[0116] See Figure 6 for a schematic diagram of misjudgment of gums in instance segmentation of the prior art; Figure 7 for a schematic diagram of accurate gum segmentation in the embodiment of the present invention.
[0117] With the high-precision instance segmentation of TeethInsReg, each tooth can be accurately segmented to determine the ownership. Therefore, the tooth contour information can be extracted, and subsequent matching work can be carried out according to the corner points of the contour.
[0118] In this embodiment, the grid correction method is used to correct the distortion of the image. The image is corrected for distortion to solve the image distortion caused by the rotation of the camera, so that the stitched image is more accurate.
[0119] In this embodiment, the four-direction template matching method is used to match the corrected image in sequence to obtain the current frame image and the next frame image that meet the corner point information matching conditions, specifically including:
[0120] Divide the current frame image into four template blocks, and perform template matching on each of the four template blocks with the next frame image to obtain the block with the highest matching rate with the next frame;
[0121] If the highest matching rate is greater than the set matching threshold, it means that there is an overlapping area between the current frame and the next frame, and subsequent corner matching can be performed; otherwise, match the current frame with the frame after the next frame.
[0122] Specifically, refer to Figure 8 and Figure 9 As shown, use the matchTemplate method of opencv to divide the current image into four template blocks (such as the four template blocks of different colors in the current frame), and perform template matching on each template block with the next frame image to obtain the block with the highest matching rate with the next frame. If the matching rate of the highest block is high, that is, greater than the set threshold, it means that there is an overlapping area between the current frame and the next frame and subsequent corner matching can be performed.
[0123] Align the tooth region features, and match the corner feature information of this frame with the corner information of the next frame. If the number of matching feature points between the current frame and the next frame exceeds a preset number such as 11, it is considered that these two frames meet the feature alignment condition.
[0124] Specifically, use the SIFI algorithm to extract the corresponding 11 corner information (the 11 corner information with the highest matching degree) (pixel x, y value information) of the current frame and the subsequent frame. And use the RANSAC matching algorithm for feature matching, as shown in the following example.
[0125] sift = cv2.SIFT_create()
[0126] # gray1 and gary2 are the front and back frame images
[0127] kp1, des1 = sift.detectAndCompute(gray1, None)
[0128] kp2, des2 = sift.detectAndCompute(gray2, None)
[0129] index_params = dict(algorithm = 0, trees = 5)
[0130] # Set 11 matching points
[0131] search_params = dict(checks = 11)
[0132] flann = cv2.FlannBasedMatcher(index_params, search_params)
[0133] matches = flann.knnMatch(des1, des2, 2)
[0134] Correct the corner point information of the next frame of image based on the gyroscope data, specifically including:
[0135] Extract the first corner point information based on the contour information of the next frame of image, calculate the second corner point information of the next frame of image based on the gyroscope data, compare the feature points of the first corner point information and the second corner point information in sequence, and replace the feature points of the first corner point information with a deviation greater than the preset pixel with the feature points of the second corner point information to obtain the corrected corner point information of the next frame of image.
[0136] Calculate the second corner point information of the next frame of image based on the gyroscope data, specifically including:
[0137] According to the attitude information of the gyroscope, calculate the actual rotation angle and speed between the current frame and the next frame of image; the attitude information includes the roll angle, yaw angle, and pitch angle of the current frame and the next frame of image;
[0138] Based on the roll angle, yaw angle, and pitch angle of the current frame and the next frame of image, calculate the rotation angle and rotation speed from the current frame of image to the next frame of image;
[0139] Calculate the corner point information of the next frame of image based on the rotation angle and rotation speed.
[0140] Calculate the homography matrix of the image based on the image corrected by the grid correction method and combined with the offset of the gyroscope on the x, y, and z axes. The specific method is as follows.
[0141] (1) Obtain the next frame of data of the gyroscope, and obtain the plane movement information of the current video stream according to the gyroscope, including the roll angle r in the left - right direction, the pitch angle p in the up - down direction, and the yaw angle y in the front - back direction.
[0142]
[0143] Among them, gx is the angular velocity of the roll angle, gy is the angular velocity of the pitch angle, and gz is the angular velocity of the yaw angle. r n is the roll angle value of the previous key frame, r n+1 is the roll angle data of the next frame. p n is the pitch angle value of the previous key frame, p n+1 is the pitch angle data of the next frame. y n is the yaw angle value of the previous key frame, y n+1It is the yaw angle data of the subsequent frame.
[0144] (2) According to the attitude information (roll angle, yaw angle, and pitch angle) provided by the gyroscope, calculate the actual rotation angle and offset between adjacent images, and calculate the corner positions of the next frame based on the actual rotation angle and offset.
[0145] Assume the rotation angle is θ, the angular velocity is w, and the offsets of the gyroscope are x, y, and z.
[0146] (a) First, find the translation direction coordinates. See Figure 10 As shown in (a), the coordinates of corner point A are (a, b), the coordinate origin is the upper left corner of the photo frame, and the displacements are x and y. The lower left frame is the position where the previous frame was taken by the camera, and the upper right frame is the position where the next frame will be taken. Then the coordinates of corner point A' are (a - x, b - y).
[0147] (b) See Figure 10 As shown in (b), calculate the camera rotation angle and obtain the rotated coordinates. Assume the coordinates of A are (a, b) and the coordinates of the next frame are A'(a', b'). The calculation formula is as follows:
[0148]
[0149] (c) Substitute the 11 corner points into steps (a) and (b) respectively, and the corner positions of the next frame of the camera can be obtained.
[0150] (3) Compare the information of the preset number of corner points (such as 11 corner points) calculated by the gyroscope for the next frame and the 11 corner points extracted by SIFI one by one. If the relative position deviation is greater than the preset number of pixels (such as 50 pixels), replace the corner point information of SIFI with the corner point information calculated by the gyroscope. To correct the errors caused by jitter during the movement.
[0151] (4) According to the corner point information of the current frame and the corrected corner point information of the next frame, use the RANSAC matching algorithm to calculate the homography matrix H2. The example code is as follows:
[0152] pts1 = np.float32([kp1[m.queryIdx].pt for min matches]).reshape(-1, 1, 2)
[0153] pts2 = np.float32([kp2[m.trainIdx].pt for min matches]).reshape(-1, 1, 2)
[0154] H2, mask = cv2.findHomography(pts1, pts2, cv2.RANSAC, 1.0)
[0155] Further, based on the homography matrix, the current frame image and the next frame image are fused to obtain a fused image, specifically as follows:
[0156] Obtain the image information of the current frame and the image information of the next frame, and based on the homography matrix, obtain the overlapping position of the two images;
[0157] Based on the overlapping position, obtain the stitching size through the perspective transformation function;
[0158] Correspond the initial position of the current frame with the initial position of the stitching size, and correspond the end position of the next frame with the end position of the next frame, and fuse the current frame image and the next frame image through linear superposition.
[0159] Use the homography matrix H2 for the image of the next frame, and use the warpPerspective method in opencv to stitch them together. The example code is as follows.
[0160] # Stitch the images according to the size of the previous frame image
[0161] result = cv2.warpPerspective(gray1, H2, (gray1.shape[1] + gray1.shape[1],
[0162] gray2.shape[0])
[0163] result[0:img2.shape[0], 0:img2.shape[1]] = img2
[0164] In this embodiment, the linear superposition technology is used to linearly superpose the previous and next frames to eliminate the segmentation line of the edge stitching caused by different light intensities. Among them, Figure 11 is the fusion schematic diagram of the direct superposition method of the prior art; Figure 12 is the fusion schematic diagram of the linear superposition method of the embodiment of the present invention.
[0165] The method for three-dimensional reconstruction of the fused image is as follows.
[0166] (1) Back-project the fused image into a relatively independent three-dimensional feature space.
[0167] (2) Directly predict the TSDF (Truncated Signed Distance Function, a common method for calculating the implicit potential surface in 3D reconstruction) of the segment in the three-dimensional feature space using sparse three-dimensional convolution.
[0168] (3) Feed the TSDFs of all segments into the GRU module (GRU is a variant of the long short-term memory network, which is also proposed to solve problems such as long-term memory and gradients in backpropagation) to make the reconstruction between segments consistent.
[0169] (4) After feeding into the multi-layer perceptron, perform image restoration and fusion to predict a complete model.
[0170] See Figure 13 As shown, a processing device for oral images includes:
[0171] A data acquisition module 1301, configured to acquire the acquired oral video images and corresponding gyroscope data;
[0172] A key frame image acquisition module 1302, configured to perform frame splitting on the oral video images to obtain key frame images;
[0173] An instance segmentation module 1303, configured to perform tooth instance segmentation based on the key frame images and extract tooth contour information with the gingival edge area as the feature points;
[0174] An image correction module 1304, configured to perform grid correction on the key frame images to obtain corrected images;
[0175] An image matching and alignment module 1305, configured to use the four-direction template matching method to sequentially match the corrected images to obtain the current frame image and the next frame image that meet the corner point information matching conditions, and perform alignment;
[0176] A corner point information correction module 1306, configured to correct the corner point information of the next frame image based on the gyroscope data;
[0177] A homography matrix acquisition module 1307, configured to acquire the homography matrix between the current frame image and the next frame image based on the corner point information of the current frame image and the corrected corner point information of the next frame image;
[0178] A fused image module 1308, configured to fuse the current frame image and the next frame image based on the homography matrix to obtain a fused image;
[0179] A 3D reconstruction module 1309, configured to perform 3D reconstruction based on the fused image.
[0180] The specific implementation of a processing device for oral images is the same as that of a processing method for oral images, and this embodiment will not be repeated here.
[0181] See Figure 14 As shown, a processing device for oral images in this embodiment includes:
[0182] The imaging device 1401 is used to collect oral video images;
[0183] The gyroscope 1402 is used to collect attitude information during movement;
[0184] The handheld part 1403 is for a mobile device to drive the imaging device 1401 and the gyroscope 1402 to move;
[0185] The control module 1404 is connected to the imaging device 1401 and the gyroscope 1402 and is configured to execute the method for processing oral images.
[0186] Specifically, the oral image processing device can be integrated into devices such as a water flosser or an endoscope for users to hold and use. The specific usage method is as follows:
[0187] Turn on the switch 1405 of the imaging device 1401, and shoot oral video images in a specified order to obtain real-time video stream images and real-time gyroscope 1402 data. The handheld device three-dimensional reconstruction device first records the upper row of teeth on the outer side of the oral cavity from right to left; then records the upper row of teeth from right to left from the inner side of the oral cavity to complete the complete oral image of the upper row of teeth; then completes the outer and inner oral cavity images of the lower row of teeth in sequence; finally, generates a visual three-dimensional reconstructed oral image to facilitate the user to observe the teeth and timely discover dental problems.
[0188] As mentioned above, it is only a preferred specific embodiment of the present invention; however, the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its improved concept, makes equivalent replacements or changes, and should be covered by the protection scope of the present invention.
Claims
1. A method for processing oral images, characterized in that, it includes: Obtain the collected oral video stream and the corresponding gyroscope data; Perform frame splitting on the oral video stream to obtain key frame images; Perform grid correction on the key frame images to obtain corrected images; Use the four-direction template matching method to sequentially match the corrected images, obtain the current frame image and the next frame image that meet the corner information matching conditions, and perform alignment; Correct the corner information of the next frame image based on the gyroscope data; Based on the corner information of the current frame image and the corrected corner information of the next frame image, obtain the homography matrix between the current frame image and the next frame image; Based on the homography matrix, fuse the current frame image and the next frame image to obtain a fused image.
2. The method for processing oral images according to claim 1, characterized in that, Performing frame splitting on the oral video stream to obtain key frame images specifically includes: S1021, disassemble each frame from the oral video stream; S1022, calculate and compare the similarity between the current frame and the next frame; S1023, if the similarity is less than the key frame minimum threshold, it means the key frame is lost, and subsequent frames are not judged; S1024, if the similarity is greater than the key frame maximum threshold, take the next next frame as the next frame, and return to S1022; S1025, if the similarity is less than or equal to the key frame maximum threshold and greater than or equal to the key frame minimum threshold, then take the next frame as the key frame image and add it to the key frame list; S1026, judge whether there are still frames not compared. If so, take the next frame as the current frame, take the next next frame as the next frame, and return to S1022.
3. The method for processing oral images according to claim 2, characterized in that, The calculation method of the similarity between the current frame and the next frame includes: Perform single-channel similarity calculation on the current frame and the next frame; Average the similarities of the three single channels to obtain the similarity between the current frame and the next frame.
4. The method for processing oral images according to claim 1, characterized in that, After performing frame splitting on the oral video stream to obtain key frame images, it further includes: performing tooth instance segmentation based on the key frame images, and extracting tooth contour information with the gingival edge area as the feature point; specifically as follows: Use the TeethInsReg model to perform tooth instance segmentation and extract tooth contour information with the gingival edge area as the feature point; the TeethInsReg model includes an encoding module encode, a feature pyramid network module FPN, a decoding module decode, a classification module category, and a mask module mask connected in sequence; The input of the TeethInsReg model is the key frame image, and the output is the category and the mask of the corresponding segmentation instance; Perform binarization on the mask to obtain the contour information of each tooth.
5. The method for processing oral images according to claim 1, characterized in that, Using the four-direction template matching method to sequentially match the corrected images to obtain the current frame image and the next frame image that meet the corner information matching conditions specifically includes: Divide the current frame image into four template blocks, and perform template matching on each of the four template blocks with the next frame image to obtain the block with the highest matching rate with the next frame; If the highest matching rate is greater than the set matching threshold, it indicates that there is an overlapping area between the current frame and the next frame, and subsequent corner matching can be performed; otherwise, match the current frame with the frame after the next frame.
6. The method for processing oral images according to claim 5, wherein, Align the current frame and the next frame images, specifically including: Match the corner information of the current frame with the corner information of the next frame; If the number of matching corner feature points between the current frame and the next frame exceeds a preset number, align the current frame and the next frame images.
7. The method for processing oral images according to claim 1, wherein, Correct the corner information of the next frame image based on the gyroscope data, specifically including: Extract the first corner information based on the contour information of the next frame image, calculate the second corner information of the next frame image based on the gyroscope data, compare the feature points of the first corner information and the second corner information in sequence, and replace the feature points of the first corner information with a deviation greater than the preset pixel with the feature points of the second corner information to obtain the corrected corner information of the next frame image.
8. The method for processing oral images according to claim 7, wherein, Calculate the second corner information of the next frame image based on the gyroscope data, specifically including: According to the attitude information of the gyroscope, calculate the actual rotation angle and speed between the current frame and the next frame image; the attitude information includes the roll angle, yaw angle, and pitch angle of the current frame and the next frame image; Based on the roll angle, yaw angle, and pitch angle of the current frame and the next frame image, calculate the rotation angle and rotation speed from the current frame image to the next frame image; Calculate the corner information of the next frame image based on the rotation angle and rotation speed.
9. The method for processing oral images according to claim 1, wherein, The corner information of the current frame image is extracted based on the contour information of the current frame image.
10. The method for processing oral images according to claim 1, wherein, Fuse the current frame image and the next frame image based on the homography matrix to obtain a fused image, specifically including: Obtain the image information of the current frame and the image information of the next frame, and based on the homography matrix, obtain the overlapping position of the two images; Based on the overlapping position, obtain the stitching size; Correspond the initial position of the current frame with the initial position of the stitching size, and correspond the end position of the next frame with the end position of the next frame, and fuse the current frame image and the next frame image by linear superposition.
11. The method for processing oral images according to claim 1, wherein, After fusing the current frame image and the next frame image based on the homography matrix to obtain a fused image, it further includes: Perform three-dimensional reconstruction based on the fused image.
12. An apparatus for processing oral images, wherein, comprising: A data acquisition module configured to acquire the acquired oral video stream and the corresponding gyroscope data; The key frame image acquisition module is configured to split frames of the oral cavity video stream to obtain key frame images; The image correction module is configured to perform grid correction on the key frame images to obtain corrected images; The image matching and alignment module is configured to sequentially perform matching on the corrected images using the four-direction template matching method to obtain the current frame image and the next frame image that meet the corner information matching conditions, and perform alignment; The corner information correction module is configured to correct the corner information of the next frame image based on the gyroscope data; The homography matrix acquisition module is configured to acquire the homography matrix between the current frame image and the next frame image based on the corner information of the current frame image and the corrected corner information of the next frame image; The fused image module is configured to fuse the current frame image and the next frame image based on the homography matrix to obtain a fused image.
13. A processing device for oral cavity images, characterized in that, it includes: A camera device for collecting the oral cavity video stream; A gyroscope for collecting attitude information during movement; A handheld part for moving the device to drive the camera device and the gyroscope to move; A control module connected to the camera device and the gyroscope, and configured to execute the method for processing oral cavity images according to any one of claims 1 to 11.