Tooth recognition method and apparatus, storage medium, and electronic device

By performing low-resolution image recognition and high-resolution local image processing on the original oral three-dimensional image, combined with image classification model and dynamic scaling technology, the problem of inaccurate tooth recognition results is solved and the accuracy of tooth 3D image construction is improved.

WO2025139463A1PCT designated stage expired Publication Date: 2025-07-03SHANGHAI EA MEDICAL INSTR CO LTD

Patent Information

Application Number
PCT/CN2024/132758
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-11-18
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The prior art When constructing 3D images of teeth, the dental recognition results are inaccurate, which may result in image information loss due to the conversion of oral images to low resolution images.

Method used

By performing low-resolution image recognition on the original oral three-dimensional image, the dental position information is determined, and then the key position information is identified and high-resolution local images are obtained. Combined with image classification model and dynamic scaling technology, the accuracy of dental recognition is improved.

Benefits of technology

It improves the accuracy of teeth recognition results and ensures the accuracy of subsequent teeth 3D image construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024132758_03072025_PF_FP_ABST
    Figure CN2024132758_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present specification discloses a tooth recognition method and apparatus, a storage medium, and an electronic device. The method comprises: firstly performing recognition on a low-resolution three-dimensional image of a target area; obtaining dentition position information of each tooth in the target area; determining key position information of the teeth from the dentition position information, so as to obtain a high-resolution local image corresponding to the key position information; performing recognition on the high-resolution local image to obtain key tooth recognition results; and on the basis of the dentition position information from the low-resolution image and the key tooth recognition results from the high-resolution image of key positions, obtaining a tooth recognition result for an original three-dimensional oral image. Recognition is firstly performed on the low-resolution three-dimensional image to obtain the dentition position information, and then tooth recognition is performed on the high-resolution three-dimensional image corresponding to the key position information, thereby improving the accuracy of tooth recognition results.
Need to check novelty before this filing date? Find Prior Art

Description

Tooth identification method, device, storage medium and electronic device

[0001] This application claims priority to the invention patent application filed with the State Intellectual Property Office on December 29, 2023, with application number 202311868228.8 and invention name “A tooth detection method, device, storage medium and electronic device”, the entire contents of which are incorporated by reference in this application. Technical Field

[0002] This specification relates to the field of medical image processing, and in particular to a tooth recognition method, device, storage medium, and electronic device. Background Art

[0003] With the development of medical image processing technology, doctors can use medical images obtained by various medical imaging devices to predict the diseases that users may have. For example, based on oral CBCT (Cone Beam Computer Tomography) images, a 3D image of the teeth can be constructed to determine whether there are any problems with the user's teeth. When constructing a 3D image of the teeth, each tooth needs to be identified to determine the position of each tooth, and then constructed based on this position. Therefore, the accuracy of the tooth recognition result affects the construction result of the 3D image of the teeth. The existing method may be to first obtain the area of ​​the teeth in the oral CBCT image, convert the image size of the area to the size of the input image of the tooth detection model, and input the converted image into the tooth recognition model to obtain the recognition result. However, when converting the image, a lot of image information may be lost, resulting in inaccurate recognition results.

[0004] Based on this, this specification provides a tooth identification method. Summary of the Invention

[0005] This specification provides a tooth recognition method, device, storage medium and electronic device to at least partially solve the above-mentioned problems existing in the prior art.

[0006] The present specification provides a tooth recognition method, comprising: determining a low-resolution three-dimensional image including a target area based on an original oral three-dimensional image, wherein the target area includes at least one of several teeth, gums, dentition, and jaw in the original oral three-dimensional image; performing image recognition on the low-resolution three-dimensional image to obtain the patient's dentition position information; determining key position information of the teeth in the original oral three-dimensional image based on the dentition position information; obtaining a local image corresponding to the key position information based on the key position information of the teeth, wherein the local image is a high-resolution three-dimensional image; performing image recognition on the high-resolution three-dimensional image to obtain a key tooth recognition result of the local image; and determining a tooth recognition result of the original oral three-dimensional image based on the dentition position information and at least one key tooth recognition result.

[0007] The above method first recognizes a low-resolution three-dimensional image of the target area to obtain dentition position information for each tooth in the target area. Key tooth positions are then determined from this dentition position information. A high-resolution local image corresponding to the key position information is then obtained. This high-resolution local image is then recognized to obtain key tooth recognition results. Based on the dentition position information in the low-resolution image and the key tooth recognition results from the high-resolution image of the key positions, tooth recognition results for the original three-dimensional oral cavity image are obtained. By first recognizing the low-resolution three-dimensional image to obtain dentition position information and then performing tooth recognition on the high-resolution three-dimensional image of the key position information, the accuracy of tooth recognition results is improved.

[0008] This specification provides a tooth identification device, the device comprising:

[0009] a low-resolution three-dimensional image acquisition module, configured to determine, based on the original three-dimensional oral image, a low-resolution three-dimensional image including a target area, wherein the target area includes at least one of a plurality of teeth, gums, dentition, and jaw in the original three-dimensional oral image;

[0010] a dentition position information determination module, configured to perform image recognition on the low-resolution three-dimensional image to obtain the patient's dentition position information;

[0011] a key position information determination module, configured to determine key position information of teeth in the original oral three-dimensional image based on the dentition position information;

[0012] A local image determination module is used to obtain a local image corresponding to the key position information according to the key position information of the tooth, wherein the local image is a high-resolution three-dimensional image;

[0013] a first recognition result determination module, configured to perform image recognition on the high-resolution three-dimensional image to obtain a key tooth recognition result of the partial image;

[0014] The second recognition result determination module is used to determine the tooth recognition result of the original oral three-dimensional image based on the dentition position information and at least one key tooth recognition result.

[0015] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned tooth recognition method is implemented.

[0016] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned tooth recognition method when executing the computer program. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] FIG1 is a schematic diagram of a process of a tooth identification method provided in this specification;

[0018] FIG2 is a schematic diagram of target tooth recognition results provided in this specification;

[0019] FIG3 is a schematic diagram of the intersection of adjacent detection frames and key frames provided in this specification;

[0020] FIG4 is a schematic diagram of the minimum external frame provided in this specification;

[0021] FIG5 is a schematic diagram of key tooth recognition results provided in this specification;

[0022] FIG6 is a schematic diagram of the detection frame fusion provided in this specification;

[0023] FIG7 is a schematic diagram of tooth recognition results of the original oral 3D image provided in this specification;

[0024] FIG8 is a schematic diagram of dimension crossing provided in this specification;

[0025] FIG9 is a schematic diagram of a negative sample provided in this specification;

[0026] FIG10 is a schematic diagram of another negative sample provided in this specification;

[0027] FIG11 is a flow chart of a tooth detection method provided in this specification;

[0028] FIG12 is a schematic diagram of a negative sample provided in this specification;

[0029] FIG13 is a schematic diagram of an electronic device provided in this specification. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solutions and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings.

[0031] Constructing a 3D image of teeth requires identifying the teeth, so accurate tooth recognition results have a huge impact on the subsequent construction of the 3D image of the teeth. Since the amount of data in a 3D image is large and the video memory of the device that constructs the 3D image is limited, when identifying teeth, one possible way may be to convert the oral image including the teeth into a low-resolution image and then perform tooth recognition. This conversion operation causes the oral image to lose a lot of image information, and thus the tooth recognition result is inaccurate, affecting the subsequent construction of the 3D image of the teeth. Therefore, this specification provides a tooth recognition method. The execution subject of this specification includes a computing device deployed for obtaining tooth recognition results, such as a server, such as a desktop computer, a laptop computer, etc., or other electronic devices that can be deployed for obtaining tooth recognition results. It can also be a computing device that can train a model, etc. This specification does not impose any restrictions on this. For the sake of ease of description, the tooth recognition method provided in this specification is described below with only the electronic device as the execution subject.

[0032] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0033] FIG1 is a flow chart of a tooth identification method provided in this specification, which specifically includes the following steps:

[0034] S100: Determine a low-resolution three-dimensional image including a target area based on an original three-dimensional oral image, where the target area includes at least one of a plurality of teeth, gums, dentition, and jaws in the original three-dimensional oral image.

[0035] In one or more embodiments of this specification, the electronic device performs tooth recognition based on an oral image of a user (patient). Therefore, the electronic device may first obtain an original 3D oral image. Specifically, the image may be obtained using various medical imaging devices. The medical imaging device obtains the original 3D oral image of the user and transmits the original 3D oral image to the electronic device, which then receives the original 3D oral image transmitted by the medical imaging device.

[0036] Because the 3D tooth recognition method based on deep learning may need to convert the target area into a fixed size, the image resolution is reduced. Similarly, in this specification, it is also necessary to determine a low-resolution three-dimensional image including the target area based on the original oral three-dimensional image. The target area includes at least one of several teeth, gums, dentition, and jaws in the original oral three-dimensional image. The target area is selected based on business needs. For example, if the business need is to identify the user's central incisor and lateral incisor so as to subsequently construct a 3D image of the central incisor and lateral incisor, then the target area is the area in the original oral image that includes the central incisor and lateral incisor.

[0037] Specifically, the electronic device can directly perform size conversion on the original oral three-dimensional image to obtain a low-resolution three-dimensional image of a preset first size. The preset first size can be set as needed, and this specification does not limit this. The original oral three-dimensional image can also be segmented using an image segmentation model to obtain a target three-dimensional image including a target area, and then the target three-dimensional image is size-converted to obtain a low-resolution three-dimensional image of a preset first size, wherein the image segmentation model has the ability to segment the target area. If the size can be (128, 128, 128), the image segmentation model can be a coding-decoding structure network such as U-Net, V-Net, etc. to achieve fully automated target area extraction. This specification does not limit the type of image segmentation model, as long as the image segmentation model can output the target area image.

[0038] It should be noted that the low resolution mentioned in this specification can be determined based on the size, resolution, image quality, etc. of the obtained three-dimensional image, or it can be determined by at least one factor such as the requirements and computing power for image recognition, the image recognition model, the accuracy of the original image and the accuracy of recognition. A low resolution value range can be determined based on a combination of multiple factors, and this specification does not make any specific restrictions.

[0039] The high resolution mentioned in this specification can be determined relative to the low resolution, or it can be determined based on the size, resolution, image quality, etc. of the acquired medical three-dimensional image, or it can be determined based on computing power and recognition requirements. A high resolution value range can be determined based on a combination of multiple factors, and this specification does not make any specific restrictions.

[0040] S102: Performing image recognition on the low-resolution three-dimensional image to obtain the patient's dentition position information.

[0041] Specifically, the electronic device may input the low-resolution 3D image into a first tooth recognition model, and obtain a target tooth recognition result output by the first tooth recognition model. Because the first tooth recognition model recognizes each tooth included in the low-resolution 3D image, the obtained target tooth recognition result includes a detection frame for each tooth in the target area. Simultaneously, the first tooth recognition model also outputs position information for each tooth detection frame, including the 3D coordinates and 3D dimensions of the detection frame.

[0042] Because the target tooth recognition result includes the position information of each tooth's detection frame, the electronic device can determine the patient's dentition position information based on the position information of each tooth's detection frame. Specifically, the electronic device can determine the patient's dentition position information based on the target tooth recognition result. Denture position information includes the position of each tooth relative to the other teeth, and dentition includes both upper and lower dentition. This specification does not limit the type of the first tooth recognition model, such as RetinaNet3D.

[0043] FIG2 is a schematic diagram of the target tooth recognition result provided in this specification, as shown in FIG2 .

[0044] Detection frame 1 is the detection result of the upper dentition, detection frame 2 is the detection result of the lower dentition, and detection frame 3 is the real frame where no teeth are detected. For intuitive understanding, the tooth instance segmentation annotation is also displayed.

[0045] S104: Determine key position information of teeth in the original oral 3D image based on the dentition position information.

[0046] Because the target tooth recognition result is based on a low-resolution 3D image, a significant amount of image information is lost during the resizing process, leading to inaccurate target tooth recognition results. For example, the detection frame may include more than one complete tooth, or the detection frame may include an incomplete tooth. To improve the accuracy of tooth recognition results, the electronic device can determine the key position information of the teeth in the original 3D oral image so that the correct detection frame can be obtained based on this key position information during the subsequent second recognition.

[0047] In one possible implementation, since the detection frames output by the first tooth recognition model are unordered, the electronic device can sort the detection frames of each tooth based on the dentition position information to obtain sorted detection frames. The sorted detection frames are consistent with the normal order of the teeth, as shown in Figure 2, which shows the sorted detection frames from left to right. The electronic device can then determine the key teeth and their key position information within the sorted detection frames.

[0048] Key teeth may be teeth with a low accuracy rate in the first recognition. Key teeth may include at least one of the following: teeth corresponding to a preset tooth number, teeth corresponding to the classification results of the sorted detection frame based on the image classification model, and teeth corresponding to a preset tooth category. This specification does not limit the specific category of key teeth, nor does it limit the specific method for determining key teeth. For example, based on experience, teeth with a high error rate may be selected as key teeth. Specifically, the teeth corresponding to the sorted detection frame may be numbered and teeth with preset tooth numbers may be selected as key teeth. Alternatively, teeth of categories such as incisors and canines may be determined as key teeth. This is because the teeth in the incisor and canine positions are densely distributed, and in the first stage of recognition, there is a high probability of incorrect target tooth recognition results.

[0049] If the electronic device determines the key teeth based on an image classification model, then, for each detection frame, the electronic device may input the image of the detection frame into the image classification model, and obtain a classification result for the image of the detection frame as output by the image classification model. The image of the detection frame may be an image of the detection frame area of ​​the original 3D oral cavity image. The image quality of the original 3D oral cavity image is higher than that of a low-resolution 3D image. Therefore, the classification result obtained using the original 3D oral cavity image is more accurate.

[0050] If the classification result indicates that the target tooth in the detection frame is not correctly enclosed, the target tooth is determined to be a key tooth, the detection frame is determined to be a key frame, and the position information of the key frame is determined to be key position information. It should be understood that correctly enclosed means that each detection frame only includes a complete tooth, and the target tooth refers to the tooth included in the detection frame. In other words, the image classification model determines whether the detection frame of each tooth in the target tooth recognition result is a correct detection frame. If not, it is determined to be a key tooth, so that the correct recognition result of the key tooth can be obtained later.

[0051] S106: Obtaining a local image corresponding to the key position information according to the key position information of the tooth, wherein the local image is a high-resolution three-dimensional image.

[0052] In order to improve the accuracy of key tooth recognition results, the key tooth can be placed at the center of the local image.

[0053] Specifically, for each key frame, the electronic device may determine two adjacent detection frames within the sorted detection frames to obtain adjacent detection frames. When the adjacent detection frames intersect the key frame, or when the adjacent detection frames intersect the key frame after being translated by a preset distance in a preset dimension, the electronic device determines the minimum bounding box between the adjacent detection frames and the key frame as a first recognition region containing the key frame. Subsequently, based on the first recognition region and the key position information of the tooth, a partial image corresponding to the key position information is obtained.

[0054] FIG3 is a schematic diagram of the intersection of adjacent detection frames and key frames provided in this specification, as shown in FIG3 .

[0055] Coarse-outlined frames 1 and 2 are key frames. Adjacent detection frames include adjacent detection frames 1 and 2. The dashed line represents the moved adjacent detection frames. If the default dimension is the X-axis, adjacent detection frame 1 can be translated along the X-axis toward coarse-outlined frame 1 by its own height before intersecting with coarse-outlined frame 1. If the default dimension is the Y-axis, adjacent detection frame 2 can be translated along the Y-axis toward coarse-outlined frame 2 by its own width before intersecting with coarse-outlined frame 2.

[0056] FIG4 is a schematic diagram of the minimum external frame provided in this specification, as shown in FIG4 .

[0057] Detection frame 4 in Figure 4 is the selected key frame. The two adjacent detection frames adjacent to detection frame 4 are detection frames 5 and 6. The minimum bounding box target-box shared by detection frames 4, 5, and 6 is determined as the first recognition area. The purpose of doing so is to include the recognition target of the second stage as much as possible within the first recognition area.

[0058] The purpose of determining the minimum bounding box is to ensure that the key box is centered in the partial image of the second tooth recognition model during the second recognition, thereby improving the accuracy of the second recognition result. The minimum bounding box between an adjacent detection box and the key box is selected as the first recognition area only when the adjacent detection box intersects the key box or intersects the key box after being translated a preset distance. This is because if there are missing teeth between the adjacent detection box and the key box, the key box and the adjacent detection box will also be adjacent but distant. In this case, the key box will not be located at the center of the first recognition area, which is not conducive to improving the accuracy of the second recognition.

[0059] When obtaining a local image corresponding to key position information, the electronic device can determine the dynamic zoom ratio based on the preset first size of the second identification area, the size of the first identification area and the image size of the original oral three-dimensional image, and determine the first center position information of the first identification area based on the size of the first identification area.

[0060] For example, in the size of the first recognition area, the minimum and maximum coordinates in the X direction are 10 and 50 respectively, the minimum and maximum coordinates in the Y direction are 20 and 80 respectively, and the minimum and maximum coordinates in the Z direction are 30 and 90 respectively. Then, the coordinate of the center point in the X direction is (10+50) / 2=30, the coordinate of the center point in the Y direction is (20+80) / 2=50, and the coordinate of the center point in the Z direction is (30+90) / 2=60. Therefore, the first center position information is (30, 50, 60).

[0061] Afterwards, the original oral three-dimensional image is scaled according to the dynamic scaling ratio to obtain a scaled oral three-dimensional image. This is because for some original oral three-dimensional images with higher resolutions, the size of the first recognition area obtained is larger and will not be completely contained in the second recognition area of ​​the preset first size, which means that the target to be identified for the second time (key teeth) is likely not to be completely contained in the second recognition area, resulting in inaccurate key tooth recognition results. For some original oral three-dimensional images with lower resolutions, the size of one dimension may be smaller than the preset first size and cannot accommodate the second recognition area. It can also be understood that the dimension of the low-resolution image is too low, and the dimension can be increased, specifically by enlarging the low-resolution image to obtain more image information. Therefore, in order to ensure that the second recognition area can completely contain the first recognition area and can be contained by the original oral three-dimensional image, it is necessary to dynamically adjust the size of the original oral three-dimensional image according to the size of the original oral three-dimensional image and the size of the first recognition area.

[0062] At the same time, the electronic device may also determine the second center position information of the scaled first recognition area based on the dynamic zoom ratio and the first center position information.

[0063] Continuing with the above example, the first center position information is (30, 50, 60). If the dynamic scaling ratio is 0.1, the second center position information is (3, 5, 6).

[0064] Next, in the scaled three-dimensional oral cavity image, the second identification area is determined based on the second center position information and the preset first size.

[0065] For example, if the preset first size is (128, 128, 128), then in the scaled three-dimensional oral image, a region with a size of (128, 128, 128) centered on the second center position information is determined as the second identification region.

[0066] Finally, the oral cavity 3D image including the second identification area is determined as the partial image corresponding to the key location information. The electronic device can then segment the scaled 3D image based on the second identification area to obtain a partial image including only the second identification area. The partial image is a high-resolution 3D image. Although the resolution of the partial image is high, the data volume of the partial image is not higher than that of the original oral cavity 3D image because the partial image only contains a portion of the original oral cavity 3D image.

[0067] S108: Perform image recognition on the high-resolution three-dimensional image to obtain a key tooth recognition result of the partial image.

[0068] Specifically, the high-resolution three-dimensional image is input into a second tooth recognition model, and the second tooth recognition model outputs a local tooth recognition result for the local image. The local tooth recognition result includes a plurality of detection frames and their position information. The second tooth recognition model may be the same as or different from the first tooth recognition model, and this specification does not limit this.

[0069] Based on the second center position information and the preset second size, a third recognition area is determined. Alternatively, this can be understood as determining a third recognition area of ​​the preset second size within the high-resolution three-dimensional image, centered on the second center position information. The high-resolution three-dimensional image is a partial image. The size of the third recognition area can be set as needed and is not limited in this specification.

[0070] When the centroid of the detection frame in the partial tooth recognition result is within the third recognition area, the detection frame in the partial tooth recognition result is scaled according to the dynamic scaling ratio and used as the key tooth recognition result for the partial image. It is understood that the partial tooth recognition result needs to be restored to its original size space because, when the final tooth recognition result is subsequently determined, the size of the detection frame in the recognition result is based on the size in the original size space.

[0071] The third recognition area is determined because the closer the detection frame is to the center of the second recognition area, the higher the confidence level. Therefore, the electronic device may only retain detection frames whose centroids are within the third recognition area. If there is more than one detection frame in the third recognition area, a preset number of detection frames may be retained, with the centers of the retained detection frames being the closest to the center of the first recognition area. Alternatively, the preset number of detection frames may be retained based on the distance between the centers of the detection frames in the partial tooth recognition results and the geometric center of the third recognition area.

[0072] FIG5 is a schematic diagram of the key teeth recognition results provided in this specification, as shown in FIG5 .

[0073] The detection results within the second recognition area include detection frame 7 and detection frame 8. The closer the detection frame within the second recognition area is to the center, the higher the confidence. If only the detection frame with the center of mass within the third recognition area is retained, only detection frame 8 is retained. Detection frame 8 is scaled to obtain the key tooth recognition results of the local image.

[0074] Of course, it is also possible to retain all detection frames in the local tooth recognition result and resize them to serve as the key tooth recognition result for that local image. Resizing the detection frames also requires determining the position information of the scaled detection frames. This can be determined based on the dynamic scaling ratio and is not detailed here.

[0075] S110: Determine a tooth recognition result of the original oral three-dimensional image based on the dentition position information and at least one key tooth recognition result.

[0076] In one or more embodiments of the present disclosure, when determining tooth recognition results for an original three-dimensional oral cavity image, the electronic device may first determine, within the detection frames of each tooth in the dentition position information, detection frames other than the key frame as first frames to be filtered, and then use at least one detection frame in the key tooth recognition result as a second frame to be filtered. Alternatively, this may be understood as determining, among the detection frames determined in the target tooth recognition result, detection frames other than those subsequently determined as key frames as first frames to be filtered.

[0077] Then, for each second box to be screened, a first intersection-and-union ratio (IoU) is determined between the second box to be screened and each of the first boxes to be screened in the plane defined by the first and second dimensions. In other words, the IoU ratio of the first and second boxes in the XY plane is determined as the first IoU ratio.

[0078] When all the first intersection-and-union ratios of the second frame to be filtered are less than the preset first intersection-and-union ratios, the second frame to be filtered is retained and determined as the tooth recognition result of the original oral three-dimensional image.

[0079] It can also be understood that the electronic device calculates the intersection-and-union ratio matrix ious of the second box to be screened and the first box to be screened, and the dimension of the intersection-and-union ratio matrix is ​​(N2, N1), N2 is the number of second boxes to be screened, and N1 is the number of first boxes to be screened, and determines max_iou=max(ious, axis=1), where max_iou=max(ious, axis=1) represents the maximum value of the intersection-and-union ratio of each second box to be screened and all first boxes to be screened. Assuming that the preset first intersection-and-union ratio is 0.5, when max_iou<0.5 is satisfied in the second box to be screened, the second box to be screened is retained.

[0080] For example, assuming the first IoU ratio is preset to 0.5, there are two second boxes to be filtered, namely box21 and box22, and three first boxes to be filtered, namely box11, box12, and box13. The resulting IoU matrix is ​​as follows:

[0081] Then, max_iou = [0.3, 0.9]. For box21, the first intersection-over-union ratios are 0.1, 0.3, and 0.2, respectively, with a maximum of 0.3, which is still less than 0.5, so box21 is retained. For box22, the first intersection-over-union ratios are 0.3, 0.9, and 0.6, respectively, with a maximum of 0.9, which is greater than 0.5, so box22 is not retained.

[0082] The electronic device can also be further integrated to determine the second intersection and union ratio of the second frame to be filtered and each of the first frames to be filtered in the third dimension. When the first intersection and union ratio of the second frame to be filtered is greater than the preset first intersection and union ratio and the second intersection and union ratio is greater than the preset second intersection and union ratio, the second frame to be filtered is retained and determined as the tooth recognition result of the original oral three-dimensional image. The third dimension can be the Z axis. Then, the determined second intersection and union ratio refers to the overlapping range of the Z axis data range of the second filtering frame and the Z axis data range of the first filtering frame. It can be understood that when the first intersection and union ratio of the second frame to be filtered is greater than or equal to the preset first intersection and union ratio and the second intersection and union ratio is less than or equal to the preset second intersection and union ratio, the first frame to be filtered is retained.

[0083] It should be noted that when the first intersection and union of the second box to be screened is greater than the preset first intersection and union, and the second intersection and union is greater than the preset second intersection and union, it is for a second box to be screened and a first box to be screened. Continuing with the above example, if the preset second intersection and union is 0.75, the first intersection and union of box22 and box13 is 0.6, which is greater than 0.5, and the second intersection and union of box22 and box13 is 0.8, then box22 is retained and box13 is deleted. The first intersection and union of box22 and box12 is 0.9, which is greater than 0.5, and the second intersection and union of box22 and box12 is 0.7, then box12 is retained.

[0084] FIG6 is a schematic diagram of the detection frame fusion provided in this specification. As shown in FIG6 , FIG6 includes (a) and (b).

[0085] Assuming the first IoU is preset to 0.5 and the second IoU is preset to 0.9, in (a), the IoU_z along the z direction is max{0, [min(rz2, gz2) - max(rz1, gz1)] / [max(rz2, gz2) - min(rz1, gz1)]}. Where rg1 is the starting coordinate of the first box to be filtered on the Z axis, rg2 is the ending coordinate of the Z axis, gz1 is the starting coordinate of the second box to be filtered on the Z axis, and gz2 is the ending coordinate of the Z axis.

[0086] Since the iou of the second to-be-screened box and the first to-be-screened box is greater than 0.5 and iou_z is greater than 0.9, the second to-be-screened box comes from the result of the second recognition and has a higher confidence than the first to-be-screened box. Therefore, the second to-be-screened box is retained and the first to-be-screened box is deleted.

[0087] The detection frames obtained after fusion are restored to the original oral 3D image to obtain a schematic diagram of the tooth recognition results of the original oral 3D image. Figure 7 is a schematic diagram of the tooth recognition results of the original oral 3D image provided in this manual, as shown in Figure 7.

[0088] Compared with Figure 2, it can be seen that Figure 7 not only removes the false positives in the target tooth recognition results, but also replenishes the two missed false negatives, thereby improving the accuracy of the tooth recognition results.

[0089] Based on the tooth recognition method shown in Figure 1, this method first identifies a low-resolution three-dimensional image of the target area to obtain dentition position information for each tooth in the target area. From this dentition position information, key tooth position information is determined. A high-resolution local image corresponding to the key position information is then obtained. This high-resolution local image is then identified to obtain key tooth recognition results. Based on the dentition position information in the low-resolution image and the key tooth recognition results from the high-resolution image of the key positions, tooth recognition results for the original oral three-dimensional image are obtained. By first identifying the low-resolution three-dimensional image to obtain dentition position information and then performing tooth recognition on the high-resolution three-dimensional image of the key position information, the accuracy of tooth recognition results is improved.

[0090] For step S102, if the target tooth recognition result is a low-resolution three-dimensional image, the electronic device can also restore the detection frame to the original size space to facilitate the subsequent determination of the tooth recognition result of the original oral three-dimensional image. The electronic device can determine a restoration ratio based on the preset first size and the size of the original oral three-dimensional image, and scale the detection frame size based on the restoration ratio and the detection frame size to obtain a scaled detection frame. For example, assuming that the preset first size is (128, 128, 128), the size of the original oral three-dimensional image is (256, 256, 256), the size of the detection frame is (64, 64, 64), and the restoration ratio of the original oral three-dimensional image size to the preset first size is 2, then the size of the detection frame after being restored to the original size space is (128, 128, 128). Then, steps S102 to S110 are executed again.

[0091] Regarding step S104, when sorting the detection frames, since the dentition includes both upper and lower dentitions, the electronic device can sort the detection frames for each dentition. First, for each dentition, the positional information of the detection frames in the dentition position information is obtained, and the first-dimensional positional data and second-dimensional positional data of each detection frame are determined. Then, based on the second-dimensional positional data of each detection frame, the detection frames for each tooth are sorted to obtain an initial sorting result. For example, if the second dimension is the Y-axis in an XYZ coordinate system, the electronic device can sort the detection frames based on the Y-axis data of each detection frame.

[0092] Because the detection frames are three-dimensional, the initial sorting results obtained by sorting based solely on the second-dimensional position data are less accurate. For example, the patient's molars are located at the beginning and end of the dentition, as shown in Figure 2. When viewed from the X-axis, the molars may be obscured by other adjacent teeth. Accordingly, the coordinates of the detection frames on the Y-axis may be misaligned, that is, the teeth that should be in the front near the left end of Figure 2 may be placed in the back, or the teeth that should be in the back near the right end may be placed in the front. For example, the position of a tooth on the left side of the midline should be the third tooth, but it becomes the fourth tooth. The coordinates may be (1, 3, 5) instead of (1, 5, 5). Therefore, adjustments can also be made based on the first-dimensional position information. In other words, the electronic device can adjust the order of the detection frames of the specified teeth in the initial sorting result based on the first-dimensional position data of each detection frame to obtain the final sorting result. The specified teeth can be set as needed, such as the molars mentioned above. The final sorting result includes the sorted detection frames of the upper dentition and the sorted detection frames of the lower dentition.

[0093] Taking Figure 2 as an example, assuming that the first dimension is the X-axis, as shown in Figure 2, the detection frames of the upper and lower teeth are sorted based on the Y-axis distance in Figure 2, and the sorted detection frame sequence is recorded as L. For the first 4 and last 4 tooth detection frames in L, the sorting order is readjusted based on the X-axis to obtain the final sorted detection frame sequence L. The elements in L are the detection frames from left to right in Figure 2, and the upper teeth are recorded as L1 and the lower teeth are recorded as L2.

[0094] This specification also provides another method for determining key teeth and key position information of the key teeth. The key teeth determined are canines and central incisors. In the medical field, central incisors are also called front teeth, and canines are also called tiger teeth. Specifically, for each dentition, the electronic device determines the midline position information in the sorted detection frame based on the second-dimensional position data of the detection frame of each tooth in the dentition. For example, in Figure 2, the Y-axis data of the first detection frame on the left is the starting point, if it is 10, and the Y-axis data of the last detection frame on the left is the ending point, if it is 20, then the midline position is 15.

[0095] Then, based on the midline position information, a first key tooth is identified, and the position information of the first key tooth is determined as the first key position information. The electronic device may select any tooth adjacent to the midline position as the first key tooth, but this specification does not impose any restrictions on this. In situations where the tooth is intact and no tooth has been missed, the tooth adjacent to the midline is generally the central incisor. The electronic device may also determine a first adjacent frame and a second adjacent frame adjacent to the first key frame of the first key tooth, and determine a second key frame within the first adjacent frame and the second adjacent frame.

[0096] If the patient's teeth are normal, but there is a missed detection in the first tooth recognition model, resulting in a tooth not being recognized, then when determining the key teeth, the electronic device can also first determine whether there is a missed detection. The electronic device can determine the first adjacent intersection and union ratio of the first key frame and the first adjacent frame, and determine the second adjacent intersection and union ratio of the first key frame and the second adjacent frame. The intersection and union ratio is the intersection and union ratio of the plane formed by the X-axis and the Y-axis, that is, the first adjacent intersection and union ratio is determined based on the XY plane data of the first key frame and the XY plane data of the first adjacent frame. Furthermore, based on the first key position information, the first adjacent intersection and union ratio and the second adjacent intersection and union ratio, the second key frame is determined in the first adjacent frame and the second adjacent frame, and the position information of the second key tooth is determined as the second key position information.

[0097] If the first and second adjacent IoU ratios are less than the preset IoU threshold, there is a high probability that an object has been missed between the key frame and the adjacent frame. The adjacent frame closer to the centerline of the first and second adjacent frames is selected as the second key frame. The preset IoU threshold can be set as needed, such as 0.05.

[0098] If the adjacent intersection-and-union ratio is not less than a preset intersection-and-union ratio threshold, a second key frame is determined in the first adjacent frame and the second adjacent frame based on the position information of each adjacent frame. Based on the first key position information and the position information of the detection frames of each tooth in the dentition, a tooth in the dentition that has a specified positional relationship with the first key tooth is determined to be a third key tooth, and the position information of the third key tooth is set as third key position information.

[0099] Continuing with the example of the upper dentition L1 and the lower dentition L2 obtained above, based on the sorted upper and lower dentition L1 and L2 already obtained, take L1 as an example, and the same applies to L2. Determine the first key tooth L1[c] adjacent to the midline position. If the first key tooth L1[c] is the central incisor on the left side of the midline, and there is no missed target between the key frame L1[c] and the adjacent frames L1[c+1] and L1[c-1], then L1[c+1] is selected as the detection frame of another central incisor, that is, L1[c+1] is determined as the second key frame. Determine the teeth L1[c-2] and L1[c+3] whose positional relationship with the first key tooth L1[c] is the specified positional relationship, and select L1[c-2] and L1[c+3] as the canine frames, and the canine frame is the third key frame.

[0100] If the center of L1[c] is to the right of the midline, then L1[c-1] is selected as the other incisor frame, and L1[c-3] and L1[c+2] are selected as the canine frames. Note that in Figure 2, L1[c+1] represents the tooth to the right of L1[c], L1[c-1] represents the tooth to the left of L1[c], and so on.

[0101] It should be noted that the electronic device may also first determine the intersection-and-union ratio of each pair of adjacent frames in the sorted detection frame, and determine whether there is any missed detection between the pair of adjacent frames based on the intersection-and-union ratio of the pair of adjacent frames. If so, the adjacent frame that is closer to the center line is selected as the key frame. After all adjacent frames have been checked for missed detection, the first key tooth can be determined based on the center line position information, and the position information of the first key tooth can be determined as the first key position information, and then the second key frame and the third key frame can be determined. This manual does not limit whether to first determine whether there is a missed detection target between each pair of adjacent frames or to first determine the first key frame. Missed detection may be due to low detection accuracy, or it may be because the user himself is missing a tooth.

[0102] For example, detection frame 1 is adjacent to detection frame 2, and the intersection-over-union (IoU) of detection frame 1 and detection frame 2 is determined. If the IoU is less than the preset IoU threshold, there is a missed frame between detection frame 1 and detection frame 2. The distance between detection frame 1 and detection frame 2 and the center line is determined, and the adjacent frame closer to the center line is selected as the key frame.

[0103] In step S106, when determining the dynamic scaling ratio, the electronic device determines the minimum scale value of the original three-dimensional oral cavity image in each dimension based on the image size of the original three-dimensional oral cavity image. If the minimum scale value is not greater than the preset first size, the electronic device determines the image magnification ratio based on the minimum scale value and the preset first size, and uses this as the dynamic scaling ratio. If the minimum scale value is not greater than the preset first size, it indicates that the original three-dimensional oral cavity image is too small to accommodate the second recognition area, and the original three-dimensional oral cavity image needs to be enlarged.

[0104] For example, if the preset first size is 128 and the image size (H_data, W_data, D_data) of the original oral 3D image is (2, 4, 6), then the minimum scale value is 2, and the comparison result r = preset first size / min(H_data, W_data, D_data) = 128 / 2 = 64 > 1, indicating that the minimum scale value is not greater than the preset first size, then the image magnification ratio r_dynamic = r = 64, and the size of the enlarged oral 3D image is (2*64=128, 4*64=256, 6*64=384).

[0105] When the minimum scale value is greater than the preset first size, the maximum scale value of the first identification area in each dimension is determined according to the size of the first identification area, and the image reduction ratio is determined according to the maximum scale value and the preset first size, and is used as the dynamic scaling ratio.

[0106] For example, if the preset first size is 128, and the image size (H_data, W_data, D_data) of the original oral three-dimensional image is (256, 512, 256), then the minimum scale value is 256, and the comparison result r=128 / 256=0.5<1, then the minimum scale value is greater than the preset first size. The size of the first recognition area (H_target, W_target, D_target) is (52, 240, 220), and it is determined that the maximum scale value of the first recognition area is 240. When the image reduction ratio r_dynamic=preset first size / max(H_target, W_target, D_target)=128 / 240=0.53<1, it means that the size of the first recognition area is large and the second recognition area cannot accommodate the first recognition area. It can be understood that in this specification, the second recognition area must also accommodate the first recognition area. Then, when the original oral three-dimensional image is reduced, the first recognition area will also be reduced accordingly. Therefore, the electronic device may need to reduce the size of the original oral three-dimensional image. The image reduction ratio = 0.25, and the size of the reduced oral 3D image is (0.25*256=64, 0.25*512=128, 0.25*256=64). The image reduction ratio may be rounded down to one decimal place.

[0107] Of course, if the image reduction ratio r_dynamic ≥ 1, it means that the second recognition area can accommodate the first recognition area and no adjustment is required, and the dynamic scaling ratio = r_dynamic = 1.

[0108] It should be noted that each key tooth can determine the corresponding first identification area, so each first identification area has a corresponding dynamic scaling ratio, a scaled three-dimensional oral image, second center position information and a second identification area.

[0109] With respect to step S108 , the electronic device may further optimize the local teeth recognition result of the partial image before determining the key teeth recognition result.

[0110] Specifically, FIG8 is a schematic diagram of dimension crossing provided in this specification, as shown in FIG8 .

[0111] When the first recognition area is close to the edge of the original oral three-dimensional image, a certain dimension coordinate of the second recognition area will be out of bounds. For example, if the length of the X-axis of the original oral three-dimensional image is 128 and the starting coordinate is 0, then the end coordinate of the original oral three-dimensional image is 127. The coordinate out of bounds in the X-axis of the second recognition area can be understood as the X-axis coordinate of the second recognition area being less than 0 or greater than 127. In other words, if the minimum or maximum coordinate of a certain dimension of the second recognition area exceeds the actual boundary of the original oral three-dimensional image, it is considered that the second recognition area is out of bounds. At this time, under the premise of ensuring that the second recognition area is the preset first size, the electronic device can develop in the opposite direction of the out-of-bounds direction of the same dimension coordinate. Although the first recognition area is not in the center position in the second recognition area and a certain degree of offset occurs, the third recognition area can still retain part of the detection box in the local tooth recognition result.

[0112] In other words, after obtaining the partial tooth recognition result, the electronic device can determine whether the coordinates of the second recognition area are out of bounds based on the size of the second recognition area and the size of the original oral 3D image. If so, the coordinates of the out-of-bounds dimensional coordinates of the second recognition area are offset until the dimensional coordinates are no longer out of bounds. After the offset, the second recognition area still includes the first recognition area, and the third recognition area still retains part of the detection frame in the partial tooth recognition result. The electronic device can then retain the detection frame in the third recognition area.

[0113] If the coordinates of the second recognition area are out of bounds and no adjustment is made, part of the second recognition area will exceed the boundary of the original oral three-dimensional image, resulting in data loss or processing errors, affecting the accuracy of the tooth recognition results of the original oral three-dimensional image.

[0114] Furthermore, the electronic device may also perform deduplication on the partial tooth recognition results in all the second recognition areas obtained in step S108. Specifically, deduplication may be performed using methods such as non-maximum suppression (NMS), which is not limited in this specification. This is because there are several key teeth, each of which may have a corresponding second recognition area, so several partial tooth recognition results can be obtained. There may be overlapping detection frames in several partial tooth recognition results. Therefore, deduplication can be performed first, and then the key tooth recognition results can be determined in the third recognition area.

[0115] Regarding step S110, this specification also provides another detection frame fusion method. In this case, the first and second detection frames may appear concentric. In this case, the intersection-over-union ratio between the two detection frames is required to be greater than a preset value. This is because the first and second detection frames are required to have similar three-dimensional proportions and similar center coordinates, thereby improving the accuracy and reliability of the recognition results.

[0116] Specifically, the electronic device can, for each first box to be screened, when the first box to be screened completely contains the second box to be screened, simultaneously reduce the size of all dimensions of the first box to be screened in equal proportion, that is, according to the dimension reduction ratio, reduce the size of the X, Y, and Z axis dimensions at the same time. The dimension reduction ratio can be set as needed, and this specification does not limit this. When the size of any dimension of the reduced first box to be screened is consistent with the size of the corresponding dimension of the included second box to be screened, determine the third intersection and union ratio of the reduced first box to be screened and the included second box to be screened. For example, the sizes of the X, Y, and Z axis dimensions are reduced at the same time, and when the size of the X dimension of the reduced first box to be screened is consistent with the size of the X dimension of the included second box to be screened, determine the third intersection and union ratio. It is understandable that when the three dimensions of the first box to be screened are reduced in equal proportion at the same time, at least one dimension must be equal to the size of the corresponding dimension of the second box to be screened. The aforementioned at least one dimension belongs to the X, Y, and Z axis dimensions, that is, the aforementioned arbitrary dimension can be any one or more of the X, Y, and Z axis dimensions.

[0117] When the third intersection-over-union ratio is greater than the preset third intersection-over-union ratio, the second frame to be filtered is retained and determined as the tooth recognition result of the original oral three-dimensional image.

[0118] As shown in Figure 6(b), the outer detection box is the first to-be-screened box, and the inner detection box is the second to-be-screened box. The second to-be-screened box is completely contained by the first to-be-screened box, and the center coordinates of the second to-be-screened box and the first to-be-screened box are the same or similar. Therefore, the teeth contained in the second to-be-screened box are also contained in the first to-be-screened box. If the first to-be-screened box contains additional teeth, since the second to-be-screened box is smaller than the first to-be-screened box, the second to-be-screened box may not contain the additional teeth, thus improving the accuracy of the tooth recognition results.

[0119] This description also provides a training method for an image classification model. The execution subject can be a computing device, electronic device, etc. that can train the model. This description does not limit this and will now be explained using electronic devices.

[0120] First, the electronic device can obtain several positive samples, which may include an image of a sample detection frame. A positive sample is a partial image captured from the sample detection frame, containing only one complete target tooth. It is understood that positive samples can be obtained through manual annotation or other annotation methods. Samples other than positive samples are negative samples. Negative samples contain more than one complete target tooth or do not contain a single complete target tooth.

[0121] Then, for each positive sample, the sample minimum bounding box of the sample detection box of the positive sample and any number of adjacent sample detection boxes is determined, and the sample minimum bounding box includes the minimum bounding box of the XY plane and the YZ plane.

[0122] The lengths of the first and second dimensions of the sample's minimum bounding box are scaled according to a preset scaling ratio to obtain a scaled sample minimum bounding box. An image of the scaled sample minimum bounding box is then determined to obtain a first negative sample. The detection frame used to obtain the first negative sample is located between or encompasses two adjacent teeth, and the preset scaling ratio can be set as needed. It should be noted that the value of the third dimension of the first negative sample can be determined based on the sample minimum bounding box in the YZ plane.

[0123] FIG9 is a schematic diagram of a negative sample provided in this specification, as shown in FIG9 .

[0124] Taking the selection of an adjacent sample detection box of a positive sample as an example, the solid lines represent the positive sample and the adjacent sample detection box 1 of the positive sample, respectively, and the dashed line represents the negative sample. The minimum bounding box of the sample obtained in the XY plane is shown by the dotted line in Figure 9. Its height and width are (h, w) and its coordinates are (X1, Y1, X2, Y2). h and w are scaled proportionally while keeping their center positions unchanged. The scaling ratio r' = random(-0.05, 0.15). The scaling method can be reduced when r'>0 and enlarged when r'<0. The generated negative sample is shown by the dashed line. In the XY plane, the coordinates of the negative sample are (x1, y1, x2, y2), so x1 = X1 + h*r, y1 = Y1 + w*r, x2 = X2 - h*r, and y2 = Y2 - w*r. The Z-axis coordinate of the negative sample is a random value between the Z-axis coordinates of the corresponding positive sample and the adjacent sample detection frame 1. If the Z-axis range of the adjacent sample detection frame 1 is (GZ1, GZ2) and the Z-axis range of the positive sample is (BZ1, BZ2), then z1 = random(min(BZ1, GZ1), max(BZ1, GZ1)), and z2 = random(min(BZ2, GZ2), max(BZ2, GZ2)). This can also be understood as determining the value of the third dimension of the negative sample based on the numerical range of the third dimension of the positive sample and the numerical range of the third dimension of the adjacent sample detection frame.

[0125] The electronic device can also translate the sample detection frame of the positive sample according to a preset translation direction, and scale the translated sample detection frame to obtain a sample detection frame after the translation and scaling operations, determine the image of the sample detection frame after the translation and scaling operations, and obtain a second negative sample, wherein the preset translation direction can be the third dimension, and the translation distance can be set as needed.

[0126] FIG10 is a schematic diagram of another negative sample provided in this specification, as shown in FIG10 .

[0127] If the positive sample is a partial image captured based on the sample detection frame in the upper dentition, it is translated downward along the Z axis by half the frame height, and then the translated sample detection frame is randomly reduced with the center height unchanged, and the scaling ratio r'=random(-0.25, 0). The same applies to the lower dentition.

[0128] The electronic device may then add the positive sample, the first negative sample, and the second negative sample to a training sample set, input the samples in the training sample set into an image classification model, and obtain a predicted classification result output by the image classification model. The image classification model is trained based on the predicted classification result, the positive sample, the first negative sample, and the second negative sample.

[0129] This specification also provides a training method for a second tooth recognition model, where an electronic device obtains a high-resolution sample image and a key sample frame of the high-resolution sample image to obtain a label, inputs the high-resolution sample image into the second tooth recognition model, obtains a sample prediction frame of the high-resolution sample image output by the second tooth recognition model, and trains the second tooth recognition model based on the label and the sample prediction frame.

[0130] Similar to the second tooth recognition process described in this manual, when training the second tooth recognition model, the labeled frames can also be sorted and grouped, and key sample frames can be identified. The first and second recognition areas of the key sample frames can then be determined, and a dynamic scaling ratio can be determined. The original oral 3D image of the sample can then be scaled to obtain a high-resolution sample image. After obtaining the sample prediction frames, several sample prediction frames can be screened to obtain key sample recognition results. Based on these key sample recognition results and labels, the second tooth recognition model can be trained.

[0131] When acquiring key sample frames, to ensure diversity and specificity, teeth can be grouped by tooth type, with a preset number of key sample frames selected from each group. For tooth types with a high empirically determined recognition error rate, a larger number of detection frames of that tooth can be selected as key sample frames. Tooth types include central incisors, lateral incisors, canines, molars, and others.

[0132] For example, divide the dentition into 5 groups, taking the L1 example obtained above, and L2 in the same way, let the center position index of L1 be idx, group 1 = (idx-1, idx, idx+1) is the incisor index, group 2 = (idx-2, idx-3) is the left canine index, group 3 = (idx+2, idx+3) is the right canine index, group 4 = range(1, idx-3) is the index of the remaining teeth on the left, group 5 = range(idx+4, len(L1)-1) is the index of the remaining teeth on the right, and select a key sample frame in each group.

[0133] It should be noted that when determining the key sample frame, part of the structure of the adjacent teeth may also be marked in the second recognition area during annotation. For example, the annotation frame is slightly larger during annotation and includes part of the adjacent teeth. If it is only a small part, the partial structure of the adjacent teeth can be ignored when training using the key sample frame. When judging whether to ignore the partial structure of the adjacent teeth, the electronic device can set the complete segmentation annotation of a certain tooth to be m, and the intercepted segmentation annotation to be m-part. When volume(m-part) / volume(m)<0.5, the partial structure of the adjacent teeth is ignored. Otherwise, the partial structure of the adjacent teeth is determined as the target tooth to be identified.

[0134] Considering that oral CBCT images are the main basis for oral diagnosis, they can provide sufficient oral information. 3D tooth construction based on CBCT images has a wide range of applications, and 3D positioning of teeth is an important component of 3D construction. The accuracy of positioning has a huge impact on the final construction result. It is possible to extract the tooth area and convert this area into a fixed size. The 3D detection method based on deep learning is used to realize the positioning of the teeth, and then restore the positioning result to the original size space. The 3D detection method based on deep learning may need to convert the target area into the same fixed size. However, due to the large size of 3D data and the current limitation of GPU memory size, it can only be converted into a relatively small low-resolution image, thereby losing a lot of image information, resulting in a decrease in accuracy and recall rate, especially for relatively small targets.

[0135] FIG11 is a flow chart of a tooth detection method provided in this specification, which specifically includes the following steps:

[0136] S1100: Acquire a low-resolution three-dimensional image of the target area based on the original three-dimensional oral cavity image.

[0137] 3D images of teeth allow for observation of teeth from multiple angles, more intuitively displaying the anatomical structure of teeth, including crowns, roots, and pulp cavities. They can also more clearly show the location, extent, and morphology of lesions, aiding doctors in their judgment. When constructing 3D images of teeth, the tooth region in an oral CBCT image may be used as input for a tooth detection model. Because the model input requires a fixed size, the tooth region in the oral CBCT image also needs to be resized. However, due to the large amount of 3D data and limitations on GPU memory size, the image of the tooth region can only be converted into a low-resolution image. However, low-resolution images can lead to inaccurate detection results output by the model. Therefore, this specification provides a tooth detection method. The execution entity of this specification includes a server deployed with a method for obtaining tooth detection results, or other electronic devices capable of being deployed with a method for obtaining tooth detection results. It can also be various models, or a server for training the model. This specification does not impose any restrictions on this.

[0138] It should be noted that the low resolution mentioned in this application can be determined based on the size, resolution, image quality, etc. of the obtained image, or it can be determined by at least one factor such as the requirements for global detection of the image and computing power, image recognition model, accuracy of the original image and recognition accuracy. A low resolution value range can be determined based on a combination of multiple factors, and this application does not make any specific limitations.

[0139] The high resolution mentioned in this application can be determined relative to the low resolution, or it can be determined based on the size, resolution, image quality, etc. of the collected medical images, or it can be determined based on computing power and detection requirements. A high resolution value range can be determined based on a combination of multiple factors, and this application does not make any specific limitations.

[0140] The electronic device may first perform a first detection on the image containing the teeth area, and then the electronic device may first obtain the image containing the teeth area.

[0141] Specifically, a medical imaging device acquires a raw 3D oral image of the user and transmits the raw 3D oral image to the electronic device, which then receives the raw 3D oral image transmitted by the medical imaging device. Since the raw 3D oral image includes not only the user's teeth but also other structures of the user's oral cavity, such as the maxilla and mandible, to reduce the impact of other structures on the detection results, the raw 3D oral image can be segmented to obtain a region containing only all of the user's teeth. Specifically, the raw 3D oral image is input into an image segmentation model, and the image segmentation model outputs a region-of-interest image of the raw 3D oral image. The region-of-interest is the region containing all of the user's teeth. Since the input image size of the global detection model is fixed, the image size of the region-of-interest image can also be converted to a preset size to obtain a low-resolution 3D image of the target area. The preset size can be set according to the needs of the global detection model and is not limited in this specification. For example, the size can be (128, 128, 128). The image segmentation model can be an encoding-decoding structure network such as U-Net, V-Net, etc. to realize fully automatic region of interest extraction. This specification does not limit the type of image segmentation model, as long as the image segmentation model can output the region of interest image.

[0142] S1102: Input the low-resolution three-dimensional image into a global detection model to obtain a detection frame of each tooth in the target area output by the global detection model.

[0143] The electronic device inputs the low-resolution 3D image into a global detection model, which then outputs a detection frame for each tooth in the target area. When using the global detection model for tooth detection, the global detection model detects each tooth in the low-resolution 3D image of the target area, determines the position of each tooth, and encloses each tooth with a detection frame. The global detection model can be any model capable of detecting teeth and generating detection results containing detection frames, such as RetinaNet3D.

[0144] S1104: Select a key frame from the detection frame of each tooth.

[0145] Because the low-resolution 3D image used to obtain the detection frames for each tooth in the target area has a low resolution, much image information may be lost, resulting in inaccurate detection frames for each tooth in the target area. For example, there may be a detection frame that does not fully encompass a single tooth, or a detection frame that fully encompasses multiple teeth. Therefore, a second detection can be performed. During this second detection, only the key areas can be selected for inspection. The electronic device can then first determine the key frames for the second detection.

[0146] Specifically, for each detection frame, the electronic device inputs the image of the detection frame into the image classification model and obtains a classification result for the image of the detection frame output by the image classification model. If the classification result indicates that the image of the detection frame is not the target image, the electronic device determines the annotated detection frame corresponding to the image of the detection frame and uses the annotated detection frame as the key frame. In other words, the image classification model determines whether the detection frames of each tooth in the target area output by the global detection model are correct detection frames. The target image is the image of the correct area corresponding to the detection frame, which refers to the area containing only one complete target tooth.

[0147] S1106: For each key frame, determine a detection area including the key frame, and acquire a high-resolution three-dimensional image of the detection area based on the original three-dimensional oral cavity image.

[0148] In one or more embodiments of the present specification, the electronic device may acquire an image of the area where the second detection is to be performed before performing the second detection.

[0149] Specifically, the detection frames of each tooth in the target area output by the global detection model are first sorted to obtain sorted detection frames. Within the sorted detection frames, for each key frame, two detection frames adjacent to the key frame are determined to obtain adjacent detection frames. When the adjacent detection frame intersects the key frame and / or intersects the key frame after being translated by a preset distance, the minimum bounding box between the adjacent detection frame and the key frame is determined. The preset distance can be the height or width of the adjacent detection frame itself. The purpose of determining the minimum bounding box is to ensure that the key frame is located at the center of the image input to the local detection model during the second detection, thereby improving the accuracy of the second detection result. The minimum bounding box between the adjacent detection frame and the key frame is selected only when the adjacent detection frame intersects the key frame or intersects the key frame after being translated by a preset distance. This is because if there is a missing tooth between the adjacent detection frame and the key frame, the key frame and the adjacent detection frame will also be adjacent, but the distance is greater. In this case, the key frame is not located at the center of the obtained minimum bounding box, which is not conducive to improving the accuracy of the second detection.

[0150] Next, based on this key frame, an image of the key frame is determined within the original 3D oral image. Alternatively, the key frame image can be determined within the previously acquired ROI image. This is because the ROI image and the original 3D oral image have higher resolutions and contain more image information. Therefore, the determined key frame image also has higher resolution, which improves the accuracy of the second test result.

[0151] Because the input to the local detection model is fixed-size, and the high-resolution key frame image may be too large to be fully accommodated by the fixed size, the detection image of the detection area may not be fully included, resulting in a reduction in the results of the second detection. Therefore, the electronic device can also convert the size of the key frame image to a preset size to obtain a converted image of the key frame.

[0152] When the minimum three-dimensional size of the key frame image is no larger than a preset size, an image magnification ratio is determined, and the size of the key frame image is magnified to the preset size based on the image magnification ratio, thereby obtaining a converted key frame image. The size of the key frame image is magnified because the minimum three-dimensional size of the key frame image is no larger than the preset size. When input into the local detection model, there is a blank area. To maximize the image information of the input image, the size of the key frame image is magnified so that the previously blank area is filled with the key frame image.

[0153] For example, if the preset size is (128, 128, 128), and the three-dimensional sizes of the key frame image are (H_data, W_data, D_data), then the minimum three-dimensional size of the key frame image is no larger than the preset size, and the image magnification ratio is r1=128 / min(H_data, W_data, D_data).

[0154] When the minimum three-dimensional size of the key frame image is greater than a preset size and the maximum three-dimensional size of the detection area image is greater than a preset size, an image reduction ratio is determined. Based on the image reduction ratio, the size of the key frame image is reduced to the preset size to obtain a converted image of the key frame.

[0155] For example, if the three dimensions of the minimum bounding box are (H_target, W_target, D_target), and the minimum three-dimensional size of the key frame image is larger than the preset size and the maximum three-dimensional size of the detection area image is larger than the preset size, that is, r = 128 / min(H_data, W_data, D_data), r < 1, and r_dynamic = 128 / max(H_target, W_target, D_target), r_dynamic < 1, the image reduction ratio = 128 / max(H_target, W_target, D_target), and r_dynamic is rounded down to one decimal place. If r_dynamic > = 1, no adjustment is required.

[0156] According to the image of the key frame after the conversion and the detection area, the detection image of the detection area is determined. That is, the image of the key frame after the conversion is obtained, and the image exists in the detection area, and the detection image of the detection area is obtained. The center coordinates of the detection image can also be adjusted accordingly. The center coordinates are adjusted to = (x0, y0, z0) * r_dynamic, (x0, y0, z0) are the center coordinates of the image of the detection area before conversion. It should be noted that the image of the detection area and the detection image of the detection area are not the same image. The image of the detection area is the image of the minimum circumscribed frame captured in the original oral three-dimensional image or the image of the area of ​​interest. The detection image of the detection area is the image of the key frame after the conversion and is within the detection area.

[0157] S1108: Input the high-resolution three-dimensional image into a local detection model to obtain a detection result of the detection image output by the local detection model.

[0158] In the detection results, the closer the detection frame is to the center, the higher the accuracy. Therefore, a central region within the detection area can be determined, and only the detection frames in this central region are retained. If there are multiple detection frames in the central region, a preset number of detection frames are retained based on their distance from the geometric center of the central region. For example, the detection frame with the shortest distance from the geometric center of the central region is determined as the detection result for the detected image. The size of the central region can be set as needed and is not limited in this manual.

[0159] When the image of the minimum bounding box is near the edge of the test image, a coordinate in one dimension may be out of bounds when determining the test image. This means that the coordinate value of that dimension is less than 0 or greater than the preset size. In this case, while maintaining the preset size of the test image, the coordinates of the same dimension move in the opposite direction of the out-of-bounds. In this case, the image of the minimum bounding box is not in the center of the test image, and there is a certain degree of offset. However, the target tooth result can still be retained in the center area.

[0160] For the detection results of all key boxes, there may be overlapping detection boxes after aggregation. Therefore, the non-maximum suppression (NMS) method can also be used to filter and remove duplicate detection results of all key boxes to obtain the final detection results.

[0161] S1110: Filter the detection frames of the teeth and the detection results according to a preset screening method to obtain tooth detection results.

[0162] Since the labeled frames corresponding to the incorrect detection frames in the detection frames of each tooth in the target area were already selected as key frames when selecting key frames, during the screening process, the detection frames other than the key frames in the detection frames of each tooth are selected as the first frames to be screened, and all detection frames in the detection results are selected as the second frames to be screened.

[0163] Afterwards, for each second frame to be screened, the first intersection over union (IoU) of the second frame to be screened and each of the first frames to be screened is determined in the plane formed by the first dimension and the second dimension (XY plane). The direction of the second dimension is from the right cheek of the mouth to the left cheek of the mouth and is perpendicular to the third dimension. The direction of the third dimension is from the lower jaw of the mouth to the upper jaw of the mouth. The direction of the first dimension is from the tip of the tongue to the root of the tongue and is perpendicular to the second dimension and the third dimension.

[0164] When all the first intersection-and-union ratios of the second frame to be screened are less than the preset first intersection-and-union ratio, the second frame to be screened and each of the first frame to be screened are retained to obtain a tooth detection result. The preset first intersection-and-union ratio may be 0.5.

[0165] In the third dimension, the second intersection-and-union ratio of the second frame to be filtered and each of the first frames to be filtered is determined, that is, iou_z=max{0, [min(rz2, gz2)-max(rz1, gz1)] / [max(rz2, gz2)-min(rz1, gz1)]}.

[0166] When the first intersection-and-union ratio of the second frame to be screened is greater than the preset first intersection-and-union ratio and the second intersection-and-union ratio is greater than the preset second intersection-and-union ratio, the second frame to be screened is retained to obtain a tooth detection result.

[0167] For each first box to be screened, when the first box to be screened completely contains the second box to be screened, the length of a certain dimension of the first box to be screened is reduced until the length of the certain dimension is consistent with the length of the certain dimension of the contained second box to be screened, and the third intersection-union ratio of the reduced first box to be screened and the contained second box to be screened is determined.

[0168] When the third intersection-over-union ratio is greater than the preset third intersection-over-union ratio, the second frame to be screened is retained to obtain a tooth detection result.

[0169] Based on the tooth detection method shown in Figure 11, this method first performs global detection on a low-resolution 3D image of the target area to obtain detection frames for each tooth in the target area. Key frames are then selected from each tooth's detection frame. For each key frame, the detection area containing the key frame is determined. Based on the original oral 3D image, a high-resolution 3D image of the detection area is obtained. This high-resolution 3D image is input into a local detection model to obtain a detection result for the detection image output by the local detection model. The tooth detection frames and the detection results are then filtered according to a preset filtering method to obtain a tooth detection result. By first performing global detection on the low-resolution 3D image and then performing local detection on the local high-resolution 3D image, the accuracy of the tooth detection results is improved.

[0170] For step S1104 , a key frame may be determined based on historical tooth detection results. The key frame is a detection frame with a higher error rate in the historical tooth detection results, including a canine detection frame and a central incisor detection frame.

[0171] Specifically, since the global detection model does not output the detection frames for each tooth in the target area according to the order in which the teeth are arranged, before determining the canine and central incisor detection frames, the detection frames for each tooth can be sorted according to the order in which the teeth are arranged, resulting in sorted detection frames. Within the sorted detection frames, the central incisor and canine detection frames are determined, and these are used as key frames. In medicine, central incisors are also known as incisors, and canines are also known as canines. If the central incisors and canines are described according to the order in which the teeth are arranged, the upper dentition will be used as an example. Within the sorted detection frames, the midline position is determined, which is the midpoint between the positions of the first and last sorted detection frames. The first detection frame for the upper dentition is positioned closer to the right cheek than the positions of the other detection frames for the upper dentition, and the last detection frame is positioned closer to the left cheek than the positions of the other detection frames for the upper dentition. The left and right in the right cheek and the left cheek of the mouth are the same as the left and right in the left and right hands of the human body.

[0172] Then, determine the detection frame with the smallest distance from the midline position as the first central incisor detection frame. If the first central incisor detection frame is to the left of the midline position, the detection frame on the right side adjacent to the first central incisor is used as the second central incisor detection frame, the detection frame to the left of the first central incisor detection frame and separated by one detection frame is used as the first canine detection frame, and the detection frame to the right of the second central incisor detection frame and separated by one detection frame is used as the second canine detection frame. If the first central incisor detection frame is to the right of the midline position, the detection frame to the left of the first central incisor is used as the second central incisor detection frame, the detection frame to the right of the first central incisor detection frame and separated by one detection frame is used as the first canine detection frame, and the detection frame to the left of the second central incisor detection frame and separated by one detection frame is used as the second canine detection frame. The two central incisor detection frames and the two canine detection frames are used as key frames. The lower dentition is similar to the upper dentition and will not be described in detail in this manual.

[0173] This specification also provides a method for training an image classification model.

[0174] Specifically, a positive sample is obtained, and based on this positive sample, a negative sample is determined. Both the positive and negative samples are added to the training sample set. A sample includes a sample detection frame and an image of the sample detection frame. A positive sample is one in which the sample detection frame contains only one complete target tooth. Positive samples can be obtained through manual annotation. Samples other than positive samples are negative samples. In other words, a negative sample's sample detection frame contains more than one complete target tooth or contains several incomplete target teeth.

[0175] FIG12 is a schematic diagram of a negative sample provided in this specification, as shown in FIG12 .

[0176] Including several incomplete target teeth means that the sample detection frame is located between two teeth and does not completely include either tooth. That is, in the plane formed by the second and third dimensions (YZ plane), the sample detection frame is between tooth 1 and tooth 2, and does not completely include either tooth 1 or tooth 2.

[0177] When determining negative samples based on the positive samples, a specified number of positive samples are selected, and the minimum bounding box of the sample detection boxes of the specified number of positive samples is scaled according to a preset scaling ratio. It should be noted that the geometric center of the minimum bounding box remains unchanged during the scaling. In the plane formed by the first dimension and the second dimension (XY plane), the length of the first dimension and the length of the second dimension of the minimum bounding box are scaled proportionally according to the preset scaling ratio.

[0178] As shown in Figures 9 and 10, in the plane formed by the second and third dimensions (YZ plane), for the positive sample of the upper dentition, the detection frame of the positive sample is translated downward by a specified length in the third dimension, and the length of the detection frame of the positive sample after the downward shift is reduced according to a preset reduction ratio to obtain a negative sample. For the positive sample of the lower dentition, the detection frame of the positive sample is translated upward by a specified length in the third dimension, and the length of the detection frame of the positive sample after the upward shift is reduced according to a preset reduction ratio to obtain a negative sample.

[0179] The length of the third dimension is randomly generated to obtain a scaled three-dimensional sample detection frame, and a sample image of the three-dimensional sample detection frame is obtained to obtain a negative sample.

[0180] As shown in Figure 9, let the height and width of the minimum bounding box be (h, w), and the coordinates of the minimum bounding box be (X1, Y1, X2, Y2). H and w are randomly scaled proportionally while keeping their center positions unchanged, with a ratio of r = random(-0.05, 0.15). When r > 0, the image is scaled down, and when r < 0, the image is scaled up. The resulting three-dimensional sample detection box is shown by the dotted line in Figure 9. In the XY plane, the coordinates of the three-dimensional sample detection box are (x1, y1, x2, y2), where x1 = X1 + h*r, y1 = Y1 + w*r, x2 = X2 - h*r, and y2 = Y2 - w*r.

[0181] For the three-dimensional sample detection box, the length of the third dimension (Z axis) is a random value, z1 = random (min (BZ1, GZ1), max (BZ1, GZ1)), z2 = random (min (BZ2, GZ2), max (BZ2, GZ2)).

[0182] The samples in the training sample set are input into the image classification model to obtain the classification results output by the image classification model. The image classification model is trained based on the classification results, the positive samples, and the negative samples. In other words, the image classification model is trained by judging whether the classification results are correct based on the positive samples and the negative samples.

[0183] This specification also provides a method for training a local detection model.

[0184] A sample key frame and an annotated frame of the sample key frame are obtained. Based on the image of the sample key frame, a sample detection image of the sample key frame and an image of the annotated frame of the sample key frame are determined in a previously acquired sample region of interest image. The sample detection image of the sample key frame is input into a local detection model to obtain a sample prediction frame and an image of the sample prediction frame output by the local detection model. The local detection model is trained based on the sample prediction frame, the image of the sample prediction frame, the annotated frame of the sample key frame, and the image of the annotated frame of the sample key frame. That is, a first difference between the annotated frame of the sample key frame and the sample prediction frame, and a second difference between the image of the annotated frame of the sample key frame and the image of the sample prediction frame are determined. The local detection model is trained with the reduction of the first difference and the second difference as the training goal.

[0185] When obtaining sample key frames, for greater diversity and specificity, teeth can be grouped by tooth type and sample key frames can be selected from each group. For tooth types with high detection error rates, a larger number of detection frames from these teeth can be selected as sample key frames. Tooth types include central incisors, lateral incisors, canines, molars, and more.

[0186] When determining the sample detection image, similar to steps S1104 to S1106, the sample key frame can be obtained first, and for each sample key frame, the sample detection area of ​​the sample key frame is determined to obtain the sample detection image of the sample detection area. This manual will not go into details about this.

[0187] For the annotation box of the sample key frame, some parts of the teeth will be within the detection area of ​​the annotation box of the sample key frame. If it is only a small part, then the small part of the tooth structure will be ignored in the sampled data. Suppose the complete segmentation annotation of a tooth is m and the intercepted segmentation annotation is m-part. The ignoring standard is volume(m-part) / volume(m)<0.5. Otherwise, it is the target to be detected.

[0188] This manual proposes a two-stage 3D tooth detection method based on deep learning. In the first stage, the entire tooth area is converted to a lower-resolution size and a coarse detection is performed on the entire area. This coarse detection can correctly detect the positions of most teeth, but it often results in missed and false detections. The coarse detection results provide good information on the overall position of the dentition. Using these coarse detection results, several key locations are selected, and localized images are captured within the original image space. A fine detection is then performed at each key location at high resolution. The coarse and fine detection results are then fused to obtain the final result.

[0189] The following example illustrates a possible implementation of the present application.

[0190] 1. One-stage low-resolution global detection

[0191] One possible implementation method is to determine the tooth area through the coarse segmentation result of the teeth, intercept the tooth area, and resize this area to (128,128,128) size. Perform a detection in the global range, restore the detection frame to the original size space, and obtain the first stage detection frame result, as shown in Figure 2. Detection frame 1 is the detection result of the upper dentition, detection frame 2 is the detection result of the lower dentition, and the white frame is the real frame of the undetected teeth. For intuitive understanding, the tooth instance segmentation annotation is also displayed.

[0192] 2. Key location selection

[0193] One possible implementation method is to combine the canine and incisor empirical position boxes with high error rates with the detection boxes with poor results screened out by the model as key boxes based on the results of the first stage. The specific steps are as follows:

[0194] (1) The results of the first stage are divided into two groups, the upper and lower dentitions, and sorted within each group. The center coordinates of the detection frame (x, y, z) are recorded. The detection frames within the group are sorted based on y, which is the Y-axis distance in Figure 2. The sorted detection frame sequence is recorded as L. The positions of the tooth sequence at the beginning and end of L are more accurate based on the X-axis distance. For the first and last 4 tooth detection frames in L, the sorting order is readjusted based on x to obtain the final sorted detection frame sequence L. The elements in L are the detection frames from left to right in Figure 2. The upper dentition is recorded as L1 and the lower dentition is recorded as L2.

[0195] (2) Model screening of poor quality detection frames. A classification model is trained to classify the images captured by each detection frame to determine whether the detection frame correctly surrounds the target tooth. The detection frame judged as incorrect by the model is selected as the key frame, such as detection frame 3 shown in Figure 2, and will be deleted in the final fusion stage. Positive samples can be easily obtained by annotation. Since there are not enough negative samples, negative samples are generated artificially for training. As shown in Figures 9 and 10, the generation method of negative samples is as follows:

[0196] (a) Simulate the situation where the prediction box is between two adjacent teeth or includes two teeth. Randomly select a tooth real frame and its adjacent tooth real frame from the annotation, and obtain the minimum bounding box of the two in the XY plane as shown in Figure 9. Let its height and width be (h, w), and the frame coordinates be (X1, Y1, X2, Y2). Randomly scale h and w with the center position unchanged, with a ratio r = random (-0.05, 0.15). The scaling method is shown in Figure 9. When r>0, it is reduced, and when r<0, it is enlarged. The target box is generated as shown by the dotted line. In the XY plane, the target box coordinates are (x1, y1, x2, y2), then x1 = X1 + h*r, y1 = Y1 + w*r, x2 = X2 - h*r, y2 = Y2 - w*r

[0197] The z-axis coordinates of the generated target frame are random values ​​between the z-axis coordinates of the corresponding real frame, as shown in Figure 9. z1 = random(min(BZ1, GZ1), max(BZ1, GZ1)), z2 = random(min(BZ2, GZ2), max(BZ2, GZ2)).

[0198] (b) Simulate the case where the prediction box is located between the upper and lower dentition. A random tooth real frame is selected. If it is the upper dentition, it is translated downward along the Z axis by half the frame height. Then, the target frame is randomly scaled down with a center height unchanged, using a scaling factor of r = random(-0.25, 0). The same applies to the lower dentition, as shown in Figure 10.

[0199] (3) Selection of the experience position frame.

[0200] Incisors and canines are densely populated, so there's a high probability of incorrect detection results in the first phase of detection. Therefore, these positions are selected as supplementary key frames by default in the second phase of detection. Based on the sorted upper and lower dentition L1 and L2, taking L1 as an example and L2 in the same way, we calculate the middle position index c, and select L1[c] as the incisor frame:

[0201] (a) If the center of L1[c] is to the left of the midline, then L1[c+1] is selected as another incisor frame, L1[c-2],

[0202] L1[c+3] was selected as the canine frame.

[0203] (b) If the center of L1[c] is to the right of the midline, then L1[c-1] is selected as another incisor frame, L1[c-3],

[0204] L1[c+2] is selected as the canine frame.

[0205] (c) If the iou (intersection-over-union) of two adjacent boxes in L1 on the xy plane is less than 0.05, there is a high probability that there is a missed target between the two boxes, and the one closer to the midline is selected as the key box.

[0206] 3. Two-stage local area detection

[0207] As shown in Figure 4, for each key box, the minimum bounding box of the target area target-box is first calculated, and the detection area detect-box (second recognition area) with a size of 128x128x128 is obtained with the center of the target-box as the center, as shown in the yellow box in Figure 4. As shown in Figure 5, the results of the two-stage local detection are performed within the detect-box, and the local detection results of all key positions are combined as the final result of the two-stage detection. The specific operations are as follows:

[0208] (1) Calculate the target-box. As shown in Figure 4, the detection box 4 is the selected key box, and the two adjacent predicted detection boxes 6 (if they exist) and (if they exist) detection box 5 in L2 are the target-box. The purpose of doing so is to include the target that the second stage local detection is responsible for detecting in the target-box as much as possible, which means that the target must appear in the detect-box and be located at its center. When there is a missing tooth, for example, there is a missing tooth between detection box 6 and detection box 4. Although the two are adjacent in sequence in L2, they are actually far apart. If the target-box is calculated at this time, the target position, that is, the area near detection box 4, will deviate from the center of the target-box, which is not conducive to detection accuracy. Therefore, when calculating the target-box, the adjacent boxes of the key box are only considered when they are close enough to the key box. The definition of being close enough to the key box is as follows, and it only needs to meet one of the conditions:

[0209] (a) The two boxes intersect.

[0210] As shown in Figure 3.

[0211] (b) The frame is translated along the X-axis toward the key frame by a distance of its own height (as shown in Figure 3 where frame 1 is moved to the dotted line position) and then intersects with the key frame (thick outline frame 1).

[0212] (c) Translate along the Y-axis toward the key frame by a distance of its own width (from the position of frame 2 to the dotted line position in Figure 3) and then intersect with the key frame (thick outline frame 2).

[0213] (2) Dynamic scale adjustment. For some images with higher resolution, the target-box size is larger and will not be completely contained in the detect-box of size (128,128,128). This means that the target to be detected may not be completely contained in the detect-box, resulting in inaccurate final results. For some images with lower resolution, the size of one dimension may be less than 128 and cannot accommodate the detect-box. In order to ensure that the detect-box can completely contain the target-box and can be contained by the image, the scale of the image can be dynamically adjusted according to the image size and the size of each target-box. The dynamic scaling ratio is denoted as r_dynamic, and its calculation method is as follows:

[0214] (a) Assume that the three dimensions of the image are H_data, W_data, and D_data, r = 128 / min (H_data, W_data, D_data),

[0215] If r>=1, it means that the image size is too small to accommodate the detect-box. The image size can be enlarged, and r_dynamic=r.

[0216] (b) If r<1, r_dynamic can be calculated dynamically based on the target-box. Let the three dimensions of the target-box be H_target, W_target, and D_target.

[0217] r_dynamic=128 / max(H_target,W_target,D_target). If r_dynamic<1, it means that the target-box size is large and the detect-box cannot accommodate the target-box. The image size can be reduced, and r_dynamic is rounded down to one decimal place. If r_dynamic>=1, it means that the detect-box can accommodate the target-box and no adjustment is required. r_dynamic=1.

[0218] According to r_dynamic, the detection of each key position is dynamically resized, and the center coordinates of the corresponding target-box (x0, y0, z0) are dynamically adjusted to tc = (x0, y0, z0) * r_dynamic.

[0219] (3) Perform local detection within the detect-box. In the data-dynamic space, the detect-box is determined with tc as the center and 128 as the size. The center area center-area is also determined with tc as the center and 64 as the size. Because it is in the data-dynamic space, the detection box boxes = boxes / r_dynamic is restored to the original size space. As shown in Figure 5, the detection results within the detect-box are detection box 7 and detection box 8. The third recognition area represents the center-area within the detect-box. The closer the detection result within the detect-box is to the center, the higher the confidence. Only the detection result with the center of mass of the box within the center-area is retained. Therefore, only the detection box 8 is retained as the local detection result of the current position. If there are multiple results in the center area, only 3 results are retained at most, and the first 3 results with the box center closest to the target-box center are taken.

[0220] As shown in Figure 8, when the target box (minimum bounding box) is close to the edge of the image, a certain dimension of the coordinate will be out of bounds when calculating the detect box. At this time, under the premise of ensuring that the detect box is (128,128,128), the coordinate of the same dimension develops in the opposite direction of the out-of-bounds. At this time, the target box is not in the center position in the detect box, and a certain degree of offset occurs. The center-area position can still retain the target tooth result.

[0221] (4) Post-processing: When all the detection results in the detect-box are aggregated, there will inevitably be overlaps. The NMS non-maximum suppression method is used to remove the overlapping results and obtain the final two-stage detection results.

[0222] (5) Generation of training samples for local detection model.

[0223] The training sample generation process is almost the same as the local detection process: group and sort L1 based on the annotation box, first group and sort L1 and L2 based on the annotation box, then select the key box, then determine the target-box, then determine r_dynamic, then determine data-dynamic / detect-box, and finally determine the crop train data / label. There are two points to note:

[0224] (a) In the selection key box, to ensure the diversity and pertinence of the sampling, the dentition is divided into five groups. Take L1 as an example, and the same is true for L2. Let the index of the center position of L1 be idx:

[0225] Group 1 = (idx-1, idx, idx+1) represents the incisor index.

[0226] Group 2 = (idx-2, idx-3) table left canine index,

[0227] Group 3 = (idx+2, idx+3) represents the right canine index.

[0228] Group 4 = range(1,idx-3) represents the indexes of the remaining teeth on the left.

[0229] Group 5 = range(idx+4,len(L1)-1) represents the indexes of the remaining teeth on the right side.

[0230] A key frame is randomly selected within each group.

[0231] (b) When cropping the tooth according to the detect-box, some parts of the tooth will be within the detect-box. If only a small part is included, the small part of the tooth structure will be ignored in the sampled data. Suppose the complete segmentation label of a tooth is m and the cropped segmentation label is m-part. The ignoring standard is volume(m-part) / volume(m)<0.5. Otherwise, it is the target to be detected.

[0232] Except for (a) and (b), the rest of the operations are the same as above.

[0233] 4. Two-stage result fusion

[0234] In one possible implementation, let the boxes filtered out by the first-stage detection box removal model (detection box 3 in Figure 2) be boxes1, and the second-stage detection box be boxes2. The fusion method of boxes1 and boxes2 is as follows:

[0235] (1) Calculate the iou matrix ious of boxes2 and boxes1. Its dimension is (N2, N1), where N2 is the number of boxes in boxes2 and N1 is the number of boxes in boxes1. max_iou = max(ious, axis = 1) represents the maximum iou value between each box in boxes2 and all boxes in boxes1. All detection boxes in boxes2 that satisfy max_iou < 0.5 are retained.

[0236] (2) As shown in (a) of Figure 6,

[0237] iou_z=max{0,[min(rz2,gz2)-max(rz1,gz1)] / [max(rz2,gz2)-min(rz1,gz1)]},

[0238] That is, the intersection-over-union ratio along the z direction. If the iou of box2 (from boxes2) and box1 (from boxes1) is greater than 0.5 and iou_z is greater than 0.9, because box2 comes from the local detection in the second stage and has a higher confidence than box1, box2 is retained and box1 is deleted.

[0239] (3) As shown in (b) of Figure 6, there will be a situation of concentric boxes, that is, for the same target, box1 gives a large range, and box2 gives a more precise range. Box1 and box2 have similar three-dimensional proportions and close center coordinates. In this case, box2 is retained and box1 is deleted. Determine whether box2 and box1 are concentric boxes:

[0240] (a) box2 is completely contained by box1

[0241] (b) Keep the center unchanged and shrink box1 until the size of a certain dimension of box1 is equal to the size of the corresponding dimension of box2 for the first time. At this time, calculate the IOU of box1 and box2, and IOU>0.75.

[0242] Just satisfy (a) and (b).

[0243] After (1)(2)(3), the detection results left in boxes1 and boxes2 can be determined and combined into the final detection results, as shown in Figure 7. Compared with Figure 2, it can be seen that not only the false positives in the first-stage results are removed, but also the two missed false negatives are supplemented.

[0244] A two-stage detection architecture combines low-resolution global detection with high-resolution local detection. A keybox selection method based on model screening and empirical position analysis of first-stage results provides key information about the detection area for subsequent second-stage detection, while also removing false positives from the first stage. The process of calculating the target-box, dynamically adjusting the image scale, and finally determining the second-stage detection area through the keybox provides a precise detection range, ensuring that the detection target falls within this range. Furthermore, limiting the central area can resolve the inherent false positive problem of second-stage detection. Two-stage result fusion methods include replacing highly overlapping first-stage results with second-stage results and processing concentric boxes.

[0245] The above is a tooth recognition method provided in one or more embodiments of this specification.

[0246] This specification also provides a computer-readable storage medium, which stores a computer program. The computer program can be used to execute the tooth recognition method provided in FIG. 1 .

[0247] This specification provides a tooth identification device, which includes:

[0248] a low-resolution three-dimensional image acquisition module, configured to determine, based on the original three-dimensional oral image, a low-resolution three-dimensional image including a target area, wherein the target area includes at least one of a plurality of teeth, gums, dentition, and jaw in the original three-dimensional oral image;

[0249] a dentition position information determination module, configured to perform image recognition on the low-resolution three-dimensional image to obtain the patient's dentition position information;

[0250] a key position information determination module, configured to determine key position information of teeth in the original oral three-dimensional image based on the dentition position information;

[0251] A local image determination module is used to obtain a local image corresponding to the key position information according to the key position information of the tooth, wherein the local image is a high-resolution three-dimensional image;

[0252] a first recognition result determination module, configured to perform image recognition on the high-resolution three-dimensional image to obtain a key tooth recognition result of the partial image;

[0253] The second recognition result determination module is used to determine the tooth recognition result of the original oral three-dimensional image based on the dentition position information and at least one key tooth recognition result.

[0254] Optionally, the dentition position information determination module is specifically used to input the low-resolution three-dimensional image into a first tooth recognition model to obtain a target tooth recognition result output by the first tooth recognition model, wherein the target tooth recognition result includes the detection frame of each tooth in the target area and the position information of each tooth detection frame; based on the target tooth recognition result, the patient's dentition position information is determined.

[0255] Optionally, the key position information determination module is specifically used to sort the detection frames of each tooth according to the dentition position information to obtain sorted detection frames; and determine the key teeth and the key position information of the key teeth in the sorted detection frames.

[0256] Optionally, the dentition includes an upper dentition and a lower dentition;

[0257] The key position information determination module is specifically used to determine the first dimension position data and the second dimension position data of each detection frame based on the position information of the detection frame in the dental position information; sort the detection frame of each tooth according to the second dimension position data of each detection frame to obtain an initial sorting result; adjust the order of the detection frames of the specified teeth in the initial sorting result according to the first dimension position data of each detection frame to obtain a final sorting result, and the final sorting result includes the sorted detection frames of the upper dental arch and the sorted detection frames of the lower dental arch.

[0258] Optionally, the key teeth include at least one of teeth corresponding to preset tooth numbers, teeth corresponding to classification results of classifying the sorted detection frames based on an image classification model, and teeth corresponding to preset tooth categories.

[0259] Optionally, the key position information determination module is specifically used to determine the midline position information of each dental arch in the sorted detection frame based on the second dimensional position data of the detection frame of each tooth in the dental arch; determine the first key tooth based on the midline position information, and determine the position information of the first key tooth as the first key position information; determine the first adjacent frame and the second adjacent frame adjacent to the first key frame of the first key tooth; determine the first adjacent intersection and union ratio of the first key frame to the first adjacent frame, and determine the second adjacent intersection and union ratio of the first key frame to the second adjacent frame; determine the second key frame in the first adjacent frame and the second adjacent frame based on the first key position information, the first adjacent intersection and union ratio, and determine the position information of the second key tooth as the second key position information; based on the first key position information and the position information of the detection frames of each tooth in the dental arch, determine the tooth in the dental arch whose positional relationship with the first key tooth is a specified positional relationship as the third key tooth, and set the position information of the third key tooth as the third key position information.

[0260] Optionally, the key position information determination module is specifically used to input the image of each detection frame into an image classification model to obtain a classification result of the image of the detection frame output by the image classification model; when the classification result is that the target tooth in the detection frame is not correctly surrounded, the target tooth is determined as a key tooth, the detection frame is determined as a key frame, and the position information of the key frame is determined as key position information, and correct encirclement means that the detection frame includes a complete tooth.

[0261] Optionally, the local image determination module is specifically used to determine, for each key frame, two detection frames adjacent to the key frame in the sorted detection frames to obtain adjacent detection frames; when the adjacent detection frames intersect with the key frame, or when the adjacent detection frames intersect with the key frame after being translated by a preset distance in a preset dimension, determine the minimum circumscribed frame of the adjacent detection frames and the key frame as the first recognition area containing the key frame; and based on the first recognition area and the key position information of the teeth, obtain the local image corresponding to the key position information.

[0262] Optionally, the local image determination module is specifically used to determine the dynamic scaling ratio based on the preset first size of the second recognition area, the size of the first recognition area and the image size of the original oral three-dimensional image; and determine the first center position information of the first recognition area according to the size of the first recognition area; scale the original oral three-dimensional image according to the dynamic scaling ratio to obtain a scaled oral three-dimensional image; and determine the second center position information of the scaled first recognition area based on the dynamic scaling ratio and the first center position information; in the scaled oral three-dimensional image, determine the second recognition area based on the second center position information and the preset first size; and determine the oral three-dimensional image including the second recognition area as the local image corresponding to the key position information.

[0263] Optionally, the local image determination module is specifically used to determine the minimum scale value of the original oral three-dimensional image in each dimension based on the image size of the original oral three-dimensional image; when the minimum scale value is not greater than the preset first size, the image magnification ratio is determined based on the minimum scale value and the preset first size, and used as the dynamic scaling ratio; when the minimum scale value is greater than the preset first size, the maximum scale value of the first recognition area in each dimension is determined based on the size of the first recognition area, and the image reduction ratio is determined based on the maximum scale value and the preset first size, and used as the dynamic scaling ratio.

[0264] Optionally, the first recognition result determination module is specifically used to input the high-resolution three-dimensional image into a second tooth recognition model to obtain a local tooth recognition result of the local image output by the second tooth recognition model, and the local tooth recognition result includes a plurality of detection frames and position information of the detection frames; according to the second center position information and the preset second size, a third recognition area is determined; when the center of mass of the detection frame in the local tooth recognition result is located within the third recognition area, the detection frame in the local tooth recognition result is scaled according to the dynamic scaling ratio as the key tooth recognition result of the local image.

[0265] Optionally, the second recognition result determination module is specifically used to determine the detection frames other than the key frame in the detection frames of each tooth in the dental position information as the first frame to be filtered, and to use at least one detection frame in the key tooth recognition result as the second frame to be filtered; for each second frame to be filtered, determine the first intersection and union ratio of the second frame to be filtered and each of the first frames to be filtered in the plane formed by the first dimension and the second dimension; when all the first intersection and union ratios of the second frame to be filtered are less than the preset first intersection and union ratio, retain the second frame to be filtered and determine it as the tooth recognition result of the original oral three-dimensional image; or determine the second intersection and union ratio of the second frame to be filtered and each of the first frames to be filtered in the third dimension; when the first intersection and union ratio of the second frame to be filtered is greater than the preset first intersection and union ratio and the second intersection and union ratio is greater than the preset second intersection and union ratio, retain the second frame to be filtered and determine it as the tooth recognition result of the original oral three-dimensional image.

[0266] Optionally, the second recognition result determination module is specifically used to determine the detection frames other than the key frame in the detection frames of each tooth in the dental position information as the first frame to be filtered, and to use at least one detection frame in the key tooth recognition result as the second frame to be filtered; for each first frame to be filtered, when the first frame to be filtered completely contains the second frame to be filtered, the sizes of all dimensions of the first frame to be filtered are reduced proportionally at the same time, and when the size of any dimension of the reduced first frame to be filtered is consistent with the size of the corresponding dimension of the included second frame to be filtered, the third intersection-and-union ratio of the reduced first frame to be filtered and the included second frame to be filtered is determined; when the third intersection-and-union ratio is greater than the preset third intersection-and-union ratio, the included second frame to be filtered is retained and determined as the tooth recognition result of the original oral three-dimensional image.

[0267] Optionally, the device further comprises:

[0268] A training module is used to obtain several positive samples, each of which includes an image of a sample detection frame, wherein the positive sample is a sample in which the image of the sample detection frame only contains one complete target tooth; for each positive sample, determine the sample minimum bounding box of the sample detection frame of the positive sample and any several adjacent sample detection frames; scale the length of the first dimension and the length of the second dimension of the sample minimum bounding box according to a preset scaling ratio to obtain the scaled sample minimum bounding box; determine the image of the scaled sample minimum bounding box to obtain a first negative sample; translate the sample detection frame of the positive sample according to a preset translation direction, and scale the translated sample detection frame to obtain a sample detection frame after the translation and scaling operations; determine the image of the sample detection frame after the translation and scaling operations to obtain a second negative sample; add the positive sample, the first negative sample, and the second negative sample to a training sample set; input the samples in the training sample set into an image classification model to obtain a predicted classification result output by the image classification model; and train the image classification model based on the predicted classification result, the positive sample, the first negative sample, and the second negative sample.

[0269] This specification also provides a structural diagram of the electronic device shown in Figure 13. As shown in Figure 13, at the hardware level, the device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and of course may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the method described in the above embodiment. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0270] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0271] The memory may include volatile memory, such as random access memory (RAM) and / or cache memory, and may further include read-only memory (ROM).

[0272] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0273] The memory may also include a program tool (or utility) having a set (at least one) of program modules, such program modules including but not limited to: an operating system, one or more application programs, other program modules and program data, each of which or some combination may include an implementation of a network environment.

[0274] The processor executes various functional applications and data processing by running the computer programs stored in the memory, such as the method in the above embodiment.

[0275] Input and / or output devices may include a scanning device, a camera interface, an input device (e.g., a mouse, keyboard, etc.), a display device (e.g., a monitor), a printer, and / or one or more other input devices. The input / output interface may receive executable instructions and / or data, which may be stored in a data storage device (e.g., a memory). For example, it may be used to receive or store oral model data or a target file.

[0276] In some embodiments, the scanning device can be configured to scan one or more physical dental models of a patient's dentition. In one or more embodiments, the scanning device can be configured to directly scan a patient's dentition and / or dental appliances. The scanning device can be configured to input the data into a computing device.

[0277] In some embodiments, the camera interface can receive input from an imaging device (e.g., a 2D or 3D imaging device), such as a digital camera, a print photo scanner, and / or other suitable imaging device. For example, the input from the imaging device can be stored in a memory.

[0278] The processor may execute instructions to provide a display of the oral model on a display, etc. For example, the computing device may be configured to allow a treating professional or other user to input treatment goals. The received input may be sent to the processor as data and / or may be stored in a memory.

[0279] Such connectivity may allow for the input and / or output of data and / or instructions and other types of information.Some embodiments may be distributed among various computing devices within one or more networks and used to collect, calculate, and / or analyze any of the methods described herein.

[0280] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the present application, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0281] This embodiment also relates to a readable medium, which is specifically a non-transitory computing device readable medium, and stores instructions that can be executed by a processor to enable the computing device to perform the above generation method.

[0282] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0283] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of this application.

Claims

1. A tooth recognition method, characterized in that: The method comprises: Determine, based on the original oral three-dimensional image, a low-resolution three-dimensional image including a target area, wherein the target area includes at least one of a plurality of teeth, gums, dentition, and jaw in the original oral three-dimensional image; Performing image recognition on the low-resolution three-dimensional image to obtain dentition position information of the patient; Determine key position information of teeth in the original three-dimensional oral image according to the dentition position information; According to the key position information of the tooth, a local image corresponding to the key position information is obtained, wherein the local image is a high-resolution three-dimensional image; Performing image recognition on the high-resolution three-dimensional image to obtain a key tooth recognition result of the local image; The tooth recognition result of the original oral three-dimensional image is determined according to the dentition position information and at least one key tooth recognition result.

2. The method according to claim 1, characterized in that Performing image recognition on the low-resolution three-dimensional image to obtain the patient's dentition position information specifically includes: Inputting the low-resolution three-dimensional image into a first tooth recognition model to obtain a target tooth recognition result output by the first tooth recognition model, wherein the target tooth recognition result includes a detection frame of each tooth in the target area and position information of each tooth detection frame; According to the target tooth recognition result, the patient's dentition position information is determined.

3. The method according to claim 2, characterized in that Determining key position information of teeth in the original three-dimensional oral cavity image according to the dentition position information specifically includes: sorting the detection frames of the teeth according to the dentition position information to obtain sorted detection frames; In the sorted detection frame, key teeth and key position information of the key teeth are determined.

4. The method according to claim 3, characterized in that The dentition includes the upper and lower dentition; According to the dentition position information, the detection frames of the teeth are sorted to obtain the sorted detection frames, which specifically includes: Determine first dimension position data and second dimension position data of each detection frame according to the position information of the detection frame in the dentition position information; Sorting the detection frames of the teeth according to the second dimensional position data of the detection frames to obtain an initial sorting result; According to the first dimensional position data of each detection frame, the order of the detection frames of the specified teeth in the initial sorting result is adjusted to obtain a final sorting result, and the final sorting result includes the sorted detection frames of the upper dentition and the sorted detection frames of the lower dentition.

5. The method according to claim 3, characterized in that The key teeth include at least one of teeth corresponding to a preset tooth number, teeth corresponding to a classification result of classifying the sorted detection frame based on an image classification model, and teeth corresponding to a preset tooth category.

6. The method according to claim 3, characterized in that In the sorted detection frame, determining the key teeth and the key position information of the key teeth specifically includes: For each dentition, in the sorted detection frame, determine the midline position information based on the second dimensional position data of the detection frame of each tooth in the dentition; Based on the midline position information, determining a first key tooth, and determining the position information of the first key tooth as first key position information; determining a first adjacent frame and a second adjacent frame adjacent to the first key frame of the first key tooth; Determine a first adjacent intersection-and-union ratio between the first key box and a first adjacent box, and determine a second adjacent intersection-and-union ratio between the first key box and a second adjacent box; Based on the first key position information, the first adjacent intersection-and-union ratio, and the second adjacent intersection-and-union ratio, a second key frame is determined in the first adjacent frame and the second adjacent frame, and the position information of the second key tooth is determined as the second key position information; Based on the first key position information and the position information of the detection frames of each tooth in the dental arch, in the dental arch, the tooth whose position relationship with the first key tooth is a specified position relationship is determined to be the third key tooth, and the position information of the third key tooth is set as the third key position information.

7. The method according to claim 3, characterized in that In the sorted detection frame, determining the key teeth and the key position information of the key teeth specifically includes: For each detection frame, input the image of the detection frame into the image classification model to obtain a classification result of the image of the detection frame output by the image classification model; When the classification result is that the target tooth in the detection frame is not correctly surrounded, the target tooth is determined as a key tooth, the detection frame is determined as a key frame, and the position information of the key frame is determined as key position information, and the correct encirclement means that the detection frame includes a complete tooth.

8. The method according to claim 7, characterized in that According to the key position information of the tooth, a local image corresponding to the key position information is obtained, specifically including: For each key frame, determine two detection frames adjacent to the key frame in the sorted detection frames to obtain adjacent detection frames; When the adjacent detection box intersects with the key box, or When the adjacent detection frame intersects with the key frame after being translated by a preset distance in a preset dimension, determining a minimum circumscribed frame of the adjacent detection frame and the key frame as a first recognition area including the key frame; According to the first recognition area and the key position information of the tooth, a local image corresponding to the key position information is obtained.

9. The method according to claim 8, characterized in that Obtaining a local image corresponding to the key position information according to the first recognition area and the key position information of the tooth specifically includes: Determine a dynamic zoom ratio based on a preset first size of the second recognition area, a size of the first recognition area, and an image size of the original oral three-dimensional image; and determine first center position information of the first recognition area according to the size of the first recognition area; scaling the original oral cavity three-dimensional image according to the dynamic scaling ratio to obtain a scaled oral cavity three-dimensional image; and determining the second center position information of the scaled first recognition area based on the dynamic scaling ratio and the first center position information; In the scaled three-dimensional image of the oral cavity, determining the second identification area based on the second center position information and the preset first size; The three-dimensional image of the oral cavity including the second identification area is determined as the local image corresponding to the key position information.

10. The method according to claim 9, characterized in that Determining a dynamic zoom ratio based on a preset first size of the second recognition area, a size of the first recognition area, and an image size of the original oral three-dimensional image specifically includes: Determining the minimum scale value of the original oral three-dimensional image in each dimension according to the image size of the original oral three-dimensional image; When the minimum scale value is not greater than the preset first size, determining the image magnification ratio according to the minimum scale value and the preset first size, and using the image magnification ratio as the dynamic scaling ratio; When the minimum scale value is greater than the preset first size, the maximum scale value of the first recognition area in each dimension is determined according to the size of the first recognition area, and the image reduction ratio is determined according to the maximum scale value and the preset first size, and is used as the dynamic scaling ratio.

11. The method according to claim 9, characterized in that Performing image recognition on the high-resolution three-dimensional image to obtain key tooth recognition results of the local image specifically includes: Inputting the high-resolution three-dimensional image into a second tooth recognition model to obtain a local tooth recognition result of the local image output by the second tooth recognition model, wherein the local tooth recognition result includes a plurality of detection frames and position information of the detection frames; Determining a third identification area according to the second center position information and a preset second size; When the centroid of the detection frame in the partial tooth recognition result is located within the third recognition area, the detection frame in the partial tooth recognition result is scaled according to the dynamic scaling ratio to serve as the key tooth recognition result of the partial image.

12. The method according to claim 6 or 7, characterized in that: Determining the tooth recognition result of the original oral three-dimensional image according to the dentition position information and at least one key tooth recognition result specifically includes: In the detection frames of each tooth of the dentition position information, determining the detection frames other than the key frame as the first frame to be screened, and taking at least one detection frame in the key tooth recognition result as the second frame to be screened; For each second box to be screened, determining a first intersection-and-union ratio of the second box to be screened and each of the first boxes to be screened in a plane formed by the first dimension and the second dimension; When all the first intersection-and-union ratios of the second frame to be screened are less than the preset first intersection-and-union ratios, retain the second frame to be screened and determine it as the tooth recognition result of the original oral three-dimensional image; or In the third dimension, determine the second intersection-and-union ratio of the second frame to be filtered and each of the first frames to be filtered; when the first intersection-and-union ratio of the second frame to be filtered is greater than the preset first intersection-and-union ratio and the second intersection-and-union ratio is greater than the preset second intersection-and-union ratio, retain the second frame to be filtered and determine it as the tooth recognition result of the original oral three-dimensional image.

13. The method according to claim 8, characterized in that Determining the tooth recognition result of the original oral three-dimensional image according to the dentition position information and at least one key tooth recognition result specifically includes: In the detection frames of each tooth of the dentition position information, determining the detection frames other than the key frame as the first frame to be screened, and taking at least one detection frame in the key tooth recognition result as the second frame to be screened; For each first box to be screened, when the first box to be screened completely contains the second box to be screened, the sizes of all dimensions of the first box to be screened are reduced in equal proportion at the same time, and when the size of any dimension of the reduced first box to be screened is consistent with the size of the corresponding dimension of the included second box to be screened, a third intersection-over-union ratio of the reduced first box to be screened and the included second box to be screened is determined; When the third intersection-over-union ratio is greater than a preset third intersection-over-union ratio, the second to-be-screened frame is retained and determined as a tooth recognition result of the original oral three-dimensional image.

14. The method according to claim 7, characterized in that Train an image classification model, including: Acquire a plurality of positive samples, wherein the samples include an image of a sample detection frame, and the positive samples are samples in which the image of the sample detection frame contains only one complete target tooth; For each positive sample, determine the sample minimum bounding box of the sample detection box of the positive sample and any number of adjacent sample detection boxes; Scaling the length of the first dimension and the length of the second dimension of the sample minimum bounding box according to a preset scaling ratio to obtain a scaled sample minimum bounding box; Determine the image of the minimum bounding box of the scaled sample to obtain a first negative sample; The sample detection frame of the positive sample is translated according to a preset translation direction, and the translated sample detection frame is scaled to obtain a sample detection frame after the translation operation and the scaling operation; Determine the image of the sample detection frame that has undergone translation and scaling operations to obtain a second negative sample; Adding the positive sample, the first negative sample and the second negative sample to a training sample set; Inputting samples in the training sample set into an image classification model to obtain a predicted classification result output by the image classification model; The image classification model is trained according to the predicted classification result, the positive sample, the first negative sample and the second negative sample.

15. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 14 is implemented.

16. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 14 when executing the computer program.

Citation Information

Patent Citations

  • Method and system for determination of object dentition layout

    CN106137414A

  • Automatic tooth arrangement method, device and equipment based on point cloud understanding and storage medium

    CN115719404A

  • Oral cavity three-dimensional image processing method

    CN116912412A

  • Manufacturing of dental implants based on digital scan data alignment

    US20230363732A1

Cited By

  • Security reinforcement method and system for operating system of credential mobile terminal

    CN121580382A