Oral cavity detection method and device and electronic equipment
Through the combination of transfer learning and image enhancement technology, the generalization ability and accuracy of oral detection models are improved, the problem of low accuracy of oral detection in the prior art is solved, and efficient and accurate oral detection is achieved.
Patent Information
- Application Number
- CN202510784081.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
AI Technical Summary
The existing oral detection model based on artificial intelligence is not very accurate in oral detection, and the traditional oral detection method relies on artificial experience and is inefficient.
By acquiring user oral images and identifying them using the target oral detection model, the target oral detection model is obtained by transfer learning of the pre-trained oral detection model through the first sample set, combining multitasking and image enhancement technology to improve the generalization ability and accuracy of the model.
It improves the efficiency and accuracy of oral detection, reduces dependence on the sample data marked by target electronic devices, reduces training costs, and can output user oral area location and disease detection results at the same time.
Smart Images

Figure CN120298409A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of oral care, and in particular, to an oral detection method, device, and electronic device. Background Art
[0002] Oral health, as a key component of public health, is of great importance to maintain. However, traditional oral detection methods rely on the experience and manual examination of dentists, and there are many limitations. With the rapid development of artificial intelligence technology and its wide application in various fields, it provides new possibilities for real-time automated oral detection. However, the oral detection implemented by the current oral detection model based on artificial intelligence still has the problem of low accuracy. Summary of the Invention
[0003] Embodiments of this application disclose an oral detection method, device, and electronic device, which can quickly obtain a target oral detection model corresponding to a target electronic device, improve the generalization ability of the oral detection model, and thus improve the efficiency and accuracy of oral detection.
[0004] Embodiments of this application disclose an oral detection method, the method includes: Obtain an oral image corresponding to a user's oral cavity, where the oral image is obtained by a target electronic device through a camera; Identify the oral image through a target oral detection model to obtain an oral detection result corresponding to the user's oral cavity; the target oral detection model is obtained by performing transfer learning on a pre-trained oral detection model through a first sample set; the first sample oral images included in the first sample set are obtained by the target electronic device through the camera, and the first sample oral images are labeled with sample oral detection results.
[0005] In the embodiments of this application, a small number of labeled first sample oral images taken by a target electronic device can be used to perform transfer training on a pre-trained oral detection model, which not only enables the pre-trained oral detection model to quickly adapt to the image acquisition characteristics of the target electronic device, but also reduces the dependence on the labeled sample data corresponding to the target electronic device, reduces the training cost of the target oral detection model, and improves the generalization ability and accuracy of the target oral detection model, thereby improving the efficiency and accuracy of oral detection using the target oral detection model.
[0006] As an optional implementation manner, the oral image is obtained by the target electronic device through a camera using target shooting parameters; the first sample oral images included in the first sample set are obtained by the target electronic device through the camera using the target shooting parameters.
[0007] In this embodiment, both the oral cavity image and the first sample oral cavity image are obtained by the target electronic device using the same camera and the same target shooting parameters, which can ensure that the target oral cavity detection model obtained by transfer learning from the first sample oral cavity image accurately identifies the oral cavity image, thereby improving the efficiency and accuracy of oral cavity detection using the target oral cavity detection model.
[0008] As an alternative embodiment, the target shooting parameters include at least one of the target magnification of the lens, the target aperture, the target shutter speed, the target shooting angle, and the target sensitivity.
[0009] In this embodiment, through a variety of shooting parameters, different shooting conditions and user requirements can be adapted, the richness of the oral cavity image and the first sample oral cavity image can be improved, so that the target oral cavity detection model can perform oral cavity detection on oral cavity images with a variety of target shooting parameters, which helps the oral cavity detection model to perform oral cavity detection more accurately to obtain more accurate oral cavity detection results.
[0010] As an alternative embodiment, the oral cavity detection result includes an oral cavity area detection result and an oral cavity disease detection result; the oral cavity area detection result includes the position information corresponding to each oral cavity area included in the user's oral cavity; the oral cavity disease detection result includes the types of diseases existing in the user's oral cavity and the probabilities of occurrence of each disease; the process of obtaining the oral cavity detection result corresponding to the user's oral cavity by identifying the oral cavity image through the target oral cavity detection model includes: identifying the oral cavity area of the oral cavity image through the target oral cavity detection model to obtain the oral cavity area detection result corresponding to the user's oral cavity, and performing disease detection on the oral cavity image to obtain the oral cavity disease detection result corresponding to the user's oral cavity.
[0011] In this embodiment, the oral cavity detection model performs oral cavity detection on the oral cavity image through a multi-task processing method, and can simultaneously output the position information of each oral cavity area in the user's oral cavity, as well as the types of diseases that may exist in the user's oral cavity and the probabilities of occurrence of each disease, thereby improving the efficiency of oral cavity detection.
[0012] As an alternative embodiment, the oral cavity area includes each tooth and each tooth surface included in each tooth; the position information corresponding to the oral cavity area includes the position of each tooth in the user's oral cavity and the tooth surface type corresponding to each tooth surface included in each tooth.
[0013] In this embodiment, by refining the oral cavity area to each tooth and its tooth surface, a more comprehensive and detailed oral cavity detection result can be provided for the user.
[0014] As an alternative embodiment, the regional recognition of the oral image by the target oral detection model to obtain the oral region detection result corresponding to the user's oral cavity includes: recognizing each tooth included in the oral image by the target oral detection model to obtain the tooth image corresponding to each tooth; matching the tooth image corresponding to each tooth with a plurality of preset tooth images, and determining the position of each tooth in the user's oral cavity according to the preset tooth image matched with the tooth image corresponding to each tooth; the preset tooth image is an image of a tooth at different positions in the user's oral cavity; recognizing the tooth surface of each of the tooth images by the target oral detection model to determine the tooth surface type corresponding to each tooth surface included in each tooth.
[0015] In this embodiment, the target oral detection model can obtain the tooth image corresponding to each tooth according to the oral image, and determine the position of each tooth in the user's oral cavity by matching the tooth image corresponding to each tooth with a plurality of preset tooth images; and by recognizing the tooth surface of the tooth image, the tooth surface type of each tooth in the oral image is determined, so as to obtain a more detailed and accurate oral detection result.
[0016] As an alternative embodiment, the target oral detection model includes an oral region detection module and an oral disease detection module; the regional recognition of the oral image by the target oral detection model to obtain the oral region detection result corresponding to the user's oral cavity, and the disease detection of the oral image to obtain the oral disease detection result corresponding to the user's oral cavity includes: performing regional recognition of the oral image by the oral region detection module to obtain the oral region detection result corresponding to the user's oral cavity; performing disease detection on the oral image in parallel by the oral disease detection module to obtain the oral disease detection result corresponding to the user's oral cavity.
[0017] In this embodiment, in the target oral detection model, the oral region is recognized by the oral region detection module and the oral disease is detected by the oral disease detection module respectively, so that the target oral detection model can process the input oral image in parallel and more efficiently, and improve the overall oral detection rate.
[0018] As an alternative embodiment, the recognition of the oral image by the target oral detection model to obtain the oral detection result corresponding to the user's oral cavity includes: marking the position information corresponding to each oral region in the oral image, and marking the disease types existing in the user's oral cavity and the probability of occurrence of each disease type in the oral image.
[0019] In this embodiment, by annotating the position information, the existing disease types, and the probabilities corresponding to each disease type on the oral cavity image, the oral cavity detection result can be intuitively informed to the user, helping the user better manage oral cavity health.
[0020] As an alternative embodiment, the pre-trained oral cavity detection model is pre-trained according to a second sample set; the second sample oral cavity images included in the second sample set are obtained by photographing with the cameras of one or more sample electronic devices, and the second sample oral cavity images are annotated with actual oral cavity detection results; the number of second sample oral cavity images included in the second sample set is greater than the number of first sample oral cavity images included in the first sample set.
[0021] In this embodiment, the pre-trained oral cavity detection model is pre-trained based on the second sample set with a larger number of images, and the second sample set includes second sample oral cavity images obtained by photographing with the cameras of one or more sample electronic devices and annotated with actual oral cavity detection results, so that in the pre-training stage, the pre-trained oral cavity detection model has learned a wide variety of features related to oral cavity detection, thus having a strong oral cavity detection ability; making use of this pre-trained oral cavity detection model as a basis, the target electronic device can perform transfer learning on the pre-trained oral cavity detection model based on a small number of first sample oral cavity images, and can quickly adjust and obtain the target oral cavity detection model to adapt to the target electronic device, thereby reducing the dependence on the labeled sample data corresponding to the target electronic device, reducing the training cost of the target oral cavity detection model, and improving the generalization ability of the target oral cavity detection model.
[0022] As an alternative embodiment, the training process of the pre-trained oral cavity detection model includes: inputting a second sample oral cavity image into the oral cavity detection model to be trained; extracting image features from the current second sample oral cavity image through the oral cavity detection model to be trained, obtaining image features, and determining the predicted oral cavity detection result corresponding to the current second sample oral cavity image according to the image features; adjusting the parameters of the oral cavity detection model to be trained according to the actual oral cavity detection result corresponding to the current second sample oral cavity image and the corresponding predicted training oral cavity detection result until the training completion condition is met, and obtaining the pre-trained oral cavity detection model.
[0023] In this embodiment, the oral cavity detection model to be trained is pre-trained with a large number of second sample oral cavity images, improving the accuracy and adaptability of the pre-trained oral cavity detection model and the performance of the pre-trained oral cavity detection model.
[0024] As an alternative implementation, determining the predicted oral detection result corresponding to the current second-sample oral image based on the image features includes: performing structural re-parameterization on the image features; and determining the predicted oral detection result corresponding to the current second-sample oral image according to the structurally re-parameterized image features.
[0025] In this implementation, the oral detection model to be trained improves the computational efficiency and parameter utilization rate of the pre-trained oral detection model by performing structural re-parameterization on the image features of the second-sample oral images, enabling the pre-trained oral detection model to be deployed on more electronic devices and improving the applicable range of the pre-trained oral detection model.
[0026] As an alternative implementation, the pre-trained oral detection model is pre-trained according to the enhanced second-sample set; the method further includes: performing image enhancement on the second-sample oral images included in the second-sample set to obtain the enhanced second-sample set.
[0027] In this implementation, by performing image enhancement processing on the second-sample oral images, more abundant and diverse training data is provided for the oral detection images to be trained, thereby improving the adaptability and generalization ability of the pre-trained oral detection model to different image conditions.
[0028] As an alternative implementation, performing image enhancement on the second-sample oral images included in the second-sample set includes: performing color enhancement on one or more second-sample oral images in the second-sample set; and / or, performing geometric transformation on one or more second-sample oral images in the second-sample set; and / or, performing mosaic enhancement on one or more second-sample oral images in the second-sample set.
[0029] In this implementation, by performing image enhancement on the second-sample oral images through multiple enhancement methods, more diverse training data can be obtained, thereby improving the generalization ability of the pre-trained oral detection model.
[0030] As an alternative implementation, performing image enhancement on the second-sample oral images included in the second-sample set to obtain the enhanced second-sample set includes: performing image enhancement on the second-sample oral images included in the second-sample set to obtain the enhanced second-sample oral images; obtaining the actual oral detection results corresponding to each of the enhanced second-sample oral images; and adding the enhanced second-sample oral images marked with the actual oral detection results to the second-sample set to obtain the enhanced second-sample set.
[0031] In this embodiment, the enhanced second sample oral image is re-annotated so that the actual oral detection result in the enhanced second sample oral image is more accurate. Thus, during the training process of the oral detection model to be trained, the parameter optimization of the model is more precise, thereby improving the accuracy of the output of the pre-trained oral detection model.
[0032] An embodiment of the present application discloses an oral detection device, which includes: An image acquisition module, configured to acquire an oral image corresponding to a user's oral cavity, where the oral image is obtained by the target electronic device through a camera; An oral detection module, configured to identify the oral image through a target oral detection model to obtain an oral detection result corresponding to the user's oral cavity; the target oral detection model is obtained by performing transfer learning on a pre-trained oral detection model through a first sample set; the first sample oral images included in the first sample set are obtained by the target electronic device through the camera, and the first sample oral images are annotated with sample oral detection results.
[0033] In the embodiment of the present application, the oral detection device can perform transfer training on the pre-trained oral detection model by using a small number of first sample oral images captured by the target electronic device with annotations. This not only enables the pre-trained oral detection model to quickly adapt to the image acquisition characteristics of the target electronic device but also reduces the dependence on the annotated sample data corresponding to the target electronic device, reduces the training cost of the target oral detection model, and improves the generalization ability and accuracy of the target oral detection model, thereby improving the efficiency and accuracy of oral detection using the target oral detection model.
[0034] An embodiment of the present application discloses an electronic device, including a memory and a processor. When a computer program stored in the memory is executed by the processor, the processor implements the method described in any of the above embodiments.
[0035] An embodiment of the present application discloses a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method described in any of the above embodiments is implemented.
[0036] An embodiment of the present application discloses a computer program product, including a computer program. When the computer program is executed by a processor, the method described in any of the above embodiments is implemented. Description of the Drawings
[0037] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0038] Figure 1A It is an application scenario diagram of the oral cavity detection method in an embodiment; Figure 1B It is a schematic diagram of the training process of the target oral cavity detection model in an embodiment; Figure 2 It is a flowchart of the oral cavity detection method in an embodiment; Figure 3A It is a flowchart of the oral cavity detection method in another embodiment; Figure 3B It is a schematic diagram of the tooth distribution in the user's oral cavity in an embodiment; Figure 4A It is a flowchart of identifying the oral cavity area from an oral cavity image in an embodiment; Figure 4B It is a schematic diagram of the mandibular teeth of an oral cavity image in an embodiment; Figure 5 It is a schematic diagram of the training process of the pre-trained oral cavity detection model in an embodiment; Figure 6 It is a schematic diagram of the training process of the pre-trained oral cavity detection model in another embodiment; Figure 7 It is a block diagram of the oral cavity detection device in an embodiment; Figure 8 It is a block diagram of the structure of an electronic device in an embodiment. Detailed implementation manners
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0040] It is understood that the terms "first", "second", etc. used in the present application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, without departing from the scope of the present application, a first sample oral image may be referred to as a second sample oral image, and similarly, a second sample oral image may be referred to as a first sample oral image. Both the first sample oral image and the second sample oral image are images containing the interior of the oral cavity and the tooth area, but they are not the same images.
[0041] When users examine their oral cavity, they may use a variety of different electronic devices to achieve the oral examination function. For example, users may use electric toothbrushes, oral endoscopes, oral scanners, etc. to examine their oral cavity, or they may use terminal devices (such as mobile phones, etc.) to examine their oral cavity. Since oral images obtained by different electronic devices based on different shooting parameters will have different image features, the oral examination model has poor adaptability to oral images collected by different electronic devices and low accuracy.
[0042] In the related art, for different electronic devices, each electronic device needs to collect a large number of sample oral images and annotate them, and then use the annotated sample oral images to train the oral detection model, thereby improving the accuracy of the oral detection model. However, annotating a large number of sample oral images leads to high costs, which wastes human resources and time, and limits the application of oral detection models.
[0043] The embodiments of the present application disclose an oral cavity detection method, device and electronic device, which can quickly obtain a target oral cavity detection model corresponding to the target electronic device, improve the generalization ability of the oral cavity detection model, and thus improve the efficiency and accuracy of oral cavity detection.
[0044] Figure 1A FIG. 1 is an application scenario diagram of an oral cavity detection method in an embodiment. Figure 1A As shown, the oral cavity detection method provided in the embodiment of the present application can be applied to the electronic device 110.
[0045] The electronic device 110 may include but is not limited to oral care equipment, terminal equipment, etc. The oral care equipment may include but is not limited to oral cleaning equipment, oral detection equipment, etc. The oral cleaning equipment may include but is not limited to electric toothbrushes, water flossers, dental scalers, etc.; oral detection equipment may include but is not limited to oral cameras, oral endoscopes, oral scanners, breath detectors, saliva detectors, etc. The terminal equipment may include but is not limited to mobile phones, smart wearable devices, tablet computers, etc., and may also be a control terminal or detection terminal used in conjunction with the oral care equipment, such as a central control device, a smart home terminal, etc., which is not limited in the embodiments of this application.
[0046] The electronic device 110 may further include a camera. For example, when the electronic device 110 is a smart phone, the camera may be the front camera or the rear camera of the smart phone.
[0047] In some embodiments, a user or a researcher adjusts the shooting parameters of the camera on the electronic device 110, and the electronic device 110 can control the camera to capture an oral image according to the adjusted shooting parameters.
[0048] In some embodiments, a target oral detection model may be stored in the electronic device 110. The electronic device 110 can capture an oral image corresponding to the user's oral cavity 120 through the camera; and then identify the oral image through the target oral detection model to obtain an oral detection result corresponding to the user's oral cavity 120. Among them, the target oral detection model is obtained by performing transfer learning on a pre-trained oral detection model through a first sample set.
[0049] Optionally, the electronic device storing the target oral detection model and the electronic device configured with a camera may be the same electronic device or different electronic devices.
[0050] In some embodiments, assuming that the electronic device storing the target oral detection model and the electronic device configured with a camera are the same electronic device, the configured camera of the electronic device can be controlled to capture an oral image of the user's oral cavity, and then the electronic device directly uses the stored target oral detection model to identify the captured oral image to obtain an oral detection result, so as to complete the acquisition of the oral image and the user's oral cavity detection on one electronic device. For example, a smart phone storing the target oral detection model can capture an oral image through the rear camera and identify the oral image through the stored target oral detection model on the smart phone to obtain an oral detection result. The integrated electronic device can reduce the time and risk of data transmission, thereby improving the efficiency and accuracy of oral cavity detection.
[0051] In some embodiments, assuming that the electronic device storing the target oral detection model and the electronic device configured with a camera are different electronic devices, the electronic device configured with a camera can first control the camera to capture an oral image of the user's oral cavity and send the oral image to the electronic device storing the target oral detection model. The electronic device storing the target oral detection model performs oral cavity detection on the received oral image through the target oral detection model, so as to distribute the acquisition of the oral image and the user's oral cavity detection on different electronic devices, enabling the user's oral cavity detection to be performed remotely, improving the convenience of oral cavity detection. Moreover, since most electronic devices are equipped with cameras, remote oral cavity detection can be applied to various scenario requirements.
[0052] In some embodiments, Figure 1B FIG. Figure 1B is a schematic diagram of the training process of the target oral detection model. Among them, the pre-trained oral detection model 160 is obtained by training the to-be-trained oral detection model 140 with the second sample oral images 150 captured by at least one camera. The electronic device 110 can perform transfer learning on the pre-trained oral detection model 160 with the first sample oral images 170 captured by the camera, and transfer the pre-trained oral detection model 160 into the electronic device 110 to obtain the target oral detection model 180, so that the target oral detection model 180 is applicable to the electronic device 110.
[0053] It should be noted that the second sample oral images 150 can be captured by different cameras of multiple identical or different electronic devices; the cameras for capturing the second sample oral images 150 and the cameras for capturing the first sample oral images 170 can be different cameras or the same camera.
[0054] As Figure 2 shown, in one embodiment, an oral detection method is provided, which can be applied to the above-mentioned electronic device. The method may include the following steps 210 to step 220.
[0055] Step 210, obtain an oral image corresponding to the user's oral cavity, where the oral image is captured by the camera of the target electronic device.
[0056] The oral image refers to an image inside the user's oral cavity. The oral image may include multiple parts inside the user's oral cavity, including but not limited to one or more of teeth, gums, tongue, oral mucosa, and other oral structures. The target electronic device is the above-mentioned Figure 1A shown electronic device.
[0057] In some embodiments, the target electronic device can guide the user to align the camera with the user's oral cavity to ensure that the camera can clearly capture the image inside the user's oral cavity. Specifically, the target electronic device can guide the user to align the camera with the user's oral cavity through visual display or voice prompt, so as to capture the oral image inside the user's oral cavity.
[0058] In some embodiments, the target electronic device can control the camera to take pictures through the take picture button. Specifically, when the target electronic device detects that the user presses the take picture button, it controls the camera to perform image acquisition to obtain multiple frames of acquired images; the target electronic device can screen out the acquired image with the largest area of the user's oral image region from the multiple frames of acquired images through image recognition of the multiple frames of acquired images as the oral image. The target electronic device can also use a frame of acquired image selected by the user as the oral image.
[0059] Optionally, the target electronic device may be configured with a display device. During the process of the camera taking pictures of the user's oral cavity, the target electronic device can display the current shooting picture of the camera on the display device in real time, so that the user can observe the shooting picture of the camera through the display device, so that the target electronic device can take an oral image containing the teeth in the user's oral cavity through the camera, and avoid non-oral areas (such as cheeks, nose, etc.) from appearing in the oral image.
[0060] Optionally, the target electronic device may also be configured with an audio playback module. After the target electronic device is started, control the audio playback module to emit a first preset audio to remind the user that the current shooting picture of the camera is not aligned with the user's oral cavity and the target electronic device needs to be moved; in the case where the target electronic device detects that the shooting picture of the camera is aligned with the user's oral cavity, control the audio playback module to emit a second preset audio to remind the user that the current shooting picture of the camera is aligned with the user's oral cavity, stop moving the target electronic device or press the shooting button to control the camera to take a picture. The first preset audio is different from the second preset audio. For example, the first preset audio can be a soothing audio, and the second preset audio can be a rapid audio.
[0061] Exemplarily, the target electronic device can also detect the distance between the camera and the user's oral cavity. According to this distance, control the prompt tone broadcast module to emit a preset audio at a frequency corresponding to this distance. For example, the smaller the distance between the camera and the user's oral cavity, the faster the frequency at which the prompt tone broadcast module emits the preset audio, and the larger the distance between the camera and the user's oral cavity, the slower the frequency at which the prompt tone broadcast module emits the preset audio, so as to remind the user to bring the camera closer to the oral cavity. Thus, the camera can accurately take an oral image of the user's oral cavity and reduce the interference of non-oral areas.
[0062] Step 220, identify the oral image through the target oral cavity detection model to obtain an oral cavity detection result corresponding to the user's oral cavity.
[0063] The target oral cavity detection model is an artificial intelligence model used to identify and analyze the input oral image, so as to realize the oral cavity detection of the user's oral cavity and obtain the oral cavity detection result. The target electronic device can input the obtained oral image into the target oral cavity detection model, and through the processing of the target oral cavity detection model, the oral cavity detection result corresponding to the user's oral cavity can be obtained.
[0064] The oral cavity detection result refers to the result output by the target oral cavity detection model after analyzing and recognizing the input oral cavity image. Optionally, the oral cavity detection result may include an oral cavity area detection result and / or an oral cavity disease detection result. Among them, the oral cavity area detection result is used to indicate the position information corresponding to each oral cavity area in the user's oral cavity, and the oral cavity disease detection result is used to indicate the types of oral cavity diseases that can be recognized in the user's oral cavity and the relevant disease information. For example, the target oral cavity detection model can output information such as the positions of each tooth in the oral cavity image in the user's oral cavity, the possible disease types and probabilities of each tooth, and whether there are missing teeth.
[0065] In the embodiments of the present application, the target oral cavity detection model is obtained by performing transfer learning on a pre-trained oral cavity detection model through a first sample set; the first sample oral cavity images included in the first sample set are obtained by the target electronic device through a camera, and the first sample oral cavity images are labeled with sample oral cavity detection results.
[0066] In some embodiments, a pre-trained oral cavity detection model can be first pre-trained through a large number of second sample oral cavity images, and then a small number of labeled first sample oral cavity images are used to perform transfer learning on the pre-trained oral cavity detection model to convert the pre-trained oral cavity detection model into a target oral cavity detection model suitable for the oral cavity images collected by the target electronic device, thereby reducing the training time and cost of the target oral cavity detection model.
[0067] Transfer learning refers to a machine learning technique that allows a model to apply the knowledge or feature parameters learned from one task to another related but different task. Transfer learning can significantly shorten the training time of a new task (i.e., training a pre-trained oral cavity detection model through a first sample set to obtain a target oral cavity detection model), and reduce the dependence on a large amount of labeled data (first sample oral cavity images). It can improve the generalization ability of the oral cavity detection model.
[0068] Specifically, the parameters of the target oral cavity detection model can be fine-tuned on the first sample set starting from the pre-trained oral cavity detection model. For example, the parameters of the target oral cavity detection model can be , where represents the parameters of the pre-trained oral cavity detection model, represents the second sample set, represents the first sample set, and the number of images in the first sample set can be much smaller than the number of images in the second sample set. By using the pre-trained oral cavity detection model trained on as a starting point and using to fine-tune the parameters of the pre-trained oral cavity detection model, the parameters of the target oral cavity detection model are obtained.
[0069] It should be noted that the pre-trained oral detection model is obtained by pre-training the oral detection model to be trained using a second sample set through one or more sample electronic devices; the second sample oral images included in the second sample set are obtained by shooting with the cameras of one or more sample electronic devices, and the second sample oral images are labeled with actual oral detection results.
[0070] Optionally, the target oral detection model may not be stored on the target electronic device. Then, the target electronic device controls the camera to shoot the user's oral cavity to obtain an oral image, and sends the oral image to the electronic device storing the target oral detection model. The electronic device storing the target oral detection model identifies the oral image to obtain the oral detection result corresponding to the user's oral cavity, and sends the oral detection result corresponding to the oral image back to the target electronic device, so that the target electronic device can obtain the oral detection result corresponding to the user's oral cavity.
[0071] In some embodiments, the oral image is obtained by the target electronic device through the camera using target shooting parameters. Among them, the target shooting parameters include at least one of the target magnification of the lens, the target aperture, the target shutter speed, the target shooting angle, the target sensitivity, etc. The target electronic device can adapt to different shooting conditions and user requirements by setting various shooting parameters, improve the richness of the oral image, so that the target oral detection model can perform oral detection on oral images with various target shooting parameters, which helps the oral detection model to perform oral detection more accurately to obtain more accurate oral detection results.
[0072] Specifically, the target magnification refers to the magnification degree of the object in the image, which is used to ensure the clarity of the object in the oral image; the target aperture refers to the amount of light allowed to pass through the camera and the depth of field, which is used to balance the image brightness and clarity; the target shutter speed refers to the length of time the camera is open, which is used to control the exposure time to avoid overexposure; the target shooting angle refers to the position and direction of the camera relative to the object being photographed, which is used to make the obtained oral image contain more internal oral structures and details; the target sensitivity refers to the sensitivity of the camera to light, which is used to enhance the image brightness in low-light conditions. By controlling the target shooting parameters of the camera, the obtained oral image is clearer and contains more internal oral elements, which helps to detect the user's oral cavity by analyzing the oral image subsequently.
[0073] In some embodiments, the first sample oral images included in the first sample set are also obtained by the target electronic device through a camera using target shooting parameters. Specifically, the first sample oral images and the above-mentioned oral images may be taken by the target electronic device through the same camera based on the same shooting parameters. This can ensure that the target oral detection model obtained by transfer learning of the first sample oral images accurately identifies oral images, thereby improving the efficiency and accuracy of oral detection using the target oral detection model.
[0074] In the embodiments of the present application, the target electronic device can use a small number of labeled first sample oral images taken by the target electronic device to perform transfer training on a pre-trained oral detection model, which not only enables the pre-trained oral detection model to quickly adapt to the image acquisition characteristics of the target electronic device, but also reduces the dependence on the labeled sample data corresponding to the target electronic device, reduces the training cost of the target oral detection model, and improves the generalization ability and accuracy of the target oral detection model, thereby improving the efficiency and accuracy of oral detection using the target oral detection model.
[0075] As Figure 3A shown, in one embodiment, an oral detection method is provided, which can be applied to the above-mentioned electronic device. The method may include the following steps 302 to step 304.
[0076] Step 302, obtain an oral image corresponding to the user's oral cavity, and the oral image is obtained by the target electronic device through a camera.
[0077] For the relevant description of step 302, reference may be made to the relevant description of step 210 in the above embodiments, and details are not described herein again.
[0078] Step 304, perform oral region recognition on the oral image through the target oral detection model to obtain an oral region detection result corresponding to the user's oral cavity, and perform disease detection on the oral image to obtain an oral disease detection result corresponding to the user's oral cavity.
[0079] In some embodiments, the oral detection result may include an oral region detection result and an oral disease detection result. Then, the target electronic device can perform oral region recognition and disease detection on the oral image through the oral detection model respectively to obtain the oral region detection result and the oral disease detection result respectively.
[0080] The oral region detection result may include the position information corresponding to each oral region included in the user's oral cavity. Among them, the oral region may include but is not limited to the tooth region, the gum region, the tongue region, the oral mucosa region, etc. Specifically, the position information corresponding to the tooth region may include the specific positions of each tooth in the oral image in the user's oral cavity, etc.
[0081] In some embodiments, the oral region detection result may further include the appearance information of each region in the user's oral cavity. The appearance information refers to the phenomena that can be directly recognized through oral images, such as whether there is redness, swelling or bleeding in the gingival region, oral ulcers, etc.
[0082] Optionally, the target electronic device can also perform oral region recognition on the oral image through the target oral detection model to identify the position and / or number of each tooth included in the oral image in the user's oral cavity; or, identify the health condition of the user's gums (such as whether there is redness, swelling, bleeding or recession, etc.); or, identify the tongue region in the oral image and analyze the color, shape and texture of the tongue in the tongue region to determine whether there is any abnormality in the user's tongue; or, identify the oral mucosa region in the user's oral cavity and check the smoothness and color of the user's oral mucosa to determine whether there is any ulcer or abnormal hyperplasia in the user's oral cavity.
[0083] Taking the oral region detection result in the oral detection result as an example, the oral region detection result in the oral detection result corresponding to a certain user's oral cavity may include: tooth region (the specific position of each tooth in the user's oral cavity, without missing teeth), gingival region (healthy, without obvious redness, swelling or bleeding), tongue region (normal color, good shape, complete edge), oral mucosa region (smooth, without abnormal hyperplasia or ulcer). Through these position information, it can help oral doctors or users to more intuitively understand the internal situation of the oral cavity, can provide more comprehensive and detailed oral detection results for users, and provide data support for subsequent treatment or health maintenance.
[0084] In some embodiments, the oral region includes each tooth and each tooth surface included in each tooth. The position information corresponding to the oral region may include the position of each tooth in the user's oral cavity and the tooth surface type corresponding to each tooth surface included in each tooth. The tooth surface refers to each surface of the tooth. The tooth surface is in direct contact with other structures in the oral cavity (such as gums, adjacent teeth, tongue or food, etc.), and the tooth surface types of teeth in different positions are different. By refining the oral region to each tooth and its tooth surface, more comprehensive and detailed oral detection results can be provided for users.
[0085] Exemplarily, as Figure 3B shown, according to the functional characteristics of the teeth, the 4 teeth in the middle of the upper jaw and the 4 teeth in the middle of the lower jaw can be called incisors, the 4 teeth adjacent to the incisors can be called canines, and the remaining teeth can be called molars. Among them, both incisors and canines may have two tooth surface types, namely the outer side surface of the tooth close to the cheek and the inner side surface close to the tongue; molars may have three tooth surface types, that is, the outer side surface close to the cheek, the inner side surface close to the tongue, and the occlusal surface of the molar that is relatively flat for grinding food.
[0086] In some embodiments, the oral disease detection result may include the types of diseases present in the user's oral cavity and the probability of occurrence of each disease type. Among them, the disease type refers to the type of oral disease, which may include but is not limited to dental caries, gingivitis, periodontitis, complete tooth loss or partial tooth loss, oropharyngeal cancer, lip cancer, oral mucosal disease, salivary gland disease, etc. The probability of occurrence of each disease type may refer to the probability of the user suffering from each disease type, or may refer to the severity of each disease type.
[0087] In some embodiments, the target electronic device may, through the target oral detection model, analyze, based on the appearance information of each region in the user's oral cavity, the types of diseases present in the user's oral cavity and the probability of occurrence of each disease type. For example, since oropharyngeal cancer may be accompanied by symptoms such as oral ulcers and masses, if the phenomena such as oral ulcers and masses are detected in the user's oral cavity at the same time, it can be considered that the probability of the user suffering from oropharyngeal cancer is relatively high or the user's oropharyngeal cancer is relatively severe.
[0088] Taking the oral disease detection result in the oral detection result as an example, the oral disease detection result in the oral detection result corresponding to a certain user's oral cavity may include: dental caries (there are signs of dental caries in two teeth, the occurrence probability is 80%), periodontal disease (mild inflammation of the gums, the occurrence probability is 60%), oral ulcer (no obvious ulcer is found). This enables the dentist to quickly identify the types of oral diseases that the user may suffer from through the oral disease detection result, and timely understand the severity or the probability of suffering from these oral diseases, so as to make a diagnosis and treatment in a timely manner.
[0089] In some embodiments, the target oral detection model may include an oral region detection module and an oral disease detection module. Among them, the oral region detection module is used to perform oral region detection on the oral image, and the oral disease detection module is used to perform disease detection on the oral image.
[0090] In some embodiments, the electronic device storing the target oral detection model may, through the oral region detection module, perform oral region recognition on the oral image to obtain the oral region detection result corresponding to the user's oral cavity; and, through the oral disease detection module, perform disease detection on the oral image in parallel to obtain the oral disease detection result corresponding to the user's oral cavity. In this target oral detection model, by performing oral region recognition through the oral region detection module and performing oral disease detection through the oral disease detection module respectively, the target oral detection model can process the input oral image in parallel and more efficiently, improving the overall oral detection rate.
[0091] It should be noted that the oral region detection module and the oral disease detection module in the target oral detection model can be located in two parallel threads. That is, after the oral image is input into the target oral detection model, the oral region detection module and the oral disease detection module can simultaneously detect the oral image to reduce the time for the target oral detection model to perform oral detection, thereby improving the efficiency of the target electronic device in detecting the user's oral cavity.
[0092] In some embodiments, since oral diseases are often accompanied by changes in tooth surface color and tooth shape. For example, dental caries may cause brown or black spots on the tooth surface, and dental caries may present depressions or holes on the tooth surface. Gingivitis may cause the gums to appear red and swollen or dark red. Oral ulcers may present as round or oval damaged areas, etc. Therefore, the target electronic device can also capture the user's oral cavity through a camera to obtain a color oral image of the user's oral cavity, and through the oral disease detection module of the target oral detection model, extract the color features and shape features corresponding to the teeth in the color oral image, and determine the types of diseases present in the user's oral cavity and the corresponding probabilities based on the color features and shape features.
[0093] In some embodiments, before the target oral detection model outputs the oral detection result corresponding to the oral image, it can integrate the oral region detection result detected by the oral region detection module and the oral disease detection result detected by the oral disease detection module, so that the output oral detection result can intuitively indicate the position information of each tooth in the oral cavity, the possible diseases and the corresponding probabilities.
[0094] Specifically, the target oral detection model can integrate the oral detection result output by the target oral detection model into a matrix form. For example, the oral detection result may include the position information B of each detected tooth, the possible diseases C and the corresponding probabilities S. Then the oral detection result can be , where k is the number after the integration of the oral detection result.
[0095] In the embodiments of the present application, the oral detection model can perform oral detection on the oral image through a multi-task processing method, and can simultaneously output the position information of each oral region in the user's oral cavity, as well as the possible diseases in the user's oral cavity and the probabilities of each disease occurring, thereby improving the efficiency of oral detection.
[0096] In some embodiments, as Figure 4A shown, the step of performing oral region recognition on the oral image through the target oral detection model in step 304 to obtain the oral region detection result corresponding to the user's oral cavity may include steps 402 to 406.
[0097] Step 402: Identify each tooth included in the oral cavity image through the target oral cavity detection model to obtain the tooth image corresponding to each tooth.
[0098] The tooth image includes the image area of a single tooth in the oral cavity image. The target oral cavity detection model can first identify the tooth image corresponding to each tooth through the oral cavity image, facilitating subsequent tooth analysis, tooth surface recognition, and other operations.
[0099] In some embodiments, the target electronic device can perform edge detection on each tooth included in the oral cavity image through the target oral cavity detection model to determine the edge contour of each tooth; then, segment the oral cavity image according to the edge contour of each tooth to obtain the tooth image corresponding to each tooth respectively.
[0100] Step 404: Match the tooth image corresponding to each tooth with multiple preset tooth images, and determine the position of each tooth in the user's oral cavity according to the preset tooth image that matches the tooth image corresponding to each tooth; the preset tooth image is the image of teeth in different positions in the user's oral cavity.
[0101] The preset tooth image includes the standard image or typical image of a single tooth in different positions. These preset tooth images can be generated based on a standard tooth model or obtained in advance in the user's oral cavity; specifically, the preset tooth image can be sourced from an existing medical image library, the user's oral cavity scan data, or a virtual tooth image generated through 3D modeling.
[0102] Optionally, the preset tooth image can include multiple views of the same tooth at different shooting angles, and can also include views of multiple teeth at the same angle but different positions. For example, the front view, top view, and 45-degree oblique front view of the same incisor; the occlusal surface view (top view / underside view) of each molar.
[0103] In some embodiments, the target oral cavity detection model can first extract tooth features from the tooth image corresponding to any tooth. The tooth features can include, but are not limited to, the shape feature, size feature, color feature, texture feature, etc. of the tooth. Similarly, extract the corresponding preset tooth features from multiple preset tooth images respectively. Calculate the similarity or distance between the tooth features corresponding to this tooth and each preset tooth feature to evaluate the matching degree between the tooth image corresponding to this tooth and each preset tooth image respectively. Take the preset tooth image with the highest matching degree as the preset tooth image that matches the tooth image corresponding to this tooth, so as to obtain the position of this tooth in the user's oral cavity. Through the similarity between the tooth features corresponding to a single tooth and each preset tooth feature, the positions of each tooth in the user's oral cavity can be determined more accurately.
[0104] Since the number and arrangement of human teeth are relatively fixed. For example, adults usually have 32 teeth (including wisdom teeth) or 28 teeth (excluding wisdom teeth), the front part of the oral cavity is incisors, and adults usually have 8 incisors, including 4 upper incisors and 4 lower incisors, etc. In some embodiments, the target electronic device can name or number the teeth in the oral cavity according to the arrangement of the teeth, so as to facilitate determining the position of each tooth in the oral cavity image in the user's oral cavity.
[0105] Specifically, the target oral detection model can determine the name or number of each tooth in the oral cavity image based on the name or number corresponding to the preset tooth image, according to the preset tooth image matched with the tooth image corresponding to each tooth, so as to determine the position of each tooth in the user's oral cavity.
[0106] Exemplarily, as Figure 3B shown, in the order from the middle to both sides, combining the position and type of the teeth, the name or number of the incisors can be set as "the first upper central incisor 311", "the second upper central incisor 313", "the first upper lateral incisor 315", "the second upper lateral incisor 317", "the first lower central incisor 312", "the second lower central incisor 314", "the first lower lateral incisor 316", "the second lower lateral incisor 318"; the name or number of the canines can be set as "the left upper canine 321", "the left lower canine 322", "the right upper canine 323", "the right lower canine 324"; the name or number of the molars can be set as "the left upper molar 331", "the left lower molar 332", "the right upper molar 333", "the right lower molar 334"... and so on, to obtain the name or number of 32 teeth. The name or number of a kind of tooth exemplified in the embodiments of the present invention is for the convenience of understanding the description of the technical solution, and does not limit the technical solution provided by the embodiments of the present invention.
[0107] In some embodiments, if there is no tooth loss in the user's oral cavity, the first position of the first tooth in the user's oral cavity and the second position of the second tooth in the user's oral cavity can be identified first. The first tooth is any tooth in the upper jaw, and the second tooth is any tooth in the lower jaw; according to the positional relationship between other teeth and the first tooth, or the positional relationship with the second tooth in the oral cavity image, the positions of other teeth in the user's oral cavity can be determined, which can reduce the calculation amount and more quickly determine the position of each tooth in the user's oral cavity, and improve the rate of the target oral detection model for identifying the oral cavity area in the oral cavity image.
[0108] Step 406, perform tooth surface recognition on each tooth image through the target oral detection model to determine the tooth surface type corresponding to each tooth surface included in each tooth.
[0109] According to the shape of the teeth, both incisors and canines can include two tooth surface types, that is, the tooth surface types corresponding to incisors and canines can be the outer side close to the cheek and the inner side close to the tongue; molars can include three tooth surface types, that is, the tooth surface types corresponding to molars can be the outer side close to the cheek, the inner side close to the tongue, and the occlusal surface that is relatively flat for grinding food.
[0110] In some embodiments, the target electronic device can determine the tooth type (incisor or canine or molar) corresponding to each tooth according to the position of each tooth in the user's oral cavity through the target oral detection model; and then determine the tooth surface types corresponding to the respective tooth surfaces included in each tooth type according to the tooth surface types included in various tooth types.
[0111] Specifically, after identifying the target tooth type of the third tooth, the target oral detection model can perform tooth surface recognition on the tooth image of the third tooth to obtain at least one tooth surface, where the third tooth is any tooth in the user's oral cavity; extract the tooth surface features of each tooth surface to obtain the tooth surface features of each tooth surface; match the tooth surface features with a plurality of preset surface features included in the target tooth type to determine the tooth surface type corresponding to the preset surface feature that matches the tooth surface features of each tooth surface, which is the tooth surface type corresponding to each tooth surface. The preset surface features are the features corresponding to the standard or typical tooth surface types.
[0112] For example, if it is recognized that a tooth surface of a certain incisor is relatively smooth and close to the tongue, then this tooth surface is the inner side. If it is recognized that a tooth surface of a certain canine is sharp and close to the cheek, then this tooth surface is the outer side. If it is recognized that a tooth surface of a certain molar has a complex shape and multiple cusps, then this tooth surface is the occlusal surface.
[0113] In the embodiments of the present application, the target oral detection model can obtain the tooth image corresponding to each tooth according to the oral image, and determine the position of each tooth in the user's oral cavity by matching the tooth image corresponding to each tooth with a plurality of preset tooth images; and determine the tooth surface type of each tooth in the oral image by performing tooth surface recognition on the tooth image, so as to obtain a more detailed and accurate oral detection result.
[0114] In some embodiments, the target oral detection model can also annotate the position information corresponding to each oral area in the oral image, and annotate the diseases existing in the user's oral cavity and the probability of occurrence of each disease in the oral image, so that the target electronic device can obtain an oral image annotated with the position information corresponding to each oral area, the existing diseases and the probability of occurrence of each disease through the target oral detection model, so as to be able to intuitively inform the user of the oral detection result and help the user better manage oral health.
[0115] Optionally, the position information corresponding to the oral region may include the position of each tooth in the user's oral cavity, and the type of tooth surface corresponding to each tooth surface included in each tooth.
[0116] For example, Figure 4B is an image of the mandibular teeth in the oral image. As Figure 4B shown, the first oral region 421 in the oral image can be labeled as "outer surface of the right mandibular canine, with signs of dental caries, occurrence probability is 80%"; the second oral region 422 can be labeled as "occlusal surface of the left mandibular wisdom tooth, without signs of dental caries"; the third oral region 423 can be labeled as "occlusal surface of the right mandibular molar (adjacent to the canine), with signs of pulp disease, occurrence probability is 60%". By labeling the position information, the existing diseases, and the corresponding probabilities of each disease on the oral image, the oral detection results can be intuitively informed to the user.
[0117] It should be noted that, Figure 4B is a part of the oral image, Figure 4B the labels in are only labels for some oral regions or some diseases, and other parts of the oral image are not exemplified here.
[0118] As Figure 5 shown, in one embodiment, the training process of the pre-trained oral detection model may include the following steps 502 to 506.
[0119] It should be noted that the pre-trained oral detection model is pre-trained according to the second sample set; the second sample oral images included in the second sample set are obtained by photographing with the cameras of one or more sample electronic devices, and the second sample oral images are labeled with actual oral detection results.
[0120] Step 502, input the second sample oral image into the oral detection model to be trained.
[0121] In some embodiments, the target oral detection model may construct the oral detection model to be trained based on the model parameters after training on the existing medical image library. The existing medical image library may include public data sets of medical schools, relevant materials of oral exhibitions, open source projects, professional medical images, etc.
[0122] It should be noted that the number of the second sample oral images included in the second sample set is greater than the number of the first sample oral images included in the first sample set. Through the pre-training of a large number of second sample oral images, the pre-trained oral detection model has learned a wide variety of oral detection-related features, thus having a strong oral detection ability. When performing transfer learning, a target oral detection model suitable for the target electronic device with higher accuracy can be obtained through a small number of first sample oral images.
[0123] Optionally, the second sample oral image may be an image containing only a single tooth, which can more detailedly display the image features of the user's oral cavity collected by the sample electronic device through the camera, so as to more accurately train the oral detection model to be trained.
[0124] Optionally, the training process of the pre-trained oral detection model can be performed on the target electronic device or on the sample electronic device. The sample electronic device for training the pre-trained oral detection model and the sample electronic device equipped with a camera (for capturing the second sample oral image) can be the same electronic device or different electronic devices, and there can be multiple sample electronic devices equipped with a camera (for capturing the second sample oral image).
[0125] Taking the example where the sample electronic device and the target electronic device are the same electronic device, the electronic device captures the first sample oral image using the target shooting parameters, which need to be different from the sample shooting parameters used for capturing the second sample oral image.
[0126] Optionally, the electronic device storing the oral detection image to be trained, the electronic device storing the pre-trained oral detection image, and the electronic device storing the target oral detection image can all be different electronic devices.
[0127] Step 504, extract features from the current second sample oral image through the oral detection model to be trained to obtain image features, and determine the predicted oral detection result corresponding to the current second sample oral image according to the image features.
[0128] The image features can reflect the state or attribute features of each tooth in the second sample oral image, and based on these image features, the predicted oral detection result can be determined more efficiently and accurately. The image features may include but are not limited to shape features (such as the contour, size, aspect ratio of the tooth), color features, texture features (the texture on the tooth surface), edge features, etc.
[0129] In some embodiments, the oral detection model to be trained may also perform structural reparameterization on the image features; and determine the predicted oral detection result corresponding to the current second sample oral image according to the image features after structural reparameterization.
[0130] Structural reparameterization can rearrange or combine some parameters in the original features to obtain a new parameter representation method, so that the expression ability of the features is stronger, and at the same time, the complexity of the features can be reduced. Therefore, it can provide higher-quality features in the model training stage to improve the calculation efficiency and parameter utilization rate, so that the finally obtained model can be deployed on more devices.
[0131] In some embodiments, the oral cavity detection model to be trained can be based on the RepViT (Reparameterized Vision Transformer) network to perform structural reparameterization on the feature map, where the feature map is obtained by extracting features from the oral cavity image.
[0132] Specifically, in the RepViT network, the feature map sequentially passes through the Token Mixture layer, the Channel Mixture layer, the Reparameterization layer, the Multi-Head Self-Attention layer, and the Feed-Forward Network layer to complete the structural reparameterization of the features.
[0133] Exemplarily, the oral cavity detection model to be trained can input the feature map (where H is the height of the feature map, W is the width of the feature map, and C is the number of channels) into the RepViT network and use 1x1 convolution for information interaction between channels; then use 3x3 depthwise separable convolution to fuse the spatial information of the features after information interaction ; remove the computational and storage overhead of skip connections through reparameterization technology for the fused features , and then use the attention algorithm to process the reparameterized features so that it can obtain global information and get the feature ; finally, perform feedback through the feed-forward network in a loop to obtain the image features after structural reparameterization . Specifically, the image features can be obtained through formulas (1) to (5).
[0134] (1); (2); (3); (4); (5); where SE is the Squeeze-and-Excitation operation for enhancing feature representation; MSA represents the multi-head self-attention operation; FFN represents the feed-forward neural network; are n learnable parameters in the network.
[0135] In this application, the oral cavity detection model to be trained improves the computational efficiency and parameter utilization rate of the pre-trained oral cavity detection model by performing structural re-parameterization on the image features of the second sample oral cavity images, enabling the pre-trained oral cavity detection model to be deployed on more electronic devices and expanding the applicable range of the pre-trained oral cavity detection model.
[0136] Step 506: According to the actual oral cavity detection result corresponding to the current second sample oral cavity image and the corresponding predicted training oral cavity detection result, adjust the parameters of the oral cavity detection model to be trained until the training completion condition is met, obtaining the pre-trained oral cavity detection model.
[0137] In some embodiments, the oral cavity detection model to be trained can determine a loss function according to the actual oral cavity detection result corresponding to the current second sample oral cavity image and the corresponding predicted training oral cavity detection result; and update the parameters of the oral cavity detection model to be trained according to the gradient descent direction of the loss function to obtain model parameters that meet the model convergence condition, and construct a pre-trained oral cavity detection model according to the model parameters.
[0138] Optionally, the model parameters that meet the model convergence condition may include that when the target loss value no longer decreases during two or more consecutive model parameter updates, such model parameters are model parameters that meet the model convergence condition.
[0139] In the embodiments of this application, the pre-trained oral cavity detection model is pre-trained based on a second sample set with a larger number of images, and the second sample set includes second sample oral cavity images obtained by shooting with the cameras of one or more sample electronic devices and marked with actual oral cavity detection results, so that in the pre-training stage, the pre-trained oral cavity detection model has learned a wide variety of oral cavity detection-related features, thus having a strong oral cavity detection ability; using this pre-trained oral cavity detection model as a basis, the target electronic device can perform transfer learning on the pre-trained oral cavity detection model based on a small number of first sample oral cavity images, and can quickly adjust and obtain the target oral cavity detection model to adapt to the target electronic device, thereby reducing the dependence on the labeled sample data corresponding to the target electronic device, reducing the training cost of the target oral cavity detection model, and improving the generalization ability of the target oral cavity detection model.
[0140] As Figure 6 shown, in one embodiment, the training process of the pre-trained oral cavity detection model may further include the following steps 602 to 604.
[0141] Step 602: Perform image enhancement on the second sample oral cavity images included in the second sample set to obtain an enhanced second sample set.
[0142] During the process of the camera capturing oral images, due to possible influences during the shooting conditions, the performance of the camera device, or the image transmission process, problems such as blurred oral images, excessive noise, or low contrast may occur. Therefore, during the training process of the pre-trained oral detection model, these problems can be reduced through image enhancement. Moreover, through image enhancement, the image features in the oral images can be highlighted, making them easier to be learned by the pre-trained oral detection model.
[0143] Optionally, the pre-trained oral detection model is pre-trained based on the enhanced second sample set.
[0144] In some embodiments, the image enhancement methods that can be adopted include, but are not limited to: performing color enhancement on one or more second sample oral images in the second sample set; and / or, performing geometric transformation on one or more second sample oral images in the second sample set; and / or, performing mosaic enhancement on one or more second sample oral images in the second sample set.
[0145] Exemplarily, the electronic device for training the pre-trained oral detection model can complete the color enhancement of the second sample oral image by adjusting the brightness, contrast, and saturation of the second sample oral image. Specifically, the electronic device for training the pre-trained oral detection model can obtain the color-enhanced second sample oral image through formula (6).
[0146] (6); Where, represents the original second sample oral image; I represents the enhanced second sample oral image; respectively represent the parameters of brightness, contrast, and saturation.
[0147] Specifically, the geometric transformations that can be adopted include, but are not limited to, perspective transformation, rotation transformation, scaling transformation, and translation transformation, etc.
[0148] Exemplarily, the electronic device for training the pre-trained oral detection model can achieve perspective transformation by adjusting the perspective of the image to complete the geometric transformation of the second sample oral image, thereby simulating the effects of capturing oral images from different angles. Specifically, the electronic device for training the pre-trained oral detection model can obtain the second sample oral image after perspective transformation through formula (7).
[0149] (7); Where, represents the pixel coordinates in the original second sample oral image; represents the pixel coordinates of the enhanced second sample oral image; represents the perspective transformation parameter.
[0150] Exemplarily, the electronic device for training the pre-trained oral detection model can complete the geometric transformation of the second sample oral image by rotating the image. Specifically, the rotated second sample oral image can be obtained through formula (8).
[0151] (8); wherein, represents the pixel coordinates in the original second sample oral image; represents the pixel coordinates of the enhanced second sample oral image; represents the rotation angle.
[0152] Exemplarily, the electronic device for training the pre-trained oral detection model can complete the geometric transformation of the second sample oral image by changing the scaling ratio of the image. Specifically, the scaled second sample oral image can be obtained through formula (9).
[0153] (9); wherein, represents the rotation angle.
[0154] Exemplarily, the electronic device for training the pre-trained oral detection model can complete the geometric transformation of the second sample oral image by translating the image. Specifically, the translated second sample oral image can be obtained through formula (10).
[0155] (10); wherein, respectively represent the translation parameters in the horizontal axis direction and the vertical axis direction.
[0156] In some embodiments, the electronic device can generate a new second sample oral image by randomly splicing different regions of the second sample oral image, thereby improving the generalization ability of the oral detection model. Specifically, the mosaic-enhanced second sample oral image can be obtained through formula (11). .
[0157] (11); wherein, represents different regions in the original second sample oral image; is the mosaic enhancement function.
[0158] In some embodiments, the electronic device for training the pre-trained oral cavity detection model may also perform image enhancement on the second sample oral cavity images included in the second sample set to obtain enhanced second sample oral cavity images; obtain the actual oral cavity detection results corresponding to each enhanced second sample oral cavity image; and add the enhanced second sample oral cavity images marked with the actual oral cavity detection results to the second sample set to obtain an enhanced second sample set. By re-labeling the enhanced second sample oral cavity images, the actual oral cavity detection results in the enhanced second sample oral cavity images are made more accurate, so that during the training process of the oral cavity detection model to be trained, the parameter optimization of the model is more precise, thereby improving the accuracy of the output of the pre-trained oral cavity detection model.
[0159] Step 604: Train the oral cavity detection model to be trained according to the enhanced second sample set.
[0160] For the relevant description of training the oral cavity detection model to be trained in step 604, reference can be made to steps 502 to 506 of the above embodiments, which will not be elaborated here.
[0161] In the embodiments of the present application, by performing image enhancement processing on the second sample oral cavity images, richer and more diverse training data is provided for the oral cavity detection image to be trained, thereby improving the adaptability of the pre-trained oral cavity detection model to different image conditions; and by performing image enhancement on the second sample oral cavity images through multiple enhancement methods, more diverse training data can be obtained, thereby improving the generalization ability of the pre-trained oral cavity detection model.
[0162] In some specific embodiments, assume that the training process of the pre-trained oral cavity detection model is on the sample electronic device a, and the sample electronic device a stores the oral cavity detection image to be trained; the target electronic device b is equipped with both a camera and stores the target oral cavity detection model.
[0163] The user can use the camera of the sample electronic device a or other sample electronic devices to take pictures of the oral cavity, so that the sample electronic device a obtains a large number of second sample oral cavity images, performs image enhancement on the second sample oral cavity images (which may include color enhancement, geometric transformation, mosaic enhancement, etc.), and inputs the enhanced second sample oral cavity images into the oral cavity detection model to be trained.
[0164] The sample electronic device a is based on the RepViT algorithm. The trained oral detection model extracts features from each second sample oral image to obtain corresponding image features. The sample electronic device a performs structural reparameterization on each image feature; and determines the predicted oral detection results corresponding to each second sample oral image according to the structurally reparameterized image features; and adjusts the parameters of the trained oral detection model according to the actual oral detection results and the corresponding predicted training oral detection results of the current second sample oral image until the training completion condition is met, obtaining a pre-trained oral detection model.
[0165] The target electronic device b captures the user's oral cavity through a camera to obtain a small number of first sample oral images, and fine-tunes the model parameters of the pre-trained oral detection model through the small number of first sample oral images to obtain a target oral detection model. The target electronic device b obtains the oral image captured by the user through the camera, and then uses the target oral detection model to identify the oral image to obtain the oral detection result corresponding to the user's oral cavity. The target electronic device b outputs the oral detection result corresponding to the oral image so that the user or dentist can perform oral diagnosis and treatment according to the oral detection result.
[0166] In the embodiment of the present application, it is possible to use a small number of first sample oral images captured by the labeled target electronic device to perform transfer training on the pre-trained oral detection model, which not only enables the pre-trained oral detection model to quickly adapt to the image acquisition characteristics of the target electronic device, but also reduces the dependence on the labeled sample data corresponding to the target electronic device, reduces the training cost of the target oral detection model, and improves the generalization ability and accuracy of the target oral detection model, thereby improving the efficiency and accuracy of oral detection using the target oral detection model.
[0167] As Figure 7 shown, in one embodiment, an oral detection device 700 is provided, which can be applied to the above-mentioned electronic device. The oral detection device 700 may include an image acquisition module 710 and an oral detection module 720.
[0168] The image acquisition module 710 is configured to obtain an oral image corresponding to the user's oral cavity, and the oral image is captured by the target electronic device through a camera.
[0169] The oral detection module 720 is configured to identify the oral image through the target oral detection model to obtain the oral detection result corresponding to the user's oral cavity; the target oral detection model is obtained by performing transfer learning on the pre-trained oral detection model through a first sample set; the first sample oral images included in the first sample set are captured by the target electronic device through a camera, and the first sample oral images are labeled with sample oral detection results.
[0170] Optionally, the oral cavity image is obtained by the target electronic device through a camera using target shooting parameters; the first sample oral cavity images included in the first sample set are obtained by the target electronic device through a camera using target shooting parameters. The target shooting parameters include at least one of the target magnification of the lens, the target aperture, the target shutter speed, the target shooting angle, and the target sensitivity.
[0171] Optionally, the oral cavity detection result includes an oral cavity area detection result and an oral cavity disease detection result; the oral cavity area detection result includes the position information corresponding to each oral cavity area included in the user's oral cavity; the oral cavity disease detection result includes the types of diseases existing in the user's oral cavity and the probabilities of occurrence of each disease type.
[0172] In some embodiments, the oral cavity detection module 720 is further configured to perform oral cavity area recognition on the oral cavity image through a target oral cavity detection model to obtain an oral cavity area detection result corresponding to the user's oral cavity, and perform disease detection on the oral cavity image to obtain an oral cavity disease detection result corresponding to the user's oral cavity.
[0173] Optionally, the oral cavity area includes each tooth and each tooth surface included in each tooth; the position information corresponding to the oral cavity area includes the position of each tooth in the user's oral cavity and the tooth surface type corresponding to each tooth surface included in each tooth.
[0174] In some embodiments, the oral cavity detection module 720 is further configured to identify each tooth included in the oral cavity image through a target oral cavity detection model to obtain a tooth image corresponding to each tooth; match the tooth image corresponding to each tooth with a plurality of preset tooth images, and determine the position of each tooth in the user's oral cavity according to the preset tooth image matched with the tooth image corresponding to each tooth; the preset tooth images are images of teeth in different positions in the user's oral cavity; perform tooth surface recognition on each tooth image through the target oral cavity detection model to determine the tooth surface type corresponding to each tooth surface included in each tooth.
[0175] Optionally, the target oral cavity detection model includes an oral cavity area detection module and an oral cavity disease detection module.
[0176] In some embodiments, the oral cavity detection module 720 is further configured to perform oral cavity area recognition on the oral cavity image through the oral cavity area detection module to obtain an oral cavity area detection result corresponding to the user's oral cavity; perform disease detection on the oral cavity image in parallel through the oral cavity disease detection module to obtain an oral cavity disease detection result corresponding to the user's oral cavity.
[0177] In some embodiments, the oral cavity detection module 720 is further configured to label the position information corresponding to each oral cavity region in the oral cavity image, and label the diseases existing in the user's oral cavity and the occurrence probabilities of each disease in the oral cavity image.
[0178] Optionally, the pre-trained oral cavity detection model is pre-trained according to a second sample set; the second sample oral cavity images included in the second sample set are obtained by photographing with the cameras of one or more sample electronic devices, and the second sample oral cavity images are labeled with actual oral cavity detection results; the number of second sample oral cavity images included in the second sample set is greater than the number of first sample oral cavity images included in the first sample set.
[0179] In some embodiments, the oral cavity detection device 700 further includes a pre-training module.
[0180] The pre-training module is configured to input the second sample oral cavity images into the oral cavity detection model to be trained; extract image features from the current second sample oral cavity image through the oral cavity detection model to be trained, obtain image features, and determine the predicted oral cavity detection results corresponding to the current second sample oral cavity image according to the image features; adjust the parameters of the oral cavity detection model to be trained according to the actual oral cavity detection results corresponding to the current second sample oral cavity image and the corresponding predicted training oral cavity detection results until the training completion condition is met, and obtain the pre-trained oral cavity detection model.
[0181] Optionally, the pre-training module is further configured to perform structural re-parameterization on the image features; and determine the predicted oral cavity detection results corresponding to the current second sample oral cavity image according to the image features after structural re-parameterization.
[0182] In some embodiments, the pre-trained oral cavity detection model is pre-trained according to the enhanced second sample set. The pre-training module further includes an image enhancement unit.
[0183] The image enhancement unit is configured to perform image enhancement on the second sample oral cavity images included in the second sample set to obtain the enhanced second sample set.
[0184] Optionally, the image enhancement unit is further configured to perform color enhancement on one or more second sample oral cavity images in the second sample set; and / or perform geometric transformation on one or more second sample oral cavity images in the second sample set; and / or perform mosaic enhancement on one or more second sample oral cavity images in the second sample set.
[0185] Optionally, the image enhancement unit is further configured to enhance the second sample oral images included in the second sample set to obtain enhanced second sample oral images; obtain actual oral detection results corresponding to each of the enhanced second sample oral images; and add the enhanced second sample oral images labeled with the actual oral detection results to the second sample set to obtain an enhanced second sample set.
[0186] In the embodiments of the present application, the oral detection device can use a small number of labeled first sample oral images captured by the target electronic device to perform transfer training on the pre-trained oral detection model, which not only enables the pre-trained oral detection model to quickly adapt to the image acquisition characteristics of the target electronic device, but also reduces the dependence on the labeled sample data corresponding to the target electronic device, reduces the training cost of the target oral detection model, and improves the generalization ability and accuracy of the target oral detection model, thereby improving the efficiency and accuracy of oral detection using the target oral detection model.
[0187] Figure 8 is a structural block diagram of an electronic device in an embodiment. As Figure 8 shown, the electronic device 110 may include one or more of the following components: a processor 810 and a memory 820 coupled to the processor 810. The memory 820 may store one or more computer programs, and the one or more computer programs may be configured to be executed by the one or more processors 810 to implement the methods described in the above embodiments.
[0188] The processor 810 may include one or more processing cores. The processor 810 connects various parts within the entire electronic device 110 using various interfaces and lines, and executes various functions of the electronic device 110 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 820, and by calling data stored in the memory 820.
[0189] The memory 820 may include a random access memory (RAM) and may also include a read-only memory (ROM). The memory 820 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 820 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the above method embodiments, etc. The data storage area may also store data created during the use of the electronic device 110.
[0190] Understandably, the electronic device 110 may include more or fewer structural elements than those shown in the above structural block diagram. For example, it may include a power supply, input buttons, a camera, a speaker, a screen, an RF (Radio Frequency) circuit, a Wi-Fi (Wireless Fidelity) module, a Bluetooth module, sensors, etc., which will not be further limited herein.
[0191] An embodiment of the present application discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the methods described in the above embodiments are implemented.
[0192] An embodiment of the present application discloses a computer program product including a computer program, and when the computer program is executed by a processor, the methods described in the above embodiments are implemented.
[0193] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium, and when the program is executed, it may include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), etc.
[0194] Any reference to a memory, storage, database, or other medium as used herein may include non-volatile and / or volatile memory. Suitable non-volatile memory may include ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which is used as an external cache memory.
[0195] The above has introduced in detail an oral detection method, device, and electronic device disclosed in the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An oral detection method, characterized in that, The method includes: Obtaining an oral image corresponding to the user's oral cavity, where the oral image is captured by a camera of a target electronic device; Identifying the oral image through a target oral detection model to obtain an oral detection result corresponding to the user's oral cavity; the target oral detection model is obtained by performing transfer learning on a pre-trained oral detection model through a first sample set; the first sample oral images included in the first sample set are captured by the target electronic device through the camera, and the first sample oral images are labeled with sample oral detection results.
2. The method according to claim 1, wherein The oral image is captured by the target electronic device through the camera using target shooting parameters; The first sample oral images included in the first sample set are captured by the target electronic device through the camera using the target shooting parameters.
3. The method according to claim 2, wherein The target shooting parameters include at least one of a target magnification ratio of the lens, a target aperture, a target shutter speed, a target shooting angle, and a target sensitivity.
4. The method according to claim 1, wherein The oral detection result includes an oral region detection result and an oral disease detection result; the oral region detection result includes position information corresponding to each oral region included in the user's oral cavity; the oral disease detection result includes the types of diseases present in the user's oral cavity and the probabilities of occurrence of each disease type; The identifying the oral image through the target oral detection model to obtain an oral detection result corresponding to the user's oral cavity includes: Identifying the oral region of the oral image through the target oral detection model to obtain an oral region detection result corresponding to the user's oral cavity, and performing disease detection on the oral image to obtain an oral disease detection result corresponding to the user's oral cavity.
5. The method according to claim 4, wherein The oral region includes each tooth and each tooth surface included in each tooth; The position information corresponding to the oral region includes the position of each tooth in the user's oral cavity and the tooth surface type corresponding to each tooth surface included in each tooth.
6. The method according to claim 5, wherein The identifying the oral region of the oral image through the target oral detection model to obtain an oral region detection result corresponding to the user's oral cavity includes: Identifying each tooth included in the oral image through the target oral detection model to obtain a tooth image corresponding to each tooth; Matching the tooth image corresponding to each tooth with a plurality of preset tooth images, and determining the position of each tooth in the user's oral cavity according to the preset tooth image matched with the tooth image corresponding to each tooth; the preset tooth images are images of teeth in different positions in the user's oral cavity; Identifying the tooth surface of each of the tooth images through the target oral detection model to determine the tooth surface type corresponding to each tooth surface included in each tooth.
7. The method according to claim 4, wherein The target oral detection model includes an oral region detection module and an oral disease detection module; the process of using the target oral detection model to perform oral region recognition on the oral image to obtain the oral region detection result corresponding to the user's oral cavity, and perform disease detection on the oral image to obtain the oral disease detection result corresponding to the user's oral cavity includes: Using the oral region detection module to perform oral region recognition on the oral image to obtain the oral region detection result corresponding to the user's oral cavity; Using the oral disease detection module to concurrently perform disease detection on the oral image to obtain the oral disease detection result corresponding to the user's oral cavity.
8. The method according to any one of claims 4 to 7, characterized in that The process of using the target oral detection model to perform recognition on the oral image to obtain the oral detection result corresponding to the user's oral cavity includes: Marking the position information corresponding to each oral region in the oral image, and marking the disease types existing in the user's oral cavity and the occurrence probability of each disease type in the oral image.
9. The method according to claim 1, wherein The pre-trained oral detection model is pre-trained according to a second sample set; the second sample oral images included in the second sample set are obtained by photographing with the cameras of one or more sample electronic devices, and the second sample oral images are marked with actual oral detection results; The number of second sample oral images included in the second sample set is greater than the number of first sample oral images included in the first sample set.
10. The method according to claim 9, wherein The training process of the pre-trained oral detection model includes: Inputting the second sample oral image into the oral detection model to be trained; Performing feature extraction on the current second sample oral image through the oral detection model to be trained to obtain image features, and determining the predicted oral detection result corresponding to the current second sample oral image according to the image features; Adjusting the parameters of the oral detection model to be trained according to the actual oral detection result corresponding to the current second sample oral image and the corresponding predicted training oral detection result until the training completion condition is met, to obtain the pre-trained oral detection model.
11. The method according to claim 10, characterized in that, The process of determining the predicted oral detection result corresponding to the current second sample oral image according to the image features includes: Performing structural re-parameterization on the image features; Determining the predicted oral detection result corresponding to the current second sample oral image according to the image features after structural re-parameterization.
12. The method according to claim 9, wherein The pre-trained oral detection model is pre-trained according to the enhanced second sample set; The training process of the pre-trained oral detection model further includes: Performing image enhancement on the second sample oral images included in the second sample set to obtain the enhanced second sample set.
13. The method according to claim 12, characterized in that, The process of performing image enhancement on the second sample oral images included in the second sample set includes: Performing color enhancement on one or more second sample oral images in the second sample set; and / or, Performing geometric transformation on one or more second sample oral images in the second sample set; and / or, Performing mosaic enhancement on one or more second sample oral images in the second sample set.
14. The method according to claim 12, wherein Performing image enhancement on the second sample oral images included in the second sample set to obtain an enhanced second sample set, including: Performing image enhancement on the second sample oral images included in the second sample set to obtain enhanced second sample oral images; Obtaining the actual oral detection results corresponding to each of the enhanced second sample oral images; Adding the enhanced second sample oral images labeled with the actual oral detection results to the second sample set to obtain an enhanced second sample set.
15. An oral detection device, characterized in that, The apparatus includes: An image acquisition module, configured to acquire an oral image corresponding to a user's oral cavity, where the oral image is obtained by a target electronic device through a camera; An oral detection module, configured to identify the oral image through a target oral detection model to obtain an oral detection result corresponding to the user's oral cavity; the target oral detection model is obtained by performing transfer learning on a pre-trained oral detection model through a first sample set; the first sample oral images included in the first sample set are obtained by the target electronic device through the camera, and the first sample oral images are labeled with sample oral detection results.
16. An electronic device, characterized in that, Including a memory and a processor, where a computer program is stored in the memory, and when the computer program is executed by the processor, the processor implements the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Deep learning method for identifying oral squamous cell carcinoma based on visual features
CN111369501A
Pressed film definition detection method, device and apparatus of shell-shaped film and medium
CN113781472A
Oral health monitoring method and device
CN114972297A
Tooth type identification method and device
CN115439409A
Oral cavity examination method, apparatus, system, and related device
WO2024108803A1