An automatic evaluation method and device for deformed central tips based on convolutional neural networks
The multimodal automatic evaluation method for deformed central cusps using convolutional neural networks, which utilizes pseudo-paired samples and feature extractor initialization, solves the problems of low efficiency and poor consistency in deformed central cusp evaluation, and achieves efficient and accurate automated evaluation.
Patent Information
- Application Number
- CN202311166625.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-11
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-09-11
AI Technical Summary
In existing technologies, the assessment of abnormal central cusps relies on manual observation, which is inefficient and lacks objectivity and consistency. The scarcity of medical resources and the shortage of professional personnel make the need for automated assessment urgent.
A multimodal automatic evaluation method for deformed central cusp based on convolutional neural networks is adopted. Training samples are constructed by using pseudo-paired samples. Combined with pseudo-paired sample generation and feature extractor initialization, a deformed central cusp evaluation model is constructed to achieve automated evaluation.
It improved the accuracy and robustness of the assessment of malformed central cusp, reduced the workload of sample collection, enhanced the model's feature representation and generalization ability, and reduced the workload of medical personnel.
Smart Images

Figure CN117274757B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a method and apparatus for processing multimodal oral medical images based on deep learning. Background Technology
[0002] In the current field of oral healthcare, accurate and rapid screening for malformed central cusps has become a crucial task. Malformed central cusps, a common dental condition, require assessment based on two different modalities of imagery: panoramic X-rays and intraoral photographs. These two types of images differ in characteristics. Generally, panoramic X-rays provide a comprehensive view of the dental arch, while intraoral photographs offer detailed local information, both of which are essential for the identification and assessment of malformed central cusps. Currently, the assessment process typically relies on experienced medical personnel to manually observe these two modalities of images to determine the presence of malformed central cusps. However, this method has significant drawbacks: not only is the assessment inefficient, but the results are often influenced by the subjective judgment of the medical personnel, lacking objectivity and consistency. Furthermore, the substantial image processing workload places a significant burden on medical personnel.
[0003] Given the scarcity of medical resources and the shortage of professional medical personnel in my country, relying on manual assessment of malformed central cusps is becoming increasingly difficult. Therefore, there is a growing demand for methods and devices capable of automating the processing and accurate assessment of multimodal images of malformed central cusps. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide an automatic evaluation method and apparatus for multimodal deformity central cusp based on convolutional neural networks. The automatic evaluation method for multimodal deformity central cusp based on convolutional neural networks includes:
[0005] Obtain X-ray images and intraoral photographs of the target to be detected;
[0006] Identify N X-ray regions in an X-ray film, denoted as X1, X2, ..., X... N And the N intraoral imaging regions corresponding to the N X-ray areas in the intraoral photograph, denoted as P1, P2, ..., P... N The X-ray area and the intraoral imaging area correspond to N positions of the teeth on the left and right sides of the upper and lower jaws, starting from the midline of the teeth. The X-ray area and the intraoral imaging area corresponding to each tooth position form a paired area.
[0007] A model for assessing abnormal central cusp is obtained, which is trained using pseudo-paired samples as training samples. The pseudo-paired samples include positive pseudo-paired samples and negative pseudo-paired samples.
[0008] Each paired region of the target to be detected is input into the abnormal central cusp evaluation model to determine whether an abnormal central cusp exists in the paired regions of the target to be detected.
[0009] The construction of training samples includes the following steps:
[0010] Determine the positive training sample set and the negative training sample set. The positive training sample set refers to the set of samples with positive labels corresponding to the X-ray area and the set of samples with positive labels corresponding to the intraoral area. The negative training sample set refers to the set of samples with negative labels corresponding to the X-ray area and the set of samples with negative labels corresponding to the intraoral area. A positive label indicates the presence of a deformed central cusp, and a negative label indicates the absence of a deformed central cusp.
[0011] Positive pseudo-paired samples and negative pseudo-paired samples are generated based on the positive training sample set and the negative training sample set, respectively. A positive pseudo-paired sample refers to a pair of samples consisting of an X-ray area and an intraoral area randomly selected from the positive training sample set, and a negative pseudo-paired sample refers to a pair of samples consisting of an X-ray area and an intraoral area randomly selected from the negative training sample set.
[0012] Training samples are generated based on positive and negative pseudo-paired samples.
[0013] Furthermore, constructing training samples also includes:
[0014] Generate a random number r between 0 and 1;
[0015] If r is less than the set threshold, a negative pseudo-paired sample is generated and the corresponding label is set to 0; otherwise, a positive pseudo-paired sample is generated and the corresponding label is set to 1.
[0016] Furthermore, the assessment model for abnormal central cusp includes an intraoral feature extractor, an X-ray feature extractor, and a feature convergence classifier, among which,
[0017] The intraoral imaging feature extractor is used to extract the first feature information contained in intraoral photographs;
[0018] The X-ray feature extractor is used to extract secondary feature information contained in X-ray images;
[0019] The feature convergence classifier is used to converge the first feature information and the second feature information and determine whether there is a deformed central cusp based on the converged feature information.
[0020] Furthermore, the intraoral imaging feature extractor employs six 3×3 convolutional layers. The first five convolutional layers also include a non-linear activation function ReLU and 2×2 max pooling, while the sixth convolutional layer is a single convolution. The number of convolutional kernels in the six convolutional layers are 16, 32, 64, 128, 256, and 512, respectively.
[0021] Furthermore, the parameters of layers 2, 3, 5, and 8 of the VGG16 pre-trained model are initially loaded into layers 3 through 6 of the six convolutional layers in the intraoral imaging feature extractor for training.
[0022] Furthermore, the X-ray feature extractor employs 11 3×3 convolutional layers, with the first 10 convolutional layers each including a non-linear activation function ReLU, and only the 3rd, 5th, 7th, and 9th convolutional layers containing 2×2 max pooling. The 11th convolutional layer contains only a single convolution, with the number of convolutional kernels being 16, 32, 32, 64, 64, 128, 128, 256, 256, 512, and 512, respectively.
[0023] Furthermore, in the X-ray feature extractor, the first three convolutional layers are randomly initialized, while the fourth to eleventh convolutional layers are trained using the convolutional parameters of layers 1, 2, 3, 4, 5, 6, 10, and 11 of the VGG16 pre-trained model as initial parameters.
[0024] The present invention also provides an automatic evaluation device for multimodal deformities with central cusps based on convolutional neural networks, the device comprising:
[0025] The acquisition unit is used to acquire X-ray images and intraoral photographs of the target to be detected;
[0026] The region determination unit is used to determine N X-ray regions in an X-ray film, denoted as X1, X2, ..., X... N And the N intraoral imaging regions corresponding to the N X-ray areas in the intraoral photograph, denoted as P1, P2, ..., P... N The X-ray area and the intraoral imaging area correspond to N positions of the teeth on the left and right sides of the upper and lower jaws, starting from the midline of the teeth. The X-ray area and the intraoral imaging area corresponding to each tooth position form a paired area.
[0027] The training sample generation unit is used to construct training samples using a pseudo-pairing method.
[0028] The model training unit is used to train the abnormal central apex evaluation model using training samples.
[0029] The evaluation unit is used to input each paired region of the target to be detected into the abnormal central cusp evaluation model to determine whether an abnormal central cusp exists in the paired regions of the target to be detected.
[0030] The construction of training samples includes:
[0031] Determine the positive training sample set and the negative training sample set. The positive training sample set refers to the set of samples with positive X-ray areas and positive intraoral imaging areas, and the negative training sample set refers to the set of samples with negative X-ray areas and negative intraoral imaging areas.
[0032] Positive pseudo-paired samples and negative pseudo-paired samples are generated based on the positive training sample set and the negative training sample set, respectively. A positive pseudo-paired sample refers to a pair of samples consisting of an X-ray area and an intraoral area randomly selected from the positive training sample set, and a negative pseudo-paired sample refers to a pair of samples consisting of an X-ray area and an intraoral area randomly selected from the negative training sample set.
[0033] Training samples are generated based on positive and negative pseudo-paired samples.
[0034] The beneficial effects of this invention are:
[0035] 1. This invention constructs training samples through pseudo-pairing, which greatly reduces the workload of sample collection on the one hand, and makes the training samples more diverse on the other hand, which is more conducive to model learning and generalization, enhances the model's feature expression ability, and thus improves the evaluation accuracy of abnormal central cusps.
[0036] 2. By setting the initialization method of parameters in the model, this invention can improve the computational efficiency of the model, as well as its robustness and generalization ability.
[0037] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0038] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will now be described in further detail with reference to the accompanying drawings, wherein:
[0039] Figure 1 This is a schematic flowchart illustrating an automatic evaluation method for multimodal deformed central apex based on a convolutional neural network according to an embodiment of this application;
[0040] Figure 2 This is a schematic flowchart illustrating a training sample construction method according to an embodiment of this application;
[0041] Figure 3 This is a schematic diagram of a deformed central apex model structure according to an embodiment of this application;
[0042] Figure 4 This is a schematic block diagram of an automatic evaluation device for multimodal deformed central tip based on a convolutional neural network, according to an embodiment of this application;
[0043] Figure 5 This is a schematic diagram of the eight regions on an X-ray and intraoral photograph. Detailed Implementation
[0044] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be understood that the preferred embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0045] This application proposes an automatic evaluation method and apparatus for multimodal deformed central tips based on convolutional neural networks.
[0046] Figure 1 This is a schematic flowchart illustrating an automatic evaluation method for multimodal deformed central tips based on convolutional neural networks according to an embodiment of this application.
[0047] like Figure 1 As shown, the method may include:
[0048] S1: Acquire intraoral photographs and X-rays of the target to be examined (e.g., a patient). The intraoral photographs may be images taken by a camera, and the X-rays may be panoramic X-rays obtained by a dental panoramic X-ray camera.
[0049] S2: Determine the matching area based on intraoral photographs and X-rays.
[0050] This step can specifically include: First, identifying N X-ray regions in the X-ray film, denoted as X1, X2, ..., X... N And N intraoral imaging regions corresponding to N X-ray film regions, denoted as P1, P2, ..., P... NThe method for defining the X-ray area and intraoral imaging area can be any image recognition method or image processing software, identifying images at specific locations to generate the X-ray area and intraoral imaging area. The X-ray area and intraoral imaging area correspond to N positions of teeth on the left and right sides of the midline of the upper and lower jaws. The selection of the tooth positions can be based on the location of teeth prone to developing malformed central cusps. For example, this position could be the location of the patient's upper and lower premolars (including the first premolar, second premolar, etc.). In some embodiments, the fourth and fifth teeth on the left and right sides of the midline of the upper and lower jaws (a total of 8 teeth) can be selected, such as... Figure 4 Eight regions in the X-ray and eight regions in the intraoral photograph (i.e. Figure 5 The eight rectangles shown represent the locations. Each tooth location's corresponding X-ray area and intraoral imaging area form a paired region; that is, one paired region corresponds to one tooth location. i With P i A pairing region can be formed by taking (i = 1-N). For example, the X-ray region corresponding to the fourth tooth from the midline of the teeth in the maxilla of the target to be detected and the intraoral region corresponding to that tooth constitute a pairing region. In some embodiments, the resolution of the intraoral region is 90×90, and the resolution of the X-ray region is 130×110.
[0051] S3: Obtain the malformed central apex evaluation model. This malformed central apex evaluation model is a convolutional neural network model. In some embodiments, pseudo-paired samples are used as training samples when training this model. The paired samples include positive pseudo-paired samples and negative pseudo-paired samples. The method for constructing training samples can be found in [link to documentation]. Figure 2 And its specific description.
[0052] S4: Input each paired region of the target to be detected into the trained abnormal central apex evaluation model to determine whether an abnormal central apex exists in the paired regions of the target to be detected. In some embodiments, by X i With P i The paired regions are input into the malformed central cusp assessment model to determine whether the tooth i of the target to be tested has a malformed central cusp.
[0053] This invention also provides an automatic evaluation device for multimodal deformed central cusp based on convolutional neural networks. For example... Figure 4 As shown, the multimodal deformed central tip automatic evaluation device (hereinafter referred to as the "evaluation device") may include an acquisition unit, a region determination unit, a training sample generation unit, a model training unit, and an evaluation unit.
[0054] The acquisition unit can be used to acquire X-ray images and intraoral photographs of the target to be detected. The intraoral photographs can be images taken by a camera, and the X-ray images can be panoramic X-ray images obtained by a dental panoramic X-ray camera.
[0055] The region determination unit can be used to determine N X-ray regions in an X-ray film, denoted as X1, X2, ..., X... N And the N intraoral imaging regions corresponding to the N X-ray areas in the intraoral photograph, denoted as P1, P2, ..., P... N The X-ray area and the intraoral imaging area correspond to N positions of the teeth on the left and right sides of the midline of the upper and lower jaws. The X-ray area and the intraoral imaging area corresponding to each tooth position form a paired area.
[0056] The training sample generation unit can be used to construct training samples using pseudo-pairing. The process of generating training samples by the training sample generation unit can be found in [link to relevant documentation]. Figure 2 And its related descriptions.
[0057] The model training unit can be used to train a model evaluating malformed central cusps using training samples. During model training, the cross-entropy loss function can be used.
[0058] The evaluation unit can be used to input each paired region of the target to be detected into the abnormal central cusp evaluation model to determine whether an abnormal central cusp exists in the paired regions of the target to be detected.
[0059] Figure 2 This is a schematic flowchart illustrating the construction of training samples according to some embodiments of this application. Figure 2 As shown, constructing training samples may include: determining a positive training sample set and a negative training sample set. The positive training sample set refers to the set of samples labeled as positive (i.e., containing a deformed central apex) corresponding to X-ray areas and samples labeled as positive (i.e., containing a deformed central apex) corresponding to intraoral imaging areas. The negative training sample set refers to the set of samples labeled as negative (i.e., not containing a deformed central apex) corresponding to X-ray areas and samples labeled as negative (i.e., not containing a deformed central apex) corresponding to intraoral imaging areas. A positive label indicates the presence of a deformed central apex, and a negative label indicates the absence of a deformed central apex. It should be noted that the determination of the positive and negative training sample sets can be completed in the same step or in different steps; this is not restricted here.
[0060] Next, positive pseudo-pairing samples and negative pseudo-pairing samples can be generated based on the positive training sample set and the negative training sample set, respectively. The positive pseudo-pairing samples and negative pseudo-pairing samples together constitute the generated training samples.
[0061] In this invention, a positive pseudo-paired sample refers to a pair of samples consisting of an X-ray region and an intraoral region randomly selected from the positive training sample set, and a negative pseudo-paired sample refers to a pair of samples consisting of an X-ray region and an intraoral region randomly selected from the negative training sample set. Positive or negative pseudo-paired samples are generated through pseudo-pairing. Specifically, a pseudo-paired sample does not require the intraoral region and the X-ray region to correspond to the same tooth location. For example, the intraoral region corresponding to tooth 2 can not only form a sample with the X-ray region corresponding to tooth 2, but also form a sample with the X-ray regions corresponding to other teeth besides tooth 2 (e.g., tooth 3). These samples can be called pseudo-paired samples. This invention generates training samples through pseudo-pairing, greatly enhancing the model's feature representation capability.
[0062] Random pseudo-pairing can reduce the workload of sample collection. For example, suppose there are 1000 positive intraoral radiograph regions and 1000 corresponding positive X-ray regions. Under normal pairing, only 1000 paired samples can be generated. However, by using random pseudo-pairing, the 1000 positive intraoral radiograph regions and 1000 corresponding positive X-ray regions can generate 1000 paired samples. 6 There are 1000 spurious paired samples (i.e., positive spurious paired samples). Similarly, assuming there are 1000 negative intraoral radiograph regions and 1000 corresponding negative X-ray film regions, under normal pairing conditions, only 1000 paired samples can be generated. However, through random spurious pairing, these 1000 negative intraoral radiograph regions and 1000 negative X-ray film regions can generate 1000 paired samples. 6 The presence of a few pseudo-paired samples (i.e., negative pseudo-paired samples) greatly increases the number of samples and also makes the samples more diverse, which is beneficial to the model's learning and generalization.
[0063] In some embodiments, the positive training sample set and the negative training set may be stored in four folders: a folder storing positive intraoral radiographs, a folder storing positive X-ray images, a folder storing negative intraoral radiographs, and a folder storing negative X-ray images. In actual manual annotation processes, class imbalance may occur; for example, some folders may contain more than 2000 samples, while other folders may contain only 1000 samples. If uniform random sampling is performed, it may result in class imbalance in the generated training samples (i.e., too many positive pairs or too many negative pairs), which will cause the model's prediction results to be skewed towards a particular class, thereby reducing the model's generalization ability and prediction performance. Therefore, to solve this problem, it is necessary to ensure, as far as possible, that the number of positive and negative pairs in each batch of input data (i.e., the input samples) is similar when constructing the samples.
[0064] In some embodiments, a random number r between 0 and 1 can be generated; then r is compared with a set threshold (e.g., 0.5). If r is less than the set threshold, an intraoral imaging region and an X-ray region are randomly selected from the negative training sample set to generate a negative pseudo-paired sample, and the corresponding label is set to 0; otherwise, if r is not less than the set threshold, an intraoral imaging region and an X-ray region are randomly selected from the positive training sample set to generate a positive pseudo-paired sample, and the corresponding label is set to 1. This can be expressed as:
[0065]
[0066] Where f(r) is the generated label, and r is a randomly generated floating-point number.
[0067] Figure 3 This is a schematic block diagram illustrating an assessment model for a deformed central cusp according to an embodiment of this application. Figure 3 As shown, the assessment model for abnormal central cusp includes an intraoral imaging feature extractor, an X-ray feature extractor, and a feature convergence classifier.
[0068] An intraoral imaging feature extractor is used to extract first feature information contained in intraoral photographs (e.g., intraoral imaging regions);
[0069] The X-ray feature extractor is used to extract secondary feature information contained in an X-ray (e.g., an X-ray region);
[0070] The feature convergence classifier is used to converge the first feature information and the second feature information and determine whether there is a deformed central cusp based on the converged feature information.
[0071] In some embodiments, image preprocessing is required before feature extraction. During the preprocessing stage, intraoral photographs and X-rays may be normalized.
[0072] In some embodiments, intraoral photographs may use an RGB color model, the normalization formula of which can be expressed as:
[0073]
[0074] Where I is the input image, and μ and σ represent the mean and standard deviation vectors of the color channels, respectively, with values of (0.485, 0.456, 0.406) and (0.229, 0.224, 0.225).
[0075] In some embodiments, the intraoral imaging feature extractor can be a 6-layer convolutional layer model, which can be represented as:
[0076] F pi =C p6(MP p5 (C p5 (…MP p1 (C p1 (P i ))…))),
[0077] Among them, C pn and MP pn These are the nth convolutional layer and the max pooling layer, respectively, F pi It is the feature representation of the i-th intraoral illumination region.
[0078] In some embodiments, the intraoral feature extractor may employ six 3×3 convolutional layers, wherein the first five convolutional layers further include a non-linear activation function ReLU and 2×2 max pooling, the sixth convolutional layer is a single convolution, and after all convolutional operations, global average pooling is performed to obtain a 512-channel 1×1 feature representation. In some embodiments, the padding of the six convolutional layers may be set to 1.
[0079] In some embodiments, the number of convolutional kernels in the six convolutional layers can be 16, 32, 64, 128, 256, and 512, respectively.
[0080] In deep neural networks, layers closer to the input capture lower-level, more specific image features, such as edges and textures, which typically exhibit good transferability and can be effectively applied to different tasks and datasets. Layers closer to the output capture higher-level, more abstract image features, such as object parts and overall shape. In VGG16, the earlier convolutional layers are located at the beginning of the network, meaning they capture moderate image features—neither too abstract nor too specific—making them ideal as initial training parameters. Furthermore, early convolutional layer parameters may be affected by overfitting and vanishing gradients during training, while later convolutional layer parameters may be overly specialized, only performing well on specific tasks during VGG16 training. The earlier convolutional layer parameters achieve a good balance between robustness and adaptability. Therefore, in some embodiments of this application, layers 3-6 of the six convolutional layers in the intraoral imaging feature extractor are initially trained using parameters from layers 2, 3, 5, and 8 of the VGGv6 pre-trained model.
[0081] In some embodiments, the X-ray feature extractor may employ an 11-layer convolutional model. After passing through the 11-layer convolutional model, a feature representation F for the i-th X-ray region can be generated. xi .
[0082] In some embodiments, it may employ 11 3×3 convolutional layers, wherein the first 10 convolutional layers each include a non-linear activation function ReLU, and only the 3rd, 5th, 7th, and 9th convolutional layers contain 2×2 max pooling, while the 11th convolutional layer contains only a single convolution. In some embodiments, the number of convolutional kernels in each layer of this 11-layer convolutional model may be 16, 32, 32, 64, 64, 128, 128, 256, 256, 512, and 512, respectively.
[0083] In some embodiments, of the 11 convolutional layers in the X-ray feature extractor, the first three convolutional layers are randomly initialized, while the 4th to 11th convolutional layers are trained using the convolutional parameters of layers 1, 2, 3, 4, 5, 6, 10, and 11 of the VGG16 pre-trained model as initial parameters. This is because the first few layers of a convolutional neural network are typically responsible for extracting low-level features of an image, such as edges and textures. These features may vary across different image datasets. For example, dental images and the ImageNet dataset used by the VGG16 pre-trained model may have significantly different visual characteristics. Therefore, using random initialization for the first three convolutional layers instead of pre-trained parameters allows the model to better adapt to the specific dataset. Typically, randomly initialized convolutional layer parameters require more computational resources and time for training. However, the first few layers of a convolutional neural network have relatively few parameters, requiring relatively less computational resources and time. Therefore, using random initialization for the first three convolutional layers instead of pre-trained parameters does not significantly impact the efficiency of the training process. Moreover, over-reliance on pre-trained parameters may lead to overfitting of the model to the training data, reducing the model's generalization ability. Using random initialization for the first three convolutional layers can increase the model's robustness and improve its generalization ability. By setting the first three convolutional layers to use random initialization, more flexibility is available to adjust the model's parameters and structure, allowing the model to better meet specific needs.
[0084] like Figure 3 As shown, after feature extraction, for the feature convergence classifier part, F can be... pi and F xi The vectors are concatenated into a feature vector (e.g., 1024-dimensional), then passed through a fully connected layer, and finally through a softmax layer for binary classification to determine whether the paired input regions have a deformed central apex. This process can be expressed by the following formula:
[0085]
[0086] in, Indicates the predicted label, [F pi F xi] represents the concatenated feature vector, FC() is a fully connected layer, and softmax represents the softmax layer function.
[0087] During the model training phase, samples generated through pseudo-pairing can be input into the model for training. During training, the cross-entropy loss function L can be used, expressed by the formula:
[0088]
[0089] Among them, y i M represents the true label, and M represents the number of training samples.
[0090] In some embodiments, the Adam optimizer can be used when training the model, and the learning rate can be set to 0.001.
[0091] After training, an optimized model for predicting malformed central cusp is generated. Based on this model, the presence of a central cusp in a patient can be automatically determined, thereby greatly reducing the workload of medical personnel.
[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. An automatic evaluation method for multimodal deformed central cusp based on convolutional neural networks, characterized in that, include: Obtain X-ray images and intraoral photographs of the target to be detected; Identify N X-ray regions in an X-ray film, denoted as X1, X2, ..., X... N And the N intraoral imaging regions corresponding to the N X-ray areas in the intraoral photograph, denoted as P1, P2, ..., P... N The X-ray area and the intraoral imaging area correspond to N positions of the teeth on the left and right sides of the upper and lower jaws, starting from the midline of the teeth. The X-ray area and the intraoral imaging area corresponding to each tooth position form a paired area. A model for assessing abnormal central cusp is obtained, which is trained using pseudo-paired samples as training samples. The pseudo-paired samples include positive pseudo-paired samples and negative pseudo-paired samples. Each paired region of the target to be detected is input into the abnormal central cusp evaluation model to determine whether an abnormal central cusp exists in the paired regions of the target to be detected. The construction of training samples includes the following steps: Determine the positive training sample set and the negative training sample set. The positive training sample set refers to the set of samples with positive labels corresponding to the X-ray area and the set of samples with positive labels corresponding to the intraoral area. The negative training sample set refers to the set of samples with negative labels corresponding to the X-ray area and the set of samples with negative labels corresponding to the intraoral area. A positive label indicates the presence of a deformed central cusp, and a negative label indicates the absence of a deformed central cusp. Positive pseudo-paired samples and negative pseudo-paired samples are generated based on the positive training sample set and the negative training sample set, respectively. A positive pseudo-paired sample refers to a pair of samples consisting of an X-ray area and an intraoral area randomly selected from the positive training sample set, and a negative pseudo-paired sample refers to a pair of samples consisting of an X-ray area and an intraoral area randomly selected from the negative training sample set. Training samples are generated based on positive and negative pseudo-pairing samples; The assessment model for abnormal central cusp includes an intraoral feature extractor, an X-ray feature extractor, and a feature convergence classifier. The intraoral imaging feature extractor is used to extract the first feature information contained in intraoral photographs; The X-ray feature extractor is used to extract secondary feature information contained in X-ray images; The feature convergence classifier is used to converge the first feature information and the second feature information and determine whether there is a deformed central cusp based on the converged feature information; The intraoral imaging feature extractor uses six 3×3 convolutional layers. The first five convolutional layers also include a non-linear activation function ReLU and 2×2 max pooling. The sixth convolutional layer is a single convolution. The number of convolutional kernels in the six convolutional layers are 16, 32, 64, 128, 256 and 512, respectively. The parameters of layers 3 to 6 of the six convolutional layers in the intraoral imaging feature extractor are initially loaded as the parameters of layers 2, 3, 5 and 8 of the VGG16 pre-trained model for training.
2. The automatic evaluation method for multimodal deformed central cusp based on convolutional neural networks according to claim 1, characterized in that, Building training samples also includes: Generate a random number r between 0 and 1; If r is less than the set threshold, a negative pseudo-paired sample is generated and the corresponding label is set to 0; otherwise, a positive pseudo-paired sample is generated and the corresponding label is set to 1.
3. The automatic evaluation method for multimodal deformed central cusp based on convolutional neural networks according to claim 1, characterized in that, The X-ray feature extractor uses 11 3×3 convolutional layers. The first 10 convolutional layers all include a non-linear activation function ReLU, and only the 3rd, 5th, 7th and 9th convolutional layers contain 2×2 max pooling. The 11th convolutional layer contains only a single convolution. The number of convolutional kernels are 16, 32, 32, 64, 64, 128, 128, 256, 256, 512 and 512, respectively.
4. The automatic evaluation method for multimodal deformed central cusp based on convolutional neural networks according to claim 3, characterized in that, In the X-ray feature extractor, the first three convolutional layers are randomly initialized, while the fourth to eleventh convolutional layers are trained using the convolutional parameters of layers 1, 2, 3, 4, 5, 6, 10, and 11 of the VGG16 pre-trained model as initial parameters.
5. An automatic evaluation device for multimodal deformed central cusp based on convolutional neural networks, characterized in that, include: The acquisition unit is used to acquire X-ray images and intraoral photographs of the target to be detected; The region determination unit is used to determine N X-ray regions in an X-ray film, denoted as X1, X2, ..., X... N And the N intraoral imaging regions corresponding to the N X-ray areas in the intraoral photograph, denoted as P1, P2, ..., P... N The X-ray area and the intraoral imaging area correspond to N positions of the teeth on the left and right sides of the upper and lower jaws, starting from the midline of the teeth. The X-ray area and the intraoral imaging area corresponding to each tooth position form a paired area. The training sample generation unit is used to construct training samples using a pseudo-pairing method. The model training unit is used to train the abnormal central apex evaluation model using training samples. The evaluation unit is used to input each paired region of the target to be detected into the abnormal central cusp evaluation model to determine whether an abnormal central cusp exists in the paired regions of the target to be detected. The construction of training samples includes: Determine the positive training sample set and the negative training sample set. The positive training sample set refers to the set of samples with positive X-ray areas and positive intraoral imaging areas, and the negative training sample set refers to the set of samples with negative X-ray areas and negative intraoral imaging areas. Positive pseudo-paired samples and negative pseudo-paired samples are generated based on the positive training sample set and the negative training sample set, respectively. A positive pseudo-paired sample refers to a pair of samples consisting of an X-ray area and an intraoral area randomly selected from the positive training sample set, and a negative pseudo-paired sample refers to a pair of samples consisting of an X-ray area and an intraoral area randomly selected from the negative training sample set. Training samples are generated based on positive and negative pseudo-pairing samples; The deformed central cusp assessment model includes an intraoral feature extractor, an X-ray feature extractor, and a feature convergence classifier. The intraoral imaging feature extractor is used to extract the first feature information contained in intraoral photographs; The X-ray feature extractor is used to extract secondary feature information contained in X-ray images; The feature convergence classifier is used to converge the first feature information and the second feature information and determine whether there is a deformed central cusp based on the converged feature information; The intraoral imaging feature extractor uses six 3×3 convolutional layers. The first five convolutional layers also include a non-linear activation function ReLU and 2×2 max pooling. The sixth convolutional layer is a single convolution. The number of convolutional kernels in the six convolutional layers are 16, 32, 64, 128, 256 and 512, respectively. The parameters of layers 3 to 6 of the six convolutional layers in the intraoral imaging feature extractor are initially loaded as the parameters of layers 2, 3, 5 and 8 of the VGG16 pre-trained model for training.
6. The automatic evaluation device for multimodal deformed central cusp based on convolutional neural networks according to claim 5, characterized in that, The training sample generation unit performs the following steps when constructing training samples: Generate a random number r between 0 and 1; If r is less than the set threshold, a negative pseudo-paired sample is generated and the corresponding label is set to 0; otherwise, a positive pseudo-paired sample is generated and the corresponding label is set to 1.
Citation Information
Patent Citations
Intelligent periodontitis detection method and system based on convolutional neural network
CN112037913A
6D attitude estimation data set migration method based on image content and style decoupling
CN114742890A