Method and system for multi-label recognition of oral digital photos based on regional symptom evidence

CN122530213BActive Publication Date: 2026-09-04SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611016132.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-04
Estimated Expiration
2046-07-09

AI Technical Summary

Technical Problem

[0010]本发明在于解决现有口腔数码照片多标签识别方法中局部病变特征易被全局表征淹没、逐标签查询结构缺乏解剖区域与症状的组合表达、以及原型异常检测缺乏分区域健康参照的技术问题,提供了一种基于区域症状证据的口腔数码照片多标签识别方法与系统,其通过将图像词元软路由为多个隐式口腔解剖区域词元,以分区域健康状态原型构建健康参考并计算病理残差,形成区域-症状证据矩阵,再经区域权重和症状权重因子化聚合得到各标签概率,使多标签识别过程体现口腔临床观察中“病变发生位置”与“病变表现形式”的组合关系,从而提高识别的准确性和可解释性

Benefits of technology

(1)本发明通过隐式口腔解剖区域软路由模块,将图像词元聚合成有限数量的区域词元,每个区域词元对应具有临床语义的解剖区域,模型能够在弱监督条件下自动学习区域分配,无需像素级分割标注,解决了现有全局池化方法难以聚焦局部病变区域的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122530213B_ABST
    Figure CN122530213B_ABST
Patent Text Reader

Abstract

The application discloses a kind of oral cavity digital photo multi-label identification method and system based on regional symptom evidence, belong to oral cavity disease intelligent screening technical field.The application obtains image token sequence by inputting oral cavity digital photo into visual token encoder, and image token is aggregated into multiple regional tokens by implicit oral cavity dissection area soft routing module;For each regional token, obtain healthy reference token from corresponding health state prototype library and calculate pathological residual error;Pathological residual error is mapped as regional-symptom evidence matrix;For each label, set regional weight and symptom weight, and obtain each label probability and determine determination threshold by factorizing and aggregating evidence matrix.When identifying, input photo into model, output probability and compare with threshold to obtain determination result.The application can realize multi-label identification under weak supervision without pixel-level regional annotation, and provide interpretable information, suitable for primary oral health screening and remote auxiliary diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent screening technology for oral diseases, specifically to a method and system for multi-label recognition of oral digital photographs based on regional symptom evidence. Background Technology

[0002] Periodontitis is a common chronic inflammatory oral disease, often accompanied by symptoms such as gingival redness and swelling, gingival bleeding, tartar buildup, gingival recession, tooth loss, or edentulous spaces. Traditional diagnosis typically relies on periodontal probing, attachment loss assessment, imaging data, and a comprehensive judgment by a dental specialist. While these methods are highly accurate, they suffer from drawbacks such as high time costs, heavy reliance on professional personnel, and limitations in examination conditions, particularly in primary screening at the grassroots level, remote consultation, follow-up management, and large-scale population screening.

[0003] To improve screening efficiency, digital oral photographs, as a non-invasive, low-cost, and readily available clinical resource, have been explored for auxiliary screening of periodontitis in recent years. However, digital oral photographs differ significantly from general natural images: their areas of interest are typically local structures such as the dentition, gingival margin, cervical region, interdental papillae, and oral mucosa; images often contain soft tissue occlusion, specular reflection, local shadows, white balance differences, slight defocus, handheld shooting angle shifts, and resolution differences between different devices; furthermore, labels such as gingival redness, tartar, and tooth pigmentation are highly dependent on color, local texture, and morphological boundaries. These unique characteristics necessitate an automatic recognition method for digital oral photographs that can extract local details while also considering global contextual information.

[0004] In recent years, deep learning technology has made significant progress in the field of medical image analysis. Compared with traditional image processing methods, deep learning methods can automatically learn multi-level, multi-scale feature representations from data, and have demonstrated superior performance to traditional methods in various medical image classification, detection, and segmentation tasks. In particular, models based on convolutional neural networks and visual Transformers, with their powerful feature extraction and pattern recognition capabilities, have been widely applied to the automatic recognition and assisted diagnosis of oral images. Deep learning-based oral image recognition methods can automatically learn the mapping relationship from images to disease labels from large amounts of labeled data without the need for manual feature design, thus becoming the mainstream technical approach for intelligent recognition of digital oral photographs.

[0005] However, existing deep learning-based oral image recognition methods still have the following shortcomings: First, global pooling classification methods. These methods compress the entire image into a single global representation, and then output multiple labels through a classification head. This method is prone to having local lesion features obscured by teeth, background, or other large normal areas, resulting in insufficient ability to identify subtle lesions in critical areas such as the gingival margin and neck of the tooth.

[0006] Second, object detection or semantic segmentation methods. These methods typically require pixel-level or bounding box annotations and often use segmentation results of teeth, gums, mucosa, etc., as the basis for subsequent detection. These methods are highly dependent on annotation costs and segmentation accuracy, and segmentation errors may propagate to the final multi-label results, limiting their applicability under weakly supervised annotation conditions.

[0007] Third, the multi-label Transformer method. Addressing the shortcomings of global pooling classification, the multi-label Transformer method sets a label query vector for each category and utilizes cross-attention to retrieve category-related features from the image feature map. This type of method can obtain class-by-class visual representations, but it still has limitations: for digital oral photographs, using only disease labels as query entry points makes it difficult to explicitly express the combined relationships between "anatomical regions, health deviations, and symptom manifestations" in oral diagnosis, leading to insufficient interpretability of the model.

[0008] Fourth, prototype learning methods. In recent years, researchers have introduced prototype learning into deep learning frameworks. Prototypes are typically constructed using normal samples or local normal fragments, and anomaly detection is based on the degree of deviation between the test image and the normal prototype. This approach can be used for anomaly localization or anomaly classification. However, simply comparing the entire image or ordinary image patches with the normal prototype cannot fully express the structured correspondence between oral labels such as gingival redness, gingival recession, and tartar, and different oral anatomical regions and symptom evidence.

[0009] In summary, existing oral image recognition methods have the following main shortcomings: global pooling is difficult to focus on local lesion areas, target detection and semantic segmentation rely too heavily on pixel-level annotation, label query methods are difficult to express the combined relationship between regions and symptoms, and ordinary prototype detection lacks regionalized structured reasoning. Summary of the Invention

[0010] This invention addresses the technical problems in existing multi-label recognition methods for oral digital photographs, such as the tendency for local lesion features to be overwhelmed by global representation, the lack of combined anatomical region and symptom expression in label-by-label query structures, and the lack of regional health references in prototype anomaly detection. It provides a multi-label recognition method and system for oral digital photographs based on regional symptom evidence. This method soft-routes image terms into multiple implicit oral anatomical region terms, constructs health references based on regional health state prototypes, calculates pathological residuals, and forms a region-symptom evidence matrix. The probability of each label is then obtained through factorization and aggregation of region weights and symptom weights. This allows the multi-label recognition process to reflect the combined relationship between "lesion location" and "lesion manifestation" in oral clinical observation, thereby improving the accuracy and interpretability of the recognition.

[0011] The objective of this invention is mainly achieved through the following technical solutions: A multi-label recognition method for oral digital photographs based on regional symptom evidence, including multi-label recognition model training steps and recognition steps; The training steps for the multi-label recognition model include: Step S11: Obtain multiple digital oral photos and generate a multi-label vector for each digital oral photo. Divide the multiple digital oral photos into a training set, a validation set, and a test set. The multi-label vector includes periodontitis labels and / or oral feature labels. Step S12: Preprocess the digital oral imaging; Step S13: Input the preprocessed oral digital photo into a visual word encoder. The visual word encoder is used to encode the input oral digital photo into an image word sequence composed of multiple image words, each image word corresponding to a local area in the oral digital photo. Step S14: Input the image word sequence into the implicit oral anatomy region soft routing module. The implicit oral anatomy region soft routing module generates the allocation weight of each image word belonging to each preset oral anatomy region, obtains the region soft allocation matrix, and aggregates multiple image words into multiple oral anatomy region words according to the region soft allocation matrix; wherein, the rows of the region soft allocation matrix correspond to image words, and the columns correspond to oral anatomy regions. Step S15: For each oral anatomical region terminology, obtain the health reference terminology from the health status prototype library corresponding to the oral anatomical region, and calculate the pathological residual terminology between the oral anatomical region terminology and the health reference terminology. Step S16: Convert the pathological residual lexical units into a region-symptom evidence matrix through a learnable linear mapping, wherein the rows of the region-symptom evidence matrix correspond to the oral anatomical regions and the columns correspond to the symptom dimensions. Step S17: Set learnable region weights and symptom weights for each label to be identified, and perform factorized aggregation based on the region weights, symptom weights, and region-symptom evidence matrix to obtain the log odds of each label to be identified. Step S18: Obtain the probability of each label to be identified based on the log odds using the sigmoid function, train the model using the multi-label binary classification loss function, and determine the judgment threshold for each label on the validation set. The identification steps include: Step S21: Process the digital oral cavity photograph to be identified using the same preprocessing method as in the training phase, and then input it into the trained multi-label recognition model; Step S22: The multi-label recognition model outputs the probability of each label, compares the probability of each label with the corresponding label judgment threshold, and obtains the positive or negative judgment result of each label; where positive indicates the existence of the oral disease or feature corresponding to the label, and negative indicates the absence of the label. Step S23: Generate a multi-label recognition report based on the judgment results.

[0012] This invention enables anatomical region localization without pixel-level annotation through regional soft routing, achieves explicit reasoning of "health-deviation-pathology" through health prototype difference, realizes structured diagnosis of "lesion location + manifestation" through region-symptom evidence matrix and factorized aggregation, and finally achieves multi-label recognition under weak supervision.

[0013] Furthermore, the oral feature tags include one or more of the following: dental calculus, tooth surface pigmentation, gingival bleeding, gingival swelling, gingival recession, edentulous spaces, and impacted teeth. These tags represent the most common oral manifestations in clinical dentistry and cover the main visual features related to periodontitis. By setting the above tags, this invention can comprehensively cover the common needs of multi-tag recognition of oral digital photographs.

[0014] Furthermore, in step S11, when generating the multi-label vector, for the same image from multiple source directories, deduplication and merging are performed by calculating the content hash value or reading the file name identifier. Multiple positive labels corresponding to the image are then merged into the same multi-label vector, and the training set, validation set, and test set are then divided. The positive label indicator indicates the presence of a label corresponding to an oral disease or oral feature. In actual data collection, the same digital photograph of an oral cavity may appear in multiple category folders due to having multiple positive labels. Without deduplication, the same image may be incorrectly used as multiple independent samples, causing data leakage and labeling errors. This invention performs deduplication by calculating the content hash value or reading the file name identifier, merging multiple positive labels of the same image into the same multi-label vector, and then dividing the training set, validation set, and test set, thereby ensuring the cleanliness of data partitioning and the accuracy of labels.

[0015] Furthermore, the preprocessing in step S12 includes input size determination, proportional scaling, and gray fill. Input size determination includes: calculating the average width and average height of the digital oral images in the training set; multiplying the average width and average height by an input scaling factor to obtain recommended width and height; scaling down proportionally when the longer side of the recommended width and height is greater than the preset maximum longer side; and rounding the scaled-down width and height to an integer multiple of the size divisor, where the size divisor matches the downsampling rate of the visual word encoder. Proportional scaling and gray fill include: calculating a proportional scaling factor based on the model input size and the original image size; scaling the original image and centering it before pasting it onto a target canvas with a neutral gray background. The original size of digital oral images varies significantly depending on the acquisition device. Directly scaling them to a fixed square size can lead to geometric deformation of structures such as the crown, gingival margin, and edentulous spaces, affecting the model's recognition of key anatomical structures. This invention determines the input size by calculating the average width and height of the training set, so that the input size matches the overall proportion of the dataset; by scaling proportionally and using gray fill, it adapts to the model input requirements while ensuring that the image is not distorted, effectively reducing the geometric distortion of the dentition and gingival margin.

[0016] Furthermore, the implementation method of the implicit oral anatomy region soft routing module in step S14 is as follows: Let the image word sequence be ,in, T It is a three-dimensional real tensor consisting of the batch size multiplied by the number of image terms multiplied by the feature dimension. B For batch size, L The number of image terms, D Let be the feature dimension; and let the preset number of implicit oral anatomical regions be . R The soft routing module contains a multilayer perceptron, which maps the features of each image word to... R The log-odds allocation of the region is calculated, and then normalized along the image lexical dimension to obtain the soft allocation matrix of the region. ,in, A b,l,r Indicates the first b The first sample l The image term belongs to the first... r The probability of a hidden region; No. r The formula for calculating the number of words in an oral anatomical region is:

[0017] in, ϵ To prevent positive constants with a denominator of zero, Indicates the first b The first sample l Feature vectors of image words,l ∈{1, 2, … , L}, r∈{1 , 2 , … , R}.

[0018] The core of the implicit oral anatomy region soft routing module lies in generating allocation weights (i.e., region soft allocation matrices) for each image lexical unit by using a multilayer perceptron to assign it to various preset regions. Then, image lexical units are weighted and aggregated according to these weights into region lexical units, resulting in a significantly smaller number of lexical units. The advantage of this design is that, on the one hand, the number of region lexical units is much smaller than the number of lexical units. R Much smaller than the number of image lexical units L This achieves feature dimensionality reduction and structuring; on the other hand, the region allocation is "soft" (i.e., each word can be allocated to multiple regions in different proportions), rather than "hard" (i.e., each word can only belong to one region), which allows the model to handle the problem of blurred boundaries of oral anatomical regions more flexibly. The rows of the soft region allocation matrix correspond to image words, and the columns correspond to oral anatomical regions. This row and column definition ensures the clarity of the matrix dimensions.

[0019] Furthermore, the preset number of implicit oral anatomical regions is 6, corresponding to the gingival margin, interdental papilla, attached gingiva or gingival body, cervical region, crown region, oral mucosa, or background, respectively; the training process of the region soft allocation matrix also includes region allocation balancing loss. L region This loss is used to suppress the concentration of the soft allocation matrix in a single region, and the specific formula is as follows:

[0020] in, q r The average weight of the r-th region in the soft-assignment matrix of all regions for all words in all samples within a training batch is calculated using the following formula:

[0021] Where B is the batch size and L is the number of image terms.

[0022] The six regions—gingival margin, interdental papilla, attached gingiva or gingival body, cervical region, crown region, and oral mucosa or background—cover the anatomical structures most relevant to periodontitis and oral feature recognition in digital oral photographs. (Region allocation equilibration loss) L region This loss is used to prevent region soft routing from degenerating into using only a single region. Without constraints, the model might assign all terms to a single region, rendering region soft routing meaningless. The region allocation balancing loss encourages the even distribution of weights across regions. q r Approaching 1 / R This enables the model to learn diverse regional representations.

[0023] Furthermore, the specific implementation of step S15 is as follows: Let the first r The health status prototype library corresponding to each oral anatomical region contains P A prototype of a health state For regional lexical units ,calculate U r The similarity to each healthy prototype is calculated using cosine similarity:

[0024] in, The second norm of a vector; H r,p Indicates the first r The first region p A prototype of a healthy state p ∈{1, 2, …, P}; By temperature parameters The similarity is scaled and normalized to obtain the prototype weight:

[0025] in 𝑤 𝑟,𝑝 This represents the weight coefficient of the nth health state prototype in the nth region; H r,q Indicates the first r The prototype of the q-th health state in region . q ∈{1 , 2 , …, P}; Health reference term 𝑁 𝑟 The weighted average of the various health prototypes:

[0026] Then calculate the pathological residuals. D r :

[0027] Where Norm(⋅) denotes L2 normalization, that is, for any input vector 𝑣, .

[0028] The principle of the regional health prototype differential module is as follows: For each oral anatomical region, an independent set of health state prototypes is maintained. These prototypes represent the typical characteristics of the region in a normal or near-normal state. For a region word in the current image, its cosine similarity with each health prototype is calculated, using a temperature parameter... After scaling, prototype weights are obtained through softmax. Then, a weighted average of these prototypes is calculated to obtain the health reference terminology. Finally, the L2 normalized residual between the region terminology and the health reference terminology is calculated as the pathological residual. The core of this design lies in "regional health references." Different regions (such as the gingival margin and crown) have significantly different normal state characteristics, and separate health references should be established for each region, rather than using a uniform global health prototype. Temperature parameter. Controlling the smoothness of the prototype weight distribution: The larger the value, the smoother the weight distribution and the more even the contribution of each prototype. The smaller the value, the sharper the weight distribution, and the more the model tends to choose the most similar prototype.

[0029] Furthermore, the prototypes in the health state prototype library are learnable parameters, and the loss is distributed through health prototypes during training. 𝐿 proto To suppress overlap between different prototypes within the same region, the formula for calculating the dispersion loss of the healthy prototype is:

[0030] Here, cos(⋅,⋅) represents the cosine similarity between two vectors.

[0031] During training, if multiple healthy prototypes within the same region gradually degenerate into the same feature vector, these prototypes will lose diversity and fail to effectively cover different variations of the normal state in that region. The healthy prototype dispersion loss, by penalizing the square of the cosine similarity between different prototypes within the same region, encourages the prototypes to separate from each other, maintains diversity, and thus more comprehensively represents the normal state space of that region.

[0032] Furthermore, the generation method of the region-symptom evidence matrix in step S16 is as follows: Let the preset number of symptom dimensions be . M Set a learnable linear mapping matrix. and bias vector ; For the r Pathological residuals in each region Calculate the symptom evidence vector:

[0033] in, For the first rEach region in the whole M Evidence vectors in each of the symptom dimensions; σ(⋅) is the element-wise sigmoid function; stacking the symptom evidence vectors of all 𝑅 regions yields the region-symptom evidence matrix. .

[0034] Pathological residual Δ r This indicates the direction and degree of deviation of a region's lexical units from a health reference, but the deviation itself does not directly correspond to clinical symptoms. This invention utilizes a learnable linear mapping (matrix)... W evid and bias b evid The pathological residuals are mapped to symptom evidence vectors, and then the evidence values ​​are constrained to between 0 and 1 using the sigmoid function. In this way, the pathological deviation in each region is transformed into evidence strength across multiple symptom dimensions, forming a region-symptom evidence matrix. E The rows of this matrix correspond to oral anatomical regions, and the columns correspond to symptom dimensions, providing a structured evidence base for subsequent disease diagnosis.

[0035] Furthermore, the number of symptom dimensions is 8, corresponding to redness, swelling, recession or root exposure, calculus or plaque deposition, bleeding tendency, ulceration or breakage, abnormal color, and abnormal texture or margins. These eight symptom dimensions cover the main clinical manifestations of periodontitis and oral cavity features, serving as intermediate variables connecting regional deviations to disease labels. By mapping pathological residuals to these interpretable symptom dimensions, this invention can not only output the probability of disease labels but also provide interpretable information on "which region exhibits which symptoms."

[0036] Furthermore, the specific implementation of factorization aggregation in step S17 is as follows: Assume there is a total C One tag to be identified, c ∈{1, 2, …, C} is the label index; for each label, set a learnable region weight vector and a symptom weight vector, and denote the region weight of the c-th label as . Symptom weighting ; The region weights for each label are subjected to softmax normalization, and the symptom weights are also subjected to softmax normalization:

[0037] in, Indicates the first c After normalization of the label, the label is at the...r Weights in each region Indicates the first c After normalization of the label, the label is at the... m Weights on each symptom dimension; Indicates the first c The first label in the r Unnormalized region weights in each region Indicates the first c The first label in the m Unnormalized symptom weights across symptom dimensions; r ′ is the variable for summing the region. Indicates the first c The first label in the r Unnormalized region weights over ′ regions; m ′ represents the summation variable for the symptom dimension. Indicates the first c The first label in the m Unnormalized symptom weights across ' symptom dimensions; Log-odds of the nth label z c The calculation formula is:

[0038] in, b c This is the bias term for the nth label; E r,m This represents the strength of evidence that the nth symptom occurs in the nth region of the region-symptom evidence matrix. The probability of the c-th label p c for: .

[0039] The core idea of ​​the disease-region-symptom factorization aggregation module of this invention is that each disease label has a different degree of dependence on different regions and symptoms. For example, "gingival redness and swelling" depends more on evidence of "redness" and "swelling" in the gingival margin or interdental papilla region, while "dental calculus" depends more on evidence of "calculus or plaque deposition" in the cervical region or crown region. This invention sets learnable region weights and symptom weights for each label, and after softmax normalization, these are combined with the region-symptom evidence matrix. E Bilinear aggregation is performed to obtain the log odds of the label. This factorized aggregation method allows the criteria for each label to be traced back to a specific region and specific symptoms, realizing a structured diagnostic logic of "lesion location + manifestation".

[0040] Furthermore, the model training step also includes a label relation residual calibration module, which receives the initial log-odds vector output by the factorization aggregation module. Through a learnable label relationship matrix The sigmoid activation function generates the residual correction vector r. res Then add this to the initial logarithmic odds to get the final logarithmic odds z:

[0041] in, This is the bias vector.

[0042] In scenarios with a large number of labels and complex co-occurrence relationships among them (e.g., periodontitis often occurs simultaneously with gingival redness and bleeding), the initial log-probability output by the factorization aggregation module may not fully utilize the correlation between labels. The label relationship residual calibration module uses a learnable label relationship matrix... W rel The module generates residual correction values ​​based on the initial probabilities of each label and adds them to the initial log-odds. This module operates in residual form and does not replace the region-symptom evidence matrix as the primary basis for judgment; it is only used as an auxiliary calibration method when there is sufficient data.

[0043] Furthermore, the total loss function used in the model training step is:

[0044] in, For multi-label binary classification, cross-entropy loss, L region To distribute the losses evenly across regions, L proto To mitigate losses in healthy prototypes, l region and l proto To balance the weights of the various loss terms, the total loss consists of three parts: multi-label binary classification cross-entropy loss, region allocation balancing loss, and healthy prototype dispersion loss. The multi-label binary classification cross-entropy loss is the main loss, responsible for optimizing the accuracy of multi-label classification; the region allocation balancing loss and the healthy prototype dispersion loss are auxiliary constraints, respectively responsible for preventing region soft route collapse and healthy prototype degradation. The three loss components are weighted and summed to form a multi-objective optimization training framework.

[0045] Furthermore, the multi-label binary classification cross-entropy loss Using a form with positive sample weights:

[0046] in, B For batch size, C For the total number of tags, y b,c For the first b The first sample c The true value of each label p b,c The probability predicted by the model. α c For the first c The weight of each label's positive sample is calculated using the following formula:

[0047] in, N pos,c and N neg,c The training set c The number of positive and negative samples for each label. β For smoothing coefficients, α min and α max These are the lower and upper bounds for the weights. In multi-label classification, the number of positive samples for different labels often differs significantly (for example, there may be far more samples of swollen gums than samples of missing teeth). Directly using the standard cross-entropy loss can lead to a model bias towards the majority class. This invention compensates for the loss of positive samples by calculating positive sample weights for each label. The calculation of positive sample weights is based on the ratio of the number of positive to negative samples, and after power-law smoothing and upper and lower bound truncation, it can compensate for the minority class without causing training instability due to excessively large weights.

[0048] Furthermore, the method for searching and determining the threshold for each label on the validation set in step S18 is as follows: Let the first c The predicted probability of each label on the validation set is given by [value], and the true label is [value]. Within a preset threshold range, the probability increases by a step size Δ. t Enumerate candidate thresholds t ; For each candidate threshold t Calculate the first c The F1 score for each label is:

[0049] Among them, Precision c ( t ) and Recall c ( t ) are respectively at the threshold t Precision and recall rates; Choose to make F1c ( t The largest threshold is used as the first c Optimal threshold for each label .

[0050] In practical applications, the optimal decision threshold often differs for different labels. Some labels require higher thresholds to ensure precision, while others require lower thresholds to ensure recall. This invention uses the F1 score as a metric on the validation set to independently search for the optimal threshold for each label, ensuring that the decision threshold for each label achieves the best balance between precision and recall. When multiple candidate thresholds correspond to the same F1 score, the threshold closest to the default threshold is selected to maintain output stability.

[0051] A multi-label recognition system for oral digital photographs based on regional symptom evidence includes: The sample construction module is used to acquire digital oral photos and generate multi-label vectors for each digital oral photo; wherein, the multi-label vectors include periodontitis labels and / or oral feature labels; The input processing module is used to preprocess digital dental photographs; A visual lexical encoder is used to encode a preprocessed digital oral image into a sequence of image lexical units, where each image lexical unit corresponds to a local region in the digital oral image. The implicit oral anatomy region soft routing module is used to generate the allocation weights of each image word belonging to each preset oral anatomy region, obtain the region soft allocation matrix, and aggregate multiple image words into multiple oral anatomy region words based on the region soft allocation matrix; wherein, the rows of the region soft allocation matrix correspond to image words, and the columns correspond to oral anatomy regions; the health prototype difference module is used to obtain health reference words from the health state prototype library corresponding to the oral anatomy region for each oral anatomy region word, and calculate the pathological residual words between the oral anatomy region word and the health reference words; The system comprises the following modules: an evidence generation module, which converts pathological residual terms into a region-symptom evidence matrix using a learnable linear mapping, where rows correspond to oral anatomical regions and columns correspond to symptom dimensions; a factorization aggregation module, which sets learnable region and symptom weights for each label to be identified, performs factorization aggregation based on the region weights, symptom weights, and the region-symptom evidence matrix to obtain the log odds of each label, and uses a sigmoid function to obtain the probability of each label; a threshold determination module, which determines a judgment threshold for each label on a validation set partitioned from the training dataset; and a report generation module, which processes the digital oral images to be identified using the same preprocessing method as in the training phase, outputs the probability of each label using the trained model, compares the probability of each label with the corresponding judgment threshold to obtain a positive or negative judgment result for each label, and generates a multi-label recognition report based on the judgment result; where a positive result indicates the presence of the oral disease or feature corresponding to the label, and a negative result indicates its absence.

[0052] In summary, the present invention has the following advantages compared with the prior art: (1) This invention uses an implicit oral anatomy region soft routing module to aggregate image lexical units into a limited number of region lexical units. Each region lexical unit corresponds to an anatomical region with clinical semantics. The model can automatically learn region allocation under weak supervision without pixel-level segmentation and labeling, which solves the problem that existing global pooling methods have difficulty focusing on local lesion regions.

[0053] (2) This invention uses a regional health prototype differential module to maintain a health status prototype library independently for each region, calculates the degree of deviation between regional terms and health references, and explicitly models the reasoning path of "health → deviation → pathology", thus solving the problem of lack of regional health references in existing prototype anomaly detection.

[0054] (3) This invention transforms pathological residuals into evidence values ​​of the symptom dimension through the region-symptom evidence matrix and factorization aggregation module, and aggregates them through the independent region weight and symptom weight of each label, so that the diagnostic process reflects the clinical logic of "lesion location + manifestation form", which solves the problem that the label-by-label query structure lacks the expression of anatomical region and symptom combination.

[0055] (4) The present invention designs regional allocation balance loss and healthy prototype dispersion loss as auxiliary training constraints, which effectively prevents regional soft route collapse and healthy prototype degradation, and ensures training stability and model generalization ability.

[0056] (5) The present invention designs a proportional gray border filling and restricted data enhancement strategy for the characteristics of oral digital photographs, which reduces tooth deformation and color clue destruction, and is more suitable for actual clinical acquisition scenarios. Attached Figure Description

[0057] The accompanying drawings, which are included to provide a further understanding of embodiments of the invention and form part of this application, do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the model structure based on regional symptom evidence in an embodiment of the present invention; Figure 2 This is a flowchart illustrating the multi-label sample construction and model training process in an embodiment of the present invention. Figure 3 This is a flowchart of the oral digital photo recognition process in an embodiment of the present invention. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.

[0059] To enable those skilled in the art to clearly understand and implement this invention, the meanings of the various terms are explained below: A visual lexical encoder is a neural network module used to encode an input digital photograph of a patient's mouth into a sequence of image lexical units. Image lexical units are the basic processing units in a visual Transformer architecture, with each unit corresponding to a local region in the input image. The visual lexical encoder can employ visual backbone networks capable of outputting spatial lexical sequences, such as the Swing Transformer, Visual Transformer, convolutional neural networks, or hybrid convolutional-transformer networks. In this invention, the visual lexical encoder is responsible for converting a two-dimensional image into an ordered sequence of lexical units, providing structured input for subsequent region soft routing.

[0060] The implicit oral anatomy region soft routing module is used to aggregate image word sequences into multiple oral anatomy region words. "Implicit" means that this module does not rely on pixel-level manual region annotations, but is entirely self-organized and learned by the model from multi-label classification tasks; "soft routing" means that each image word can be assigned to multiple regions with different weights, rather than being forced to belong to a single region; "oral anatomy region" refers to a pre-defined anatomical location with clinical semantics, such as the gingival margin or cervical region. The core function of this module is to aggregate a large number of image words (e.g., hundreds) into a smaller number of region words (e.g., 6), achieving feature dimensionality reduction and structuring.

[0061] Region soft assignment matrix: This refers to a matrix used to record the probability of each image word belonging to various preset regions. The rows of this matrix correspond to image words, the columns correspond to oral anatomical regions, and the values ​​of the matrix elements represent the probability of the corresponding word belonging to the corresponding region. All element values ​​are between 0 and 1, and the sum of the probabilities of each row (i.e., each word) across all regions is 1.

[0062] The health status prototype library refers to a set of learnable vectors maintained independently for each oral anatomical region. Each vector represents a typical feature pattern of that region in a normal or near-normal state. The prototype library for each region contains multiple prototypes (e.g., 4) to cover different variations of the normal state of that region (such as differences in normal gingival color among different individuals).

[0063] Health reference lexicon: This refers to the vector obtained by weighted averaging of health prototypes based on the similarity between the current region lexicon and each health prototype. The health reference lexicon represents the normal state features that the model "expects" to see under the current region lexicon features. The difference between the health reference lexicon and the region lexicon reflects the degree of deviation from health.

[0064] Pathological residuals: These are the difference vectors between region terms and healthy reference terms, obtained after L2 normalization. The direction and magnitude of the pathological residuals reflect the deviation pattern of region terms relative to the healthy reference and serve as direct input for subsequent symptom evidence generation.

[0065] Health deviation distance: This refers to the squared Euclidean distance between a regional term and a health reference term. It is a scalar value that reflects the overall degree of deviation of a regional term from its health reference. This distance can serve as a quantitative indicator of the health status of a region.

[0066] Region-Symptom Evidence Matrix: This matrix is ​​arranged with oral anatomical regions as rows and symptom dimensions as columns. The values ​​of the matrix elements represent the strength of evidence for the corresponding symptoms in the corresponding regions. This matrix transforms pathological residuals (abstract vector deviations) into clinically interpretable symptom evidence (such as "the evidence strength for redness in the gingival margin region is 0.85"), serving as an intermediate representation connecting low-level features and high-level disease labels.

[0067] Factorization aggregation refers to the process of performing a bilinear weighted summation of the region-symptom evidence matrix by assigning independent region and symptom weights to each label. This process "factorizes" the two-dimensional evidence matrix into a one-dimensional label log-odds, where the weight distribution of each label reflects the degree to which that label depends on different regions and symptoms.

[0068] Example: like Figure 1As shown, a multi-label recognition system for oral digital photographs based on regional symptom evidence includes: a sample construction module for acquiring oral digital photographs and generating multi-label vectors for each photograph; wherein the multi-label vectors include periodontitis labels and / or oral feature labels; an input processing module for preprocessing the oral digital photographs; a visual lexicon encoder for encoding the preprocessed oral digital photographs into a sequence of image lexicons, each image lexicon corresponding to a local region in the oral digital photograph; an implicit oral anatomy region soft routing module for generating allocation weights for each image lexicon belonging to various preset oral anatomy regions, obtaining a region soft allocation matrix, and weighting and aggregating multiple image lexicons into multiple oral anatomy region lexicons according to the region soft allocation matrix; wherein the rows of the region soft allocation matrix correspond to image lexicons, and the columns correspond to oral anatomy regions; and a health prototype difference module for obtaining health reference lexicons from a health state prototype library corresponding to the oral anatomy region for each oral anatomy region lexicon, and calculating the symptom difference between the oral anatomy region lexicons and the health reference lexicons. The system comprises the following modules: a residual lexical unit; an evidence generation module, which converts pathological residual lexical units into a region-symptom evidence matrix through a learnable linear mapping, wherein the rows of the region-symptom evidence matrix correspond to oral anatomical regions and the columns correspond to symptom dimensions; a factorization aggregation module, which sets learnable region weights and symptom weights for each label to be identified, performs factorization aggregation based on the region weights, symptom weights, and the region-symptom evidence matrix to obtain the log odds of each label to be identified, and obtains the probability of each label to be identified by using a sigmoid function based on the log odds; a threshold determination module, which determines a judgment threshold for each label on a validation set partitioned from the training dataset; and a report generation module, which processes the digital oral images to be identified using the same preprocessing method as in the training phase, outputs the probability of each label through the trained model, compares the probability of each label with the judgment threshold of the corresponding label to obtain a positive or negative judgment result for each label, and generates a multi-label recognition report based on the judgment result; wherein a positive result indicates the existence of the oral disease or feature corresponding to the label, and a negative result indicates its absence.

[0069] A multi-label recognition method for oral digital photographs based on regional symptom evidence includes a multi-label recognition model training step and a recognition step; wherein, the multi-label recognition model training step is as follows: Figure 2As shown, the process includes: Step S11, acquiring multiple digital oral photos and generating multi-label vectors for each digital oral photo, dividing the multiple digital oral photos into a training set, a validation set, and a test set; wherein, the multi-label vectors include periodontitis labels and / or oral feature labels; Step S12, preprocessing the digital oral photos; Step S13, inputting the preprocessed digital oral photos into a visual word encoder, the visual word encoder being used to encode the input digital oral photos into a sequence of image words composed of multiple image words, each image word corresponding to a local region in the digital oral photos; Step S14, inputting the image word sequence into an implicit oral anatomy region soft routing module, the implicit oral anatomy region soft routing module generating allocation weights for each image word belonging to various preset oral anatomy regions, obtaining a region soft allocation matrix, and weighting and aggregating multiple image words into multiple oral anatomy region words according to the region soft allocation matrix; wherein In step S15, for each oral anatomy region term, obtain a health reference term from the health status prototype library corresponding to that oral anatomy region, and calculate the pathological residual term between the oral anatomy region term and the health reference term; in step S16, convert the pathological residual term into a region-symptom evidence matrix through a learnable linear mapping, where the rows of the region-symptom evidence matrix correspond to the oral anatomy region and the columns correspond to the symptom dimension; in step S17, set learnable region weights and symptom weights for each label to be identified, and perform factorized aggregation based on the region weights, symptom weights, and the region-symptom evidence matrix to obtain the log odds of each label to be identified; in step S18, obtain the probability of each label to be identified by the sigmoid function based on the log odds, train the model using a multi-label binary classification loss function, and determine the judgment threshold for each label on the validation set.

[0070] The identification steps in this embodiment are as follows: Figure 3 As shown, the process includes: Step S21, processing the digital photo of the oral cavity to be identified using the same preprocessing method as in the training phase and then inputting it into the trained multi-label recognition model; Step S22, the multi-label recognition model outputs the probability of each label, compares the probability of each label with the judgment threshold of the corresponding label, and obtains the positive or negative judgment result of each label; wherein, positive indicates the existence of the oral disease or feature corresponding to the label, and negative indicates the absence of the label; Step S23, generating a multi-label recognition report based on the judgment result.

[0071] The oral feature labels in this embodiment include one or more of the following: dental calculus, tooth surface pigmentation, gingival bleeding, gingival swelling, gingival recession, edentulous spaces, and impacted teeth. In this embodiment, training samples are first obtained from a digital oral image dataset, which is based on intraoral digital photographs, mobile phone oral photographs, oral camera images, or other similar color oral images. Each image may have one or more positive labels, and the label set includes periodontitis, dental calculus, tooth surface pigmentation, gingival bleeding, gingival swelling, gingival recession, edentulous spaces, and impacted teeth. In specific implementations, the label set can be added, reduced, or replaced according to the actual screening task; for example, labels such as oral ulcers, suspected caries, or crowded teeth can be added. This flexible configuration of the label set allows this solution to adapt to the needs of different screening scenarios.

[0072] In step S11 of this embodiment, when generating multi-label vectors, for the same image from multiple source directories, deduplication and merging are performed by calculating the content hash value or reading the file name identifier. Multiple positive labels corresponding to the image are merged into the same multi-label vector, and then training, validation, and test sets are defined. The positive label indicator indicates the presence of a label corresponding to an oral disease or oral feature. If the data source uses a folder organization method for each category, the same image may appear in multiple category folders due to multiple oral manifestations. To avoid incorrectly processing the same image into multiple independent single-label samples, this embodiment can calculate a content hash value (e.g., SHA-256) for each image file. When multiple files have the same hash value, they are merged into one sample, and all positive categories from their sources are merged into one multi-label vector. The deduplicated samples are first randomly shuffled and then divided into training, validation, and test sets according to a preset ratio. A sample source table can also be output, recording the retained path, label vector, label name, and original source path of each merged sample, facilitating the tracing of the labeling reasons.

[0073] In this embodiment, multiple acquired digital images of the oral cavity and their corresponding multi-label vectors are divided into a training set, a validation set, and a test set according to a preset ratio. The training set is used to train the multi-label recognition model, i.e., to update the model's network parameters; the validation set is used to adjust hyperparameters, monitor the model's training status, and determine the judgment threshold for each label during training; the test set is used to evaluate the model's final recognition performance after training is complete, but it does not participate in the model's training process or the determination of the judgment threshold to ensure the objectivity of the performance evaluation. The division ratio can be adjusted according to the total number of samples and actual needs. In one specific embodiment, all samples are randomly divided into the training set, validation set, and test set in a 7:1:2 ratio.

[0074] In step S12 of this embodiment, the preprocessing includes input size determination, proportional scaling, and gray fill. The input size determination includes: calculating the average width and average height of the oral digital photos in the training set, multiplying the average width and average height by the input scaling factor to obtain the recommended width and height, scaling down proportionally when the long side of the recommended width and height is greater than the preset maximum long side, and rounding the scaled-down width and height to an integer multiple of the size divisor, wherein the size divisor matches the downsampling factor of the visual word encoder. The proportional scaling and gray fill include: calculating the proportional scaling factor based on the model input size and the original image size, scaling the original image and pasting it in the center onto the target canvas with a neutral gray background.

[0075] This embodiment first reads the width and height of each oral digital photograph in the training set and calculates the average width. W avg and average height H avg Then, based on the scaling factor... s Recommended width and height W rec = W avg × s and H rec = H avg × s Scaling factor s The value can be selected between 0.5 and 1.0; this embodiment uses 0.75 to achieve a balance between preserving image details and controlling video memory. If max( W rec , H rec Exceeding the preset maximum length side L max Then W rec and H rec Reduce by the same proportion, so that the longer side does not exceed L max In this embodiment L max Take 768 pixels. Then, round the reduced width and height to the divisor of the dimensions. D size In this embodiment, it is an integer multiple of the total number of units. D size The value is set to 32 to fit the multi-stage downsampling structure of the visual lexical encoder; a minimum size is also set. S min =224 pixels, ensuring the input size is not too small. The final input size is [ H ,W ].

[0076] For any image to be processed, calculate the scaling factor using the following formula:

[0077] in, W orig and H orig This represents the width and height of the original image. After scaling the image proportionally, center it to a size of [size missing]. W × H On the RGB canvas, uncovered areas are filled with neutral gray. In this embodiment, the gray pixel values ​​are set to [114, 114, 114], which is close to the mean color of the ImageNet dataset, helping to maintain the consistency of the input distribution of the pre-trained model. This proportional gray edge filling method can reduce the change in the crown width-to-height ratio caused by direct stretching and avoid the loss of gingival margin information caused by random cropping.

[0078] In this embodiment, data augmentation is performed after preprocessing. The data augmentation in this embodiment is selected and constrained based on common acquisition perturbations in digital oral imaging. Augmentation is performed after proportional gray border filling to ensure that the boundary conditions seen by the model are consistent with those in the inference stage. The augmentation methods in this embodiment include geometric augmentation, lighting augmentation, color augmentation, and quality augmentation. Geometric augmentation includes: small-angle rotation: performed with a preset probability (e.g., 0.5), with rotation angles ranging from [−12...]. ∘ ,12 ∘Uniform random sampling within the specified range; small-range translation: the translation amount does not exceed 4% of the image width or height to simulate slight shifts caused by handheld shooting or mouth mirror angles; weak perspective transformation: the perturbation amplitude at the four corners does not exceed 3% of the image width or height. The above enhancements are used to simulate slight posture changes caused by handheld shooting, mouth mirror angles, and patient mouth opening postures. Illumination enhancements include: brightness perturbation: brightness enhancement coefficients are sampled in the range of 0.75 to 1.25; contrast perturbation: contrast enhancement coefficients are sampled in the range of 0.80 to 1.20; gamma perturbation: gamma coefficients are sampled in the range of 0.80 to 1.25; elliptical shadow simulation: an elliptical grayscale mask is constructed and Gaussian blurred, with a shadow intensity between 0.75 and 0.95 to attenuate the pixel values ​​of local areas of the image, used to simulate local shadows caused by soft tissue of the lips and cheeks, oral instruments, or light source angles. Color enhancements include mild saturation perturbation, with saturation enhancement coefficients sampled in the range of 0.90 to 1.10. Since the identification of gingival redness, tooth discoloration, and tartar is highly dependent on color cues, this embodiment disables hue perturbation by default (even if enabled, the hue shift is limited to a small range of [−0.02, 0.02]). Quality enhancements include slight Gaussian blur and low-intensity noise. The Gaussian blur radius is sampled from 0.2 to 1.0, and the noise standard deviation is sampled from 2.0 to 8.0 to simulate imaging differences caused by different digital devices, focal lengths, and lighting conditions. This embodiment disables vertical flip and strong random cropping by default, and disables ImageNet general automatic enhancement; the horizontal flip probability is set to 0 to preserve the left-right alignment of the teeth and the shooting direction information.

[0079] In this embodiment, a Swing Transformer Small structure is preferably used as the visual lexical encoder. The input image size is... H × W ×3, after passing through the image block segmentation layer, the image is divided into 4×4 pixel blocks, and the initial number of words is obtained. Each word has a dimension of 96. After passing through a four-stage Swin Transformer module (with depths of 2, 2, 18, and 2 respectively), and with 3, 6, 12, and 24 multi-head self-attention heads respectively, the final output is an image word sequence. ,in , D =768. In this embodiment, the visual word encoder can be initialized using weights pre-trained on ImageNet.

[0080] The implementation method of the implicit oral anatomy region soft routing module in step S14 of this embodiment is as follows: Let the image word sequence be... ,in, T It is a three-dimensional real tensor consisting of the batch size multiplied by the number of image terms multiplied by the feature dimension.B For batch size, L The number of image terms, D Let be the feature dimension; and let the preset number of implicit oral anatomical regions be . R The soft routing module contains a two-layer multilayer perceptron: the first layer combines the features of each word. Mapped to 256 dimensions using the GELU activation function; the second layer maps the 256 dimensions to... R The system outputs the log-odds of assignment along the lexical dimension. Then, it performs softmax normalization along the lexical dimension to obtain the region soft assignment matrix. ,in, A b,l,r Indicates the first b The first sample l The image term belongs to the first... r The probability of a hidden region; No. r Oral anatomical region terminology The calculation formula is:

[0081] in, ϵ To prevent positive constants with a denominator of zero, this embodiment sets it to... ϵ =10 -8 ; Indicates the first b The first sample l Feature vectors of image words, l ∈{1, 2, … , L}, r∈{1 , 2 , … , R}.

[0082] It should be noted that the region allocation of the regional soft routing module is completely implicit. During the training process, the model learns these regions indirectly only through multi-label classification loss and region allocation balancing loss, without the need for manual pixel-level region segmentation annotation.

[0083] In this embodiment, the preset number of implicit oral anatomical regions is 6, corresponding to the gingival margin, interdental papilla, attached gingiva or gingival body, cervical region, crown region, and oral mucosa or background. These six regions cover the anatomical structures most relevant to periodontitis and oral feature recognition in digital oral photographs. The training process of the region soft allocation matrix in this embodiment also includes a region allocation balancing loss. L region This loss is used to suppress the concentration of the soft allocation matrix in a single region, and the specific formula is as follows:

[0084] in, q rThe average weight of the r-th region in the soft-assignment matrix of all regions for all words in all samples within a training batch is calculated using the following formula:

[0085] Where B is the batch size and L is the number of image terms. In this embodiment, under the constraint of region allocation equalization loss, the model can spontaneously learn region terms that have semantic correspondence with clinical regions.

[0086] The specific implementation method of step S15 in this embodiment is as follows: Let the first... r The health status prototype library corresponding to each oral anatomical region contains P A prototype of a health state In this embodiment, each region is configured with P=4 health status prototypes, and the prototype dimension is the same as the word feature dimension. D =768. Health Prototype Matrix These are learnable parameters. During initialization, region-specific word features are extracted from all negative samples in the training set, and k-means clustering is performed on each region. k = P The initial values ​​for the prototype are set using cluster centers. If there are not enough pure healthy samples, Xavier random initialization is used.

[0087] For regional words ,calculate U r Similarity with each healthy prototype, using cosine similarity:

[0088] in, The second norm of a vector; H r,p Indicates the first r The first region p A prototype of a healthy state p ∈{1, 2, …, P}; By temperature parameters The similarity is scaled and normalized to obtain the prototype weight:

[0089] in 𝑤 𝑟,𝑝 This represents the weight coefficient of the nth health state prototype in the nth region; H r,q Indicates the first r The prototype of the q-th health state in region . q∈{1 , 2 , …, P} The temperature parameter is used to control the smoothness of the prototype weight distribution. The larger the value, the smoother the weight distribution of the softmax output, and the smaller the weight difference between the prototypes of each health state; The smaller the value, the sharper the weight distribution, and the more the model tends to select the health state prototype most similar to the current region's word units. In this embodiment, the temperature parameter for all regions... All can be set to 0.1, the temperature parameters for each region. It can also be set as a learnable parameter, which can be optimized during training using gradient descent.

[0090] Health reference term 𝑁 𝑟 The weighted average of the various health prototypes:

[0091] Then calculate the pathological residuals. D r :

[0092] Where Norm(⋅) denotes L2 normalization, that is, for any input vector 𝑣, This normalization gives the pathological residuals a unit length, facilitating subsequent mapping to the symptom evidence space.

[0093] In addition, this embodiment can also calculate the health deviation distance. As a quantitative indicator of the degree of deviation in the health status of the region, it can be used to generate interpretable information.

[0094] In this embodiment, the prototypes in the health status prototype library are learnable parameters, and the loss is distributed through health prototypes during the training process. 𝐿 proto To suppress overlap between different prototypes within the same region, the formula for calculating the dispersion loss of the healthy prototype is:

[0095] Here, cos(⋅,⋅) represents the cosine similarity between two vectors.

[0096] In step S16 of this embodiment, the region-symptom evidence matrix is ​​generated as follows: Let the preset number of symptom dimensions be... M This embodiment sets the number of symptom dimensions. M =8, corresponding to redness, swelling, recession or root surface exposure, calculus or plaque buildup, bleeding tendency, ulceration or breakage, abnormal color, and abnormal texture or margins, respectively. A learnable linear mapping matrix is ​​set. and bias vector ; For the r Pathological residuals in each region Calculate the symptom evidence vector:

[0097] in, For the first r Each region in the whole M Evidence vectors for each symptom dimension; σ(⋅) is an element-wise sigmoid function, ensuring that each symptom evidence value is between 0 and 1; stacking the symptom evidence vectors of all 𝑅 regions yields the region-symptom evidence matrix. .

[0098] The generation method of the region-symptom evidence matrix in this embodiment is based on the region-symptom evidence generation module. The design principle of this module is: pathological residual Δ r It is an abstract vector representing the direction and degree of deviation of a region term from a health reference, but the deviation itself does not directly correspond to an understandable clinical symptom. Through a learnable linear mapping, the abstract deviation is transformed into concrete symptom evidence (such as the evidence strength of 0.85 for redness in the gingival margin region), making the model's reasoning process interpretable.

[0099] The specific implementation method of factorization aggregation in step S17 of this embodiment is as follows: Assume there is a total C One tag to be identified, c ∈{1 , 2 , … , C} is the label index; for each label, set a learnable region weight vector and a symptom weight vector, and denote the region weight of the c-th label as . Symptom weighting These weights are learnable parameters and are initialized with zero.

[0100] To make the weights have a probabilistic interpretation, softmax normalization is applied to the region weights of each label, and softmax normalization is also applied to the symptom weights:

[0101] in, Indicates the first c After normalization of the label, the label is at the... r Weights in each region Indicates the first c After normalization of the label, the label is at the... m Weights on each symptom dimension; Indicates the first c The first label in ther Unnormalized region weights in each region Indicates the first c The first label in the m Unnormalized symptom weights across symptom dimensions; r ′ is the variable for summing the region. Indicates the first c The first label in the r Unnormalized region weights over ′ regions; m ′ represents the summation variable for the symptom dimension. Indicates the first c The first label in the m Unnormalized symptom weights across ' symptom dimensions; Log-odds of the nth label z c The calculation formula is:

[0102] in, b c This is the bias term for the nth label; E r,m This represents the strength of evidence that the nth symptom occurs in the nth region of the region-symptom evidence matrix. The probability of the c-th label p c for: .

[0103] Different disease labels rely on different regions and symptoms to varying degrees. For example, the "red and swollen gums" label relies more on evidence of "redness" and "swelling" in the gingival margin or interdental papilla area; therefore, its corresponding regional weight is higher in the gingival margin and interdental papilla areas, while its symptom weight is higher in the "redness" and "swelling" dimension. Similarly, the "tartar" label relies more on evidence of "tartar or plaque deposition" in the cervical or crown area; therefore, its corresponding regional weight is higher in the cervical and crown areas, while its symptom weight is higher in the "tartar or plaque deposition" dimension. This two-factor decomposition of location and manifestation allows the criteria for each label to be traced back to a specific anatomical region and specific symptom presentation.

[0104] The model training steps in this embodiment also include an optional label relationship residual calibration module. This module can be enabled when the amount of data is sufficient (e.g., the training set exceeds 5000 images) and the label co-occurrence relationships are complex. This module receives the initial log-odds vector output by the factorization aggregation module. Through a learnable label relationship matrix The sigmoid activation function generates the residual correction vector r.res Then add this to the initial logarithmic odds to get the final logarithmic odds z:

[0105] in, This is the bias vector.

[0106] This module operates in residual form and does not replace the region-symptom evidence matrix as the primary criterion; it serves only as an auxiliary calibration tool when label co-occurrence relationships are complex. Label Relationship Matrix W rel Initialize to the identity matrix, bias b rel Initialize to 0.

[0107] This embodiment employs a phased training strategy to ensure convergence stability.

[0108] Phase 1: Train only the visual lexical encoder and the temporary multi-label classification head (without using modules such as region soft routing), with a learning rate set to 1×10⁻⁶. −4 Using the AdamW optimizer, weight decay is set to 1×10. −5 Train for 20 cycles.

[0109] Phase 2: Freeze the parameters of the visual lexical encoder and train the region soft routing module, health prototype differential module, symptom evidence generation module, and factorization aggregation module separately. The learning rate is set to 1×10⁻⁶. −3 Train for 10 epochs. Then use the total loss function:

[0110] in, For multi-label binary classification, cross-entropy loss, L region To distribute the losses evenly across regions, L proto To mitigate losses in healthy prototypes, l region and l proto To balance the weighting coefficients of each loss term, this embodiment sets them to 0.01.

[0111] Phase 3: Unfreeze all parameters, fine-tune end-to-end, and set the learning rate to 1×10. -5 Train for 10 cycles. At this point... l region and l proto It can be reduced appropriately (e.g., each reduced by half).

[0112] In this embodiment, when using pre-trained weights, only the pre-trained weights of the visual lexical encoder (such as patchbedding, each encoding layer, normalization layer, and positional encoding) are loaded, ignoring the original classification head; the newly added region soft routing module, health prototype difference module, symptom evidence generation module, and factorization aggregation module are initialized separately. During the frozen training phase, only the visual lexical encoder can be frozen, keeping the newly added modules trainable, so as to first learn the regions, health deviations, and symptom evidence space relevant to oral tasks.

[0113] This embodiment uses multi-label binary classification cross-entropy loss. Using a form with positive sample weights:

[0114] in, B For batch size, C For the total number of tags, y b,c For the first b The first sample c The true value of each label p b,c The probability predicted by the model. α c For the first c The positive sample weights of each label.

[0115] Positive sample weights α c The calculation formula is:

[0116] in, N pos,c and N neg,c The training set c The number of positive and negative samples for each label; β The smoothing coefficient is set to 0.5 in this embodiment; α min and α max As the lower and upper limits of the weight, in this embodiment, α min Set to 1.0, α max Set to 4.0.

[0117] After training, a threshold is searched for each label on the validation set. The threshold search range is set to 0.05 to 0.95, with a step size of 0.01. Let the... c The predicted probability of each label on the validation set is: The real label is Within the preset threshold range Inner step size Δ t Enumerate candidate thresholds t ; For each candidate threshold t Calculate the first c The F1 score for each label is:

[0118] Among them, Precision c ( t ) and Recall c ( t ) are respectively at the threshold t Precision and recall rates; selection of F1 c ( t The largest threshold is used as the first c Optimal threshold for each label If multiple candidate thresholds have the same F1 score, the candidate threshold closest to the default threshold of 0.5 is selected. The optimal threshold for each label is saved as a threshold file by label name and read by category name during testing or actual screening to avoid errors caused by changes in label order.

[0119] In this embodiment, during the recognition steps and report generation, the digital photo of the oral cavity to be recognized is obtained. First, the same preprocessing as in the training phase is performed: converting it to an RGB image, and then performing proportional gray border filling and normalization according to the input size determined during training (using the mean and standard deviation of the ImageNet dataset: mean [0.485, 0.456, 0.406], standard deviation [0.229, 0.224, 0.225]).

[0120] The preprocessed image is input into the trained multi-label recognition model, and the model outputs the probability of each label. p c Read the judgment threshold corresponding to each label from the threshold file. Generate the judgment result:

[0121] In this context, 1 indicates a positive result (the presence of the oral disease or characteristic corresponding to the label) and 0 indicates a negative result (the absence of the label).

[0122] When generating the recognition report in this embodiment, the report content includes: the input image number or file name; the probability value of each label (periodontitis, tartar, tooth surface pigmentation, gingival bleeding, gingival redness and swelling, gingival recession, edentulous gaps, impacted teeth, etc.); the judgment threshold for each label; a list of positive labels; an optional soft-assignment heatmap of the region (showing the spatial response of each region to the image); a region-symptom evidence matrix heatmap (showing the evidence strength of each region in each symptom dimension); and prompts for doctors to review.

[0123] This embodiment performs a final performance evaluation of the trained model on a test set. The test set is not involved in the model's training process or the determination of the decision threshold; therefore, the performance metrics on the test set can objectively reflect the model's generalization ability. Evaluation metrics may include accuracy, precision, recall, F1 score, etc.

[0124] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-label recognition method for oral digital photographs based on regional symptom evidence, characterized in that, This includes the training and recognition steps of the multi-label recognition model; The training steps for the multi-label recognition model include: Step S11: Obtain multiple digital oral photos and generate a multi-label vector for each digital oral photo. Divide the multiple digital oral photos into a training set, a validation set, and a test set. The multi-label vector includes periodontitis labels and / or oral feature labels. Step S12: Preprocess the digital images of the oral cavity; Step S13: Input the preprocessed oral digital photo into a visual word encoder. The visual word encoder is used to encode the input oral digital photo into an image word sequence composed of multiple image words, each image word corresponding to a local area in the oral digital photo. Step S14: Input the image word sequence into the implicit oral anatomy region soft routing module. The implicit oral anatomy region soft routing module generates the allocation weight of each image word belonging to each preset oral anatomy region, obtains the region soft allocation matrix, and aggregates multiple image words into multiple oral anatomy region words according to the region soft allocation matrix; wherein, the rows of the region soft allocation matrix correspond to image words, and the columns correspond to oral anatomy regions. Step S15: For each oral anatomical region terminology, obtain the health reference terminology from the health status prototype library corresponding to that oral anatomical region, and calculate the pathological residual terminology between the oral anatomical region terminology and the health reference terminology. Step S16: Convert the pathological residual lexical units into a region-symptom evidence matrix through a learnable linear mapping, wherein the rows of the region-symptom evidence matrix correspond to the oral anatomical regions and the columns correspond to the symptom dimensions. Step S17: Set learnable region weights and symptom weights for each label to be identified, and perform factorized aggregation based on the region weights, symptom weights, and region-symptom evidence matrix to obtain the log odds of each label to be identified. Step S18: Obtain the probability of each label to be identified based on the log odds using the sigmoid function, train the model using the multi-label binary classification loss function, and determine the judgment threshold for each label on the validation set. The identification steps include: Step S21: Process the digital oral cavity photograph to be identified using the same preprocessing method as in the training phase, and then input it into the trained multi-label recognition model; Step S22: The multi-label recognition model outputs the probability of each label, compares the probability of each label with the corresponding label judgment threshold, and obtains the positive or negative judgment result of each label; where positive indicates the existence of the oral disease or feature corresponding to the label, and negative indicates the absence of the label. Step S23: Generate a multi-label recognition report based on the judgment results.

2. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 1, characterized in that, The oral cavity feature labels include one or more of the following: dental calculus, tooth surface pigmentation, gingival bleeding, gingival redness and swelling, gingival recession, edentulous gaps, and impacted teeth.

3. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 1, characterized in that, In step S11, when generating the multi-label vector, for the same image from multiple source directories, deduplication and merging are performed by calculating the content hash value or reading the file name identifier, and multiple positive labels corresponding to the image are merged into the same multi-label vector, and then the training set, validation set and test set are divided; wherein, the positive label index annotation result is a label that corresponds to an oral disease or oral feature.

4. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 1, characterized in that, The preprocessing in step S12 includes input size determination, proportional scaling, and gray fill. The input size determination includes: calculating the average width and average height of the oral digital photos in the training set, multiplying the average width and average height by the input scaling factor to obtain the recommended width and height, scaling down proportionally when the long side of the recommended width and height is greater than the preset maximum long side, and rounding the scaled-down width and height to an integer multiple of the size divisor, wherein the size divisor matches the downsampling factor of the visual word encoder. The proportional scaling and gray fill include: calculating the proportional scaling factor based on the model input size and the original image size, scaling the original image and centering it before pasting it onto the target canvas with a neutral gray background.

5. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 1, characterized in that, The implementation method of the implicit oral anatomy region soft routing module in step S14 is as follows: Let the image word sequence be ,in, T It is a three-dimensional real tensor consisting of the batch size multiplied by the number of image terms multiplied by the feature dimension. B For batch size, L The number of image terms, D Let be the feature dimension; and let the preset number of implicit oral anatomical regions be . R The soft routing module contains a multilayer perceptron, which maps the features of each image word to... R The log-odds allocation of the region is calculated, and then normalized along the image lexical dimension to obtain the soft allocation matrix of the region. ,in, A b,l,r Indicates the first b The first sample l The image term belongs to the first... r The probability of a hidden region; No. r Oral anatomical region terminology The calculation formula is: ; in, ϵ To prevent positive constants with a denominator of zero, Indicates the first b The first sample l Feature vectors of image words, l ∈{1, 2, … , L }, r∈{1 , 2 , … , R}.

6. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 5, characterized in that, The preset number of implicit oral anatomical regions is 6, corresponding to the gingival margin, interdental papilla, attached gingiva or gingival body, cervical region, crown region, oral mucosa or background; the training process of the region soft allocation matrix also includes region allocation balancing loss. L region This loss is used to suppress the concentration of the soft allocation matrix in a single region, and the specific formula is as follows: ; in, q r The average weight of the r-th region in the soft-assignment matrix of all regions for all words in all samples within a training batch is calculated using the following formula: ; Where B is the batch size and L is the number of image terms.

7. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 5, characterized in that, The specific implementation method of step S15 is as follows: Let the first r The health status prototype library corresponding to each oral anatomical region contains P A prototype of a health state For regional lexical units ,calculate U r The similarity to each healthy prototype is calculated using cosine similarity: ; in, The second norm of a vector; H r,p Indicates the first r The first region p A prototype of a healthy state p ∈{1, 2, …, P }; By temperature parameters The similarity is scaled and normalized to obtain the prototype weight: ; in 𝑤 𝑟,𝑝 This represents the weight coefficient of the nth health state prototype in the nth region; H r,q Indicates the first r The prototype of the q-th health state in region . q ∈{1 , 2 , …, P }; Health reference term 𝑁 𝑟 The weighted average of the various health prototypes: ; Then calculate the pathological residuals. Δ r : ; Where Norm(⋅) denotes L2 normalization, that is, for any input vector 𝑣, .

8. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 7, characterized in that, The prototypes in the health state prototype library are learnable parameters, and the loss is distributed through health prototypes during training. 𝐿 proto To suppress overlap between different prototypes within the same region, the formula for calculating the dispersion loss of the healthy prototype is: ; Here, cos(⋅,⋅) represents the cosine similarity between two vectors.

9. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 7, characterized in that, The region-symptom evidence matrix in step S16 is generated as follows: Let the preset number of symptom dimensions be . M Set a learnable linear mapping matrix. and bias vector ; For the r Pathological residuals in each region Calculate the symptom evidence vector: ; in, For the first r Each region in the whole M Evidence vectors in each of the symptom dimensions; σ(⋅) is the element-wise sigmoid function; stacking the symptom evidence vectors of all 𝑅 regions yields the region-symptom evidence matrix. .

10. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 9, characterized in that, The number of symptom dimensions is 8, corresponding to redness, swelling, recession or root surface exposure, calculus or plaque deposition, bleeding tendency, ulceration or damage, abnormal color, and abnormal texture or edges.

11. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 9, characterized in that, The specific implementation method of factorization aggregation in step S17 is as follows: Assume there is a total C One tag to be identified, c ∈{1 , 2 , … , C } is the label index; for each label, set a learnable region weight vector and a symptom weight vector, and denote the region weight of the c-th label as . Symptom weighting ; The region weights for each label are subjected to softmax normalization, and the symptom weights are also subjected to softmax normalization: ; in, Indicates the first c After normalization of the label, the label is at the... r Weights in each region Indicates the first c After normalization of the label, the label is at the... m Weights on each symptom dimension; Indicates the first c The first label in the r Unnormalized region weights in each region Indicates the first c The first label in the m Unnormalized symptom weights across symptom dimensions; r ′ is the variable for summing the region. Indicates the first c The first label in the r Unnormalized region weights over ′ regions; m ′ represents the summation variable for the symptom dimension. Indicates the first c The first label in the m Unnormalized symptom weights across ' symptom dimensions; Log-odds of the nth label z c The calculation formula is: ; in, b c This is the bias term for the nth label; The probability of the c-th label p c for: 。 12. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 11, characterized in that, The model training step also includes a label relation residual calibration module, which receives the initial log-odds vector output by the factorization aggregation module. Through a learnable label relationship matrix The sigmoid activation function generates the residual correction vector r. res Then add this to the initial logarithmic odds to get the final logarithmic odds z: ; in, This is the bias vector.

13. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 1, characterized in that, The total loss function used in the model training step is: ; in, For multi-label binary classification, cross-entropy loss, L region To distribute the losses evenly across regions, L proto To mitigate losses in healthy prototypes, λ region and λ proto To balance the weighting coefficients of each loss term.

14. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 13, characterized in that, The multi-label binary classification cross-entropy loss Using a form with positive sample weights: ; in, B For batch size, C For the total number of tags, y b,c For the first b The first sample c The true value of each label p b,c The probability predicted by the model. α c For the first c The weight of each label's positive sample is calculated using the following formula: ; in, N pos,c and N neg,c The training set c The number of positive and negative samples for each label. β For smoothing coefficients, α min and α max These represent the lower and upper limits of the weight.

15. The method for multi-label recognition of oral digital photographs based on regional symptom evidence according to claim 1, characterized in that, The method for searching and determining the threshold for each label on the validation set in step S18 is as follows: Let the first c The predicted probability of each label on the validation set is: The real label is Within the preset threshold range Inner step size Δ t Enumerate candidate thresholds t ; For each candidate threshold t Calculate the first c The F1 score for each label is: ; Among them, Precision c ( t ) and Recall c ( t ) are respectively at the threshold t Precision and recall rates; Selecting F1 c ( t The largest threshold is used as the first c Optimal threshold for each label .

16. A multi-label recognition system for oral digital photographs based on regional symptom evidence, characterized in that, include: The sample construction module is used to acquire digital oral photos and generate multi-label vectors for each digital oral photo; wherein, the multi-label vectors include periodontitis labels and / or oral feature labels; The input processing module is used to preprocess digital dental photographs; A visual lexical encoder is used to encode a preprocessed digital oral image into a sequence of image lexical units, where each image lexical unit corresponds to a local region in the digital oral image. The implicit oral anatomy region soft routing module is used to generate the allocation weight of each image word belonging to each preset oral anatomy region, obtain the region soft allocation matrix, and aggregate multiple image words into multiple oral anatomy region words based on the region soft allocation matrix; wherein, the rows of the region soft allocation matrix correspond to image words, and the columns correspond to oral anatomy regions; the health prototype difference module is used to obtain health reference words from the health state prototype library corresponding to the oral anatomy region for each oral anatomy region word, and calculate the pathological residual words between the oral anatomy region word and the health reference words; The system comprises the following modules: an evidence generation module, which converts pathological residual terms into a region-symptom evidence matrix using a learnable linear mapping, where rows correspond to oral anatomical regions and columns correspond to symptom dimensions; a factorization aggregation module, which sets learnable region and symptom weights for each label to be identified, performs factorization aggregation based on the region weights, symptom weights, and the region-symptom evidence matrix to obtain the log odds of each label, and uses a sigmoid function to obtain the probability of each label; a threshold determination module, which determines a judgment threshold for each label on a validation set partitioned from the training dataset; and a report generation module, which processes the digital oral images to be identified using the same preprocessing method as in the training phase, outputs the probability of each label using the trained model, compares the probability of each label with the corresponding judgment threshold to obtain a positive or negative judgment result for each label, and generates a multi-label recognition report based on the judgment result; where a positive result indicates the presence of the oral disease or feature corresponding to the label, and a negative result indicates its absence.

Citation Information

Patent Citations

  • Household auxiliary diagnosis system based on oral cavity scanning and disease prediction

    CN121331348A

  • Tooth health preliminary screening method and device based on image recognition, equipment and medium

    CN121504870A