Predicting bacterial vaginosis development using artificial neural networks

WO2026207454A1PCT designated stage Publication Date: 2026-10-01BOARD OF SUPERVISORS OF LOUISIANA STATE UNIV & AGRI & MECHANICAL COLLEGE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/021303
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-03-27
Publication Date
2026-10-01

Smart Images

  • Figure US2026021303_01102026_PF_FP_ABST
    Figure US2026021303_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method for predicting development of bacterial vaginosis, wherein a vaginal sample of the subject is used to detect levels of a plurality or organisms in the subject using a sequencing analysis; applying an artificial neural network model to determine a probability that the subject develops bacterial vaginosis over a first period of time using information of the subject's detected levels; treating the subject for bacterial vaginosis with an antibiotic when the probability is greater than a threshold value.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Docket No. 2932719-000286-W01

[0002] Filed: March 27, 2026 PREDICTING BACTERIAL VAGINOSIS DEVELOPMENT USING ARTIFICIAL NEURAL NETWORKS RELATED APPLICATIONS

[0003] This application claims the benefit of US Application No. 63 / 778,989, filed March 27, 2025.

[0004] TECHNICAL FIELD

[0005] The present disclosure relates to a method for predicting development of bacterial vaginosis, under an embodiment.

[0006] INCORPORATION BY REFERENCE

[0007] Each patent, patent application, and / or publication mentioned in this specification is herein incorporated by reference in its entirety to the same extent as if each individual patent, patent application, and / or publication was specifically and individually indicated to be incorporated by reference.

[0008] BRIEF DESCRIPTION OF THE FIGURES

[0009] Figure 1A shows sample collection and preparation under an embodiment.

[0010] Figure IB shows ANN training and application, under an embodiment.

[0011] Figures 2A-2D show the predictive accuracy of the ANN, under an embodiment.

[0012] Figures 3A and 3B show the predictive accuracy of the ANN, under an embodiment. Figure 4A-4B shows the importance of features in their predictive contributions to the ANN, under an embodiment.

[0013] Figure 5A-5D how a minimum feature analysis, under an embodiment.

[0014] Figures 6A-6E illustrate minimum feature accuracy and loss curves, under an embodiment.

[0015] Figure 7A-7E show ShAP analysis of minimum feature models, under an embodiment.

[0016] Figures 8A-8D show race specific relationships with BV, under an embodiment.

[0017] Figure 9A-9B show Accuracy and Loss Curves of Race Specific ANN Models.

[0018] Figures 10A-10D show race specific modeling metrics AUD curves, under an embodiment.

[0019] Figures 11A-11E show accuracy and loss curves of black specific minimum features models, under an embodiment.

[0020] Figures 12A-12E show accuracy and loss curves of white specific minimum features models, under an embodiment.Docket No. 2932719-000286-W01

[0021] Filed: March 27, 2026 Figure 13A-13E show ShAP Analysis of Minimum Feature Models trained on Black Participants, under an embodiment.

[0022] Figures 14A-14E show ShAP Analysis of Minimum Feature Models trained on White Participants, under an embodiment.

[0023] Figure 15 shows model specific parameters, under an embodiment.

[0024] Figure 16A shows sample collection and preparation, under an embodiment.

[0025] Figure 16B shows ANN training and application, under an embodiment.

[0026] Figure 17 shows incident bacterial vaginosis cohorts included in the study, under and embodiment.

[0027] Figure 18A-18C shows a summary of ANN modeling performance trained using 16S relative abundance data of 20 common vaginal bacterial taxa, under an embodiment.

[0028] Figure 19A-19B shows 20-feature model training and validation curves, under an embodiment.

[0029] Figure 20 shows 20-feature model performance predicting early and late pre-iBV, under an embodiment.

[0030] Figure 21A-21B shows feature performance and feature use determined from SHAP analysis, under an embodiment.

[0031] Figure 22A-22D shows N-feature models receiver operating curves and precision-recall curves, under an embodiment.

[0032] Figure 23A-23E shows N-feature model training and validation curves, under an embodiment.

[0033] Figure 24 shows taxa used to train race-specific feature subset models, under an embodiment.

[0034] Figure 25A-25E shows subset model performance stratified by cohort, under an embodiment.

[0035] Figure 26A-26E shows SHAP analysis of n-feature models, under an embodiment.

[0036] Figure 27A-27D shows classification performance and SHAP analysis of models trained on race-specific data, under an embodiment.

[0037] Figure 28A-28D shows receiver operating curves of race-specific n-feature models, under an embodiment.Docket No. 2932719-000286-W01

[0038] Filed: March 27, 2026 Figure 29A-29E shows SHAP analysis of models trained on top features set in Black participants.

[0039] Figure 30A-30E shows SHAP analysis of models trained on top features set in White participants.

[0040] Figure 31 shows model specific parameters, under an embodiment.

[0041] DETAILED DESCRIPTION

[0042] Despite considerable efforts to understand its pathogenesis, bacterial vaginosis (BV) remains the most common vaginal infection, affecting approximately 30% of reproductive-age women (1-3). Specific risk factors associated with an increased predisposition to BV have been identified, including colonization by high-risk organisms, certain sexual activities, and menstruation (4-6). However, the presence of these factors does not guarantee the development of BV. This suggests that additional, yet unidentified, factors are required for the development of BV.

[0043] The vaginal microbiota is a community of many different micro-organisms creating an intricate ecosystem. Recent advancements in computational power have facilitated the development of machine learning algorithms capable of analyzing large amounts of data (7). An artificial neural network (ANN) is one such supervised machine learning tool, designed to identify complex patterns in data. An ANN consists of an interconnected, feed-forward network of input, hidden, and output neurons (8,9). For microbiome analysis using ANNs, quantitative data generated through sequencing are used at the input layer, with separate neurons representing each bacterial taxonomic classification (7). These machine learning techniques facilitate the investigation of the complex vaginal microbiota by processing and interpreting vast amounts of microbial data. By applying ANNs to vaginal microbiome studies, researchers can analyze subtle changes in the microbiome, potentially uncovering dynamic changes associated with incidence of BV. This approach enables the identification of specific microbial signatures and patterns that may be critical in understanding and diagnosing BV.

[0044] In the case of incidence BV (iBV), our goal is to utilize ANNs and machine learning techniques to predict future iBV development through tracking the polymicrobial composition over time. We applied specialized ANNs to data from an ongoing prospective iBV pathogenesis study (10). While the parent study was designed to track and assess changes in the microbial communities over time, astonishingly we found the ANN modeling allowed us to accuratelyDocket No. 2932719-000286-W01

[0045] Filed: March 27, 2026 predict from individual samples whether the participant would develop iBV. From our study we have developed models that are capable of detecting iBV 2-weeks prior to onset. Through analysis of the model, we have gained additional insight into the complex microbial signatures contributing to iBV onset.

[0046] RESULTS

[0047] Clinical sampling and modeling workflow

[0048] For this study, we analyzed data from 8 iBV cases and 8 comparable healthy controls from an ongoing parent study (31). Table 1 displays the demographic data for these 16 participants. Vaginal samples were Nugent scored for iBV diagnostic purposes and sequenced using 16S rRNA gene sequencing to characterize the longitudinal dynamics of the vaginal microbiome (Figure 1A).

[0049] Taxonomic classification identified 20 vaginal organisms, listed in general order of relative abundance across all specimens: L. iners, L. cri spatus, L. jensenii, Gardnerella, L. gasseri, Prevotella, Lactobacillus, Streptococcus, Aerococcus, Sneathia, Dialister, BVAB1, Megasphaera, Fannyhessea, Fastidiosipila, Mobiluncus, Parvimonas, Limosilactobacillus, L. intestinalis, and Ligilactobacillus.

[0050] Sequencing was conducted in conjunction with qPCR to calculate inferred absolute abundance (IAA) of vaginal taxa (11). IAA is calculated as follows:

[0051] / A4 or qani ,sm. ( i \ - IhsrRN A gen

[0052]

[0053] svvni; -e cop - -ied=ora

[0054] / &anism(%) *

[0055] This data was used to train artificial neural network (ANN) models to predict whether a sample was collected within 14 days of the onset of iBV (Figure IB). In total, we analyzed 495 vaginal samples: 212 from the 8 iBV cases (137 pre-iBV and 60 during iBV) and 283 from the 8 comparable controls. Our final trained model was designed to predict whether a participant would develop iBV or remain healthy using a single vaginal specimen. The model demonstrated high accuracy, sensitivity, and specificity in distinguishing at-risk individuals from healthy controls.

[0056] Predictive Modeling for Early Detection of iBV

[0057] To assess if the vaginal microbiome could aid in early detection of iBV we initially trained ANN models using the 20 identified vaginal taxa in our cohort. Following parameter optimization, the final architecture consisted of 2 layers. Hidden layer 1 contained 23 nodes and hidden layer 2 contained 28 hidden nodes. The final selected model achieved a peak accuracy of 100% onDocket No. 2932719-000286-W01

[0058] Filed: March 27, 2026 training data and 97.7% on testing data. Tn addition to attaining high accuracy, the 20-feature model also achieved high sensitivity and specificity (Figure 2A-2D, Table 2, Figure 3A-3B).

[0059] Table 1

[0060] Characteristic Total N=16 iBV cases (n=8) Controls (n=8) p-value Age (years) 0.798 Median (QI. Q3) 29.2 ± 8.4 29.8 ± 8.6 28.6 ± 8.7

[0061] Race (self-identified) 0.285 White 8 (50.0) 3 (37.5) 5 (62.5)

[0062] African American 6 (37.5) 3 (37.5) 3 (37.5)

[0063] Asian 2 (12.5) 2 (25.0) 0 (0.00)

[0064] Ethnicity (self-identified) 1.000 Hispanic 1 (6.25) 0 (0.00) 1 (12.5)

[0065] Non-Hispanic 15 (93.75) 8 (100.0) 7 (87.5)

[0066] Education 0.219 Any post-graduate studies 5 (31.2) 4 (50.0) 1 (12.5)

[0067] Bachelor’s degree 3 (18.8) 1 (12.5) 2 (25.0)

[0068] Some college / Associate’s degree 7 (43.8) 2 (25.0) 5 (62.5)

[0069] High school / GED 1 (6.2) 1 (12.5) 0 (0.00)

[0070] Less than high school 0 (0.00) 0 (0.00) 0 (0.00)

[0071] Current Smoker 1.000 Yes 0 (00.0) 0 (0.00) 0 (0.00)

[0072] No 16 (100.0) 8 (100.0) 8 (100.0)

[0073] Douched within the Past 3 Months 1.000 0 (0.00) 0 (0.00) 0 (0.00)

[0074] STI History* 0.842 >2 STIs 4 (25.0) 2 (25.0) 2 (25.0)

[0075] 1 STI 5 (31.2) 3 (37.5) 2 (25.0)

[0076] None 7 (43.8) 3 (37.5) 4 (50.0)

[0077] History of BV 1.000 Yes 6 (37.5) 3 (37.5) 3 (37.5)

[0078] No 10 (62.5) 5 (62.5) 5 (62.5)

[0079] Current Contraception Use 0.911 Hormonal Contraception# 10 (62.5) 5 (62.5) 5 (62.5)

[0080] Birth Control Pills 3 (18.8) 2 (25.0) 1 (12.5)

[0081] Hormonal IUD 5 (31.2) 2 (25.0) 3 (37.5)

[0082] Implants 2 (12.5) 1 (12.5) 1 (12.5)

[0083] Non-Hormonal Contraceptions 0 (0.00) 0 (0.00) 0 (0.00)

[0084] Copper IUD 0 (0.00) 0 (0.00) 0 (0.00)

[0085] None 6 (37.5) 3 (37.5) 3 (37.5)

[0086] History of Contraception Use& 1.000 Hormonal Contraception# 5 (31.2) 2 (25.0) 3 (37.5)

[0087] Non-Hormonal Contraceptions 0 (0.00) 0 (0.00) 0 (0.00)

[0088]

[0089] None 1 (6.20) 1 (12.5) 0 (0.00)

[0090] Role of Microbial Features in Pre-iBV Classification

[0091] SHAP analysis was performed on the 20-feature model to assess how each taxa (feature) contributed to model predictions (Figure 4A-4B). This analysis revealed that vaginal Lactobacillus species were found to be the most important for model accuracy (12). Specifically,Docket No. 2932719-000286-W01

[0092] Filed: March 27, 2026 high IAA of L. crispatus and . jensenii ere strongly associated with the model predicting a participant would remain healthy. Similarly, lower levels of L. iners contributed to healthy predictions. Interestingly, low IAA of L. gasseri favored a healthy classification, whereas high IAA of this organism contributed to samples being classified as pre-iBV. Higher abundances of many BV-associated bacteria such as Gardnerella Sneathia, BVAB1, and Aerococcus were found to contribute predictions toward the pre-iBV category. These insights demonstrate how the ANN leverages specific microbial features, particularly the balance between the Lactobacillus species and the BV-associated bacteria to accurately distinguish between healthy and pre-iBV states. Table 2 Model Accuracy Metrics

[0093] Metric Train Test

[0094]

[0095] Accuracy 1.00 0.976

[0096] Sensitivity 1.00 0.96

[0097] Specificity 1.00 0.98

[0098] Fl Score 1.00 0.96

[0099] AUC 1.00 0.99

[0100] PPV 1.00 0.96

[0101] NPV 1.00 0.98

[0102]

[0103] Cohen Kappa 1.00 0.94

[0104] Table 2. ANN Model Performance. ANN Model Accuracy Metrics. The highest preforming epoch was selected to assess model accuracy. (306)

[0105] Minimum Feature Analysis

[0106] To evaluate the number of vaginal taxa required to attain accurate iBV predictions, we trained models using different numbers of top taxa (features). These features were selected based on their importance as determined by mean absolute SHAP values. Models were built using the top 3, 5, 7, 9, and 12 features (Figure 5A-5D, Figure 6A-6E, Table 3). The final model trained on the top 3 features, (L. jensenii, L. crispatus, and L. gasseri) had a peak training accuracy of 91.0% and testing accuracy of 94.0%. In training and testing data the three-feature model had decreased sensitivity (training data = 0.819, testing data = 0.884) compared to specificity (training data = 0.955, testing data =0.965). This model struggled to accurately classify pre-iBV samples, as indicated by decreased sensitivity. When trained using the top 7, 9 and 12 features, these models achieved peak training accuracies > 99%, demonstrating improved performance with the inclusion of additional features. These findings indicate that while only a few key vaginal taxa may be sufficient for accurate iBV prediction, incorporating more features enhances the model's sensitivityDocket No. 2932719-000286-W01

[0107] Filed: March 27, 2026 in detecting pre-iBV samples. This type of analysis is paramount for consideration of potential development of clinical diagnostic testing that could predict future development of iBV.

[0108] Top 3 Top 5 Top 7 Top 9 Top 12 Metric Train Test Train Test Train Test

[0109] Accuracy 0.910 0.940 0973 0.976 0.997 0.976 0.991 0.988 1.00 0.988 Sensitivity 0.819 0.884 0945 0.961 1.00 0.923 0.981 0.961 1.00 0.961 Specificity 0.955 0.965 0986 0.982 0.995 1.00 0.995 1.00 1.00 1.00 Fl_Score 0.858 0.901 0958 0.961 0.995 0.960 0.986 0.980 1.00 0.980 AUC 0.968 0.954 0997 0.981 0.999 0.996 0.999 0.994 1.00 0.998 PPV 0.900 0.920 0972 0.961 0.991 1.00 0.990 1.00 1.00 1.00 NPV 0 914 0949 0973 0 982 1 00 0 966 0 991 0983 1 00 0 983

[0110]

[0111] Cohen Kappa

[0112] Table 3. Performance of Minimum Feature ANN Models. Performance metrics of ANN Models trained following feature selection.

[0113] SHAP analysis was applied to determine how features were utilized by the n-feature models (n = top 3, 5, 7, 9, and 12 features, Figure 7A-7E). Across all n-feature models, Lactobacillus species consistently emerged as the most important taxa, similar to findings from the 20-feature model. In models where Gardnerella was included as a feature, Gardnerella ranked as the 4th most important feature, following the Lactobacillus species. Like the 20-feature model, high IAA of Gardnerella, BVAB1, and Sneathia contributed to classifying samples as pre-iBV, while lower IAA levels favored classification as healthy. These results emphasize that key microbial features, particularly Lactobacillus species and Gardnerella, consistently drive model predictions across varying feature sets.

[0114] Race-Specific Models: Performance and Feature Utilization

[0115] Considering the hypothesis that the onset of iBV may differ based on race, we trained models separately on Black (B-Model) and White (W -Model) participants (6). Both the Models trained, B-Model and W-Model, achieved 100% accuracy on both training and testing data (Figures 8A-12E and Table 4).

[0116] Black Participants White Participants

[0117] Metric

[0118] Accuracy 1 1 i i Sensitivity 1 1 i i Specificity 1 1 i i

[0119] Fl Score 1 1 i i

[0120] AUC 1 1 i i

[0121] PPV 1 1 i i

[0122] NPV 1 1 i i

[0123]

[0124] Cohen Kappa 1 1 i iDocket No. 2932719-000286-W01

[0125] Filed: March 27, 2026

[0126] Table 4. Performance of Race specific Models. ANN Model Accuracy Metrics from models trained on cohort subset based on participants race using IAA of 20 vaginal taxa.

[0127] The ranking of features differed between the two models. While in both models, Lactobacillus species ranked as the top 3 most important features, the specific species and how the species contributed to model predictions differed between the two models. In the B-Model, L. gasseri was the most important feature and higher abundance contributed to pre-iBV predictions. Whereas in the W -Model, jensenii was the most important feature. Additionally, high abundance of L. iners was associated with pre-iBV predictions only in the W-model. In both models, high relative IAA of BV-associated bacteria including Prevotella, BVAB1, Gardnerella, and Sneathia were found to be predictive of pre-iBV, and low relative IAA were predictive of healthy. These findings indicate that while overall model performance is excellent in both race cohorts, the role of L. iners and L. gasseri in the onset of iBV may differ depending on participants’ race. This highlights the potential population-specific differences in iBV onset.

[0128] To determine the minimum number of taxa (features) required to accurately classify a sample as pre-iBV or healthy for race-specific models, similarly, models were trained using the top 3, 5, 7, and 12 most important features. Both the 3-feature B and W models were highly accurate. 3-feature models accurately categorized 97.1% of black participants and 98.3% of white participants in the training data and 100% in the testing data (Figures 11A-11E and 12A-12E and Table 5). Race specific models achieved 100% accuracy in training and testing data when trained on a minimum of 5 features in Black participants and 7 in White participants.

[0129] A. Black-Subset Models

[0130] Top 3 Top 5 Top 7 Top 9 Top 12 Metric Train Test Train Test Train Test Train Test Train Test Accuracy 0.971 1.00 1 00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 Sensitivity 0.938 1.00 1 00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 Specificity 0.989 1.00 1 00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 Fl_Score 0.958 1.00 1 00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 AUC 0.996 1.00 1 00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 PPV 0.978 1.00 1 00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 NPV 0.967 1.00 1 00 1.00 1.00 1.00 1.00 1.00 1.00 1.00

[0131]

[0132] Cohen Kappa

[0133] B. White-Subset Models

[0134] Top 3 Top 5 Top 7 Top 9 Top 12 Metric Train Test Train Test Train Test Train Test Train Test Accuracy 0.983 1.00 0989 1.00 1.00 1.00 1.00 1.00 1.00 1.00 Sensitivity 0.981 1.00 0981 1.00 1.00 1.00 1.00 1.00 1.00 1.00 Specificity 0.984 1.00 0992 1.00 1.00 1.00 1.00 1.00 1.00 1.00

[0135]

[0136] Fl_Score 0.972 1.00 0981 1.00 1.00 1.00 1.00 1.00 1.00 1.00Docket No. 2932719-000286-W01

[0137] Filed: March 27, 2026 AUC 0.997 1.00 0998 1.00 1.00 1.00 1.00 1.00 1.00 1.00 PPV 0.964 1.00 0981 1.00 1.00 1.00 1.00 1.00 1.00 1.00 NPV 0 992 1 00 0992 1 00 1 00 1 00 1 00 1 00 1 00 1 00

[0138]

[0139] Cohen Kappa

[0140] Table 5. Performance of Minimum Feature, Race Specific ANN Models. Performance metrics of ANN Models trained separately for (A) black and (B) white participants using IAA of 3 - 12 of the most import features found to predict preiBV.

[0141] To further assess the contribution of each feature to model accuracy, SHAP analysis was performed on each race-specific model trained on subsets of the most important features (Figures 13A-13E and 14A-14D) Across all models, independent of race or feature subset, Lactobacillus species consistently ranked among the most important for model predictions. In both Black and White participant models, high IAA of L. crispatus and L. jensenii was predictive of healthy outcomes. Interestingly, in both race-specific models, high IAA of L. gasseri was predictive of pre-iBV, while low abundance predicted whether a sample would remain healthy. Also, L. gasseri was the most important feature for predicting pre-iBV / healthy status in Black participants across all model subsets, but it consistently ranked as one of the least important features in models trained on White participants. In all models, high relative IAA of typical BV-associated bacteria including Gardnerella, Sneathia, and BVAB1 was associated with predictions of pre-iBV. These findings suggest that training models separately on Black and White participants allows the identification of race-specific patterns that enhance classification accuracy.

[0142] DISCUSSION

[0143] In this study, we developed an artificial neural network (ANN) capable of accurately predicting the onset of incident bacterial vaginosis (iBV) up to 14 days in advance from a single vaginal sample. Our model achieved high accuracy, sensitivity, and specificity using inferred absolute abundances (IAA) of vaginal taxa to make predictions (11). This novel predictive capability enables early detection of iBV, allowing for proactive intervention and potentially altering treatment strategies, particularly for high-risk or recurrent cases.

[0144] By analyzing our model to understand how it accurately classified samples, we revealed novel microbial patterns correlating with pre-iBV states. From SHAP analysis, we found common vaginal Lactobacillus species were the most important taxa to accurately classify samples. As we expected, L. crispatus and L. jensenii were indicative of stable vaginal communities while genera containing BV-associated bacteria such as Gardnerella, Sneathia, and BVAB1 were indicative ofDocket No. 2932719-000286-W01

[0145] Filed: March 27, 2026 pre-iBV communities (6,12,13). Somewhat to our surprise, L. gasseri abundance positively correlated with pre-iBV predictions, although prior groups have shown L. gasseri strains produce less L-lactic acid and hydrogen peroxide, compared to L. crispatus (14,15). The relationship between L. gasseri and iBV may be poorly understood due to prior work primarily relying on compositional approaches to quantify vaginal microflora (12) . While SHAP analysis assesses correlative relationships between features and predictions our findings suggest additional emphasis should be placed on further assessing the role of L. gasseri in iBV, using non-compositional methods to measure L. gasseri abundance (ie qPCR / IAA). Furthermore, our findings reflect that L. crispatus is critical to iBV protection, while less common species such as L. gasseri and L. jensemi may vary in their protective capacity due to strain variation and hosts' features, including their racial / ethnic background.

[0146] Minimum feature analysis revealed that only a few taxa are necessary for highly accurate predictions. Models trained on the top 3 features (L. jensenii, L. crispatus, and L. gasseri) achieved >90% accuracy. While additional features improved accuracy, the strong performance of minimal models suggests that simplified approaches to characterizing the vaginal microbiome may be sufficient for specific research and diagnostic purposes. For example, using targeted qPCR panels on these three organisms could be feasible alternatives to complex microbial profiling for early iBV detection.

[0147] Models developed from race-specific subgroups improved model accuracy and revealed distinct microbial signatures contributing to pre-iBV predictions in Black and White participants. L. gasseri was found to be highly important to pre-iBV predictions in models trained on the Black subgroup, white this was not reflected in models trained on the White subgroup. Additionally, L. iners was a key predictor of pre-iBV in white-specific models but ranked low in importance for predicting pre-iBV in black participants. These findings support prior work suggesting vaginal microbiome stability differs based on race (6). The implications of these race-specific differences are profound and suggest that personalized diagnostic approaches may be necessary for the early detection of iBV. In future work, we hope to validate these findings using larger cohorts.

[0148] Early detection of iBV could have significant clinical implications, allowing for early intervention and improved clinical advice that may take the form of prophylactic antibiotics, probiotics, and behavioral modifications. Furthermore, high accuracy observed on race-specificDocket No. 2932719-000286-W01

[0149] Filed: March 27, 2026 models underscores the importance of considering demographic and biological diversity in developing microbial models and potentially in clinical practice.

[0150] While we have rigorously conducted our analysis and added valuable insight into early detection of iBV, several limitations are associated with our study. Although our data is rich in longitudinal samples our cohort is relatively small. Future studies externally validating models with large cohorts from multiple studies will be required to verify model validity. Furthermore, SHAP analysis identifies correlations between model predictions and microbial taxa, future studies will be required to verify relationships between vaginal microbiota and pre-iBV classification. Similarly, future studies will be required to identify the underlying biological mechanism mediating the microbial relationships found through SHAP analysis. Although our model was trained on a relatively small cohort, the high accuracy achieved by our model reflects the strong biological patterns underlying the pre-iBV state. The findings warrant additional investigation, with future studies dedicated to validating and expanding upon these findings in larger cohorts. Conclusion

[0151] Our study has demonstrated that machine-learning approaches are highly effective in detecting iBV prior to clinical onset, which has the potential to be applied to early diagnosis of iBV from a single vaginal sample. Our findings emphasize the role of Lactobacillus species play in regulating vaginal microbiome dynamics and emphasize race-specific differences in microbial predictors of pre-iBV states. Integrating personalized microbiome-based diagnostics capable of early detection of vaginal disorders could revolutionize the management of bacterial vaginosis and other vaginal health disorders.

[0152] Methods

[0153] Clinical Enrollment

[0154] In a prospective, longitudinal study, we enrolled an ethnically diverse group of cisgendered women who have sex with men (WSM) from the UAB Sexual Health Research Clinic (UAB IRB-300004547). The enrollment criteria are described in the published protocol (10). In brief, cisgendered women were enrolled following the criteria, aged 18-45 with a current male sexual partner, no antibiotic usage within the last 14 days, no HIV infection, and not pregnant. In the clinic, vaginal swabs were collected for Amsel criteria, participants must have no positive criteria, and Nugent score, participants cannot have any Gardnerella morphotypes (16,17). ParticipantsDocket No. 2932719-000286-W01

[0155] Filed: March 27, 2026 must also be negative on STI screening for Trichomonas vaginalis Chlamydia trachomatis, Neisseria gonorrhoeae, and Mycoplasma genitalium by nucleic acid amplification testing (10).

[0156] The study design is similar to our previous study (13), with the addition of participants self-collected tw ice-daily vaginal specimens for 60 days. Along with specimen collection, these women kept detailed diaries of daily sexual and physiologic activities (e g., menses, douching, and sexual partners). Specimens were collected at home and delivered to the clinic weekly to the study site to be Gram stained, Nugent scored, and preserved at -80°C.

[0157] Vaginal specimens for microbiome characterization

[0158] Participants were being followed for iBV development (Nugent score of 7-10 on >4 consecutive specimens) based on their twice-daily, self-collected vaginal specimens. Specimens were selected from the day of and 14 days prior to iBV (30 total), as well as up to 5 samples post iBV (17). Specimens were also selected from comparable healthy controls by matching at least three out of four categories: age, race, menstrual cycle, and birth control methods. All specimens were shipped to the Microbial Genomics Resource Group (MGRG) at LSUHSC. The MGRG isolated DNA from vaginal specimens, as described previously, using a modified QIAamp DNA Mini Kit (QIAGEN) (13,18).

[0159] Molecular Methods

[0160] As shown in Figure 1A, the isolated DNA from twice-daily vaginal specimens was used to perform 16S ribosomal RNA (rRNA) gene sequencing on vaginal specimens from women who develop iBV and comparable healthy controls to determine changes in the relative abundance of bacteria over time. Quantitative PCR (qPCR) for the universal 16S rRNA was used gene to determine bacterial burden and calculate inferred absolute abundance (IAA) concentrations of vaginal organisms, as outlined in Tettamanti Boshier et al. (11).

[0161] As in our previous study (13), the MGRG at LSUHSC performed 16S seq. First, PCR amplicon libraries spanning the taxonomically informative 4th hypervariable (V4) region of the 16S rRNA were generated and sequenced on the Illumina MiSeq platform (18). Sequencing data was analyzed in R using DADA2 vl.16 (19), with taxonomic classification using SILVA vl38. Further species classification of the Lactobacillus genus were performed using curated BLAST queries (20). We used the G:P:L mix standard as described previously to calculate the total bacterial load via broad range 16S rRNA gene qPCR within the vaginal specimens (21). @e calculated the IAA of vaginal bacterial in our specimens using the calculation: IAA (16S rRNADocket No. 2932719-000286-W01

[0162] Filed: March 27, 2026 gene copies / specimen) = Relative Abundance (%) * Total Bacterial Load (universal 16S rRNA gene copies / specimen) (11).

[0163] Modeling approach and parameters

[0164] An ANN is a type of statistical model similar to nonlinear regression models (22). Input for the model was the organisms measured through IAA. We used specimens from iBV cases and healthy controls as input. Using the study data collected from the twice daily vaginal specimens, we built and compared several ANNs using the software packages TensorFlow (v2.16.1) and Keras (v3.0) (23,24). First, we grouped the iBV cases and healthy control specimens into two groupings: “pre-iBV” from the days leading to the onset of iBV and “Healthy” from the women who did not develop iBV. The data was then split in which 80% was input into the network with the labels iBV and Healthy to train the ANN. The remaining 20% was used for validation, without these labels, to assess how well the network was trained. Also, each ANN was trained over 600 iterations (epochs) of the training-and-validation process. We compared several architectures, combinations of hidden layers, to achieve the most optimal accuracy. This range of input structures and model architectures is known to be optimal for similar omics datasets (25). Stochastic gradient descent was used as the optimization algorithm, and neurons used a Rectified Linear Unit (ReLU) activation function. Furthermore, a grid search approach was used to determine optimal number of hidden layers, node combinations, dropout percentages, batch size, and regularization to maximize the accuracy of predicting influential BVAB in the development of iBV. Additionally, we experimented with alternative optimizers and activation functions. Final models were selected based on training and validation accuracy. ANNs were implemented in Python3 in the visual studio code (VS code vl.96.2). Visualizations were created using R v4.2.1.

[0165] In certain embodiments, the predictive model comprises a feedforward artificial neural network configured for binary classification of bacterial vaginosis status. Input features include microbial taxa abundances derived from sequencing data, optionally transformed using relative abundance, centered log-ratio (CLR) transformation, or other normalization techniques.

[0166] The neural network may include an input layer corresponding to the number of selected features (e.g., top N taxa), followed by one or more fully connected hidden layers comprising approximately 16-128 neurons per layer. In some implementations, regularization techniques are applied, including Gaussian noise injection, dropout layers, and / or L1 / L2 weight penalties.Docket No. 2932719-000286-W01

[0167] Filed: March 27, 2026 The model is trained using a binary cross-entropy loss function and an optimization algorithm such as Adam, with a learning rate on the order of le-3 to le-5. Training may proceed for a fixed number of epochs or until convergence based on validation performance.

[0168] The output layer comprises a single neuron with a sigmoid activation function, producing a probability score corresponding to BV status. Classification may be performed using a predefined threshold (e.g., 0.5).

[0169] In certain embodiments, data are partitioned at the participant level to prevent leakage between training and testing datasets, ensuring that all samples from a given individual are assigned exclusively to a single partition.

[0170] In certain embodiments, regularization includes combined LI and L2 penalties applied to the model weights. The reported L1 / L2 value (e.g., 0.001) represents equal coefficients applied to both LI and L2 regularization terms. In general, these coefficients may be on the order of 104to I02, with both LI and L2 terms set to the same value.

[0171] Determine influential BVAB in the development of iBV

[0172] To account for repeated sampling within individuals, data partitioning was performed at the participant level such that all samples from a given individual were assigned exclusively to either the training or testing dataset. This approach mitigates within-subject correlation without explicitly modeling random effects. In some embodiments, repeated measures may alternatively be addressed using mixed-effects modeling approaches, including neural network architectures incorporating random effects.

[0173] Once the ANN was trained and validated, we deconstructed the model using SHapley Additive exPlanations, using the python package (SHAP v0.46.0) to determine the “relative importance” of specific organisms that are key for iBV determination and development (25). The data used to train each model was included as the background while testing data was used to compute SHAP values. For models trained on 100 or more samples, SHAP explainer background was limited to 100 training samples. This involved first evaluating baseline performance and then permuting features (organisms) by removing that bacterium and measuring the resulting change in accuracy. After, we calculated the difference in performance for each organism (or combination of organisms) and compared it to the baseline. A larger drop means the accuracy indicates this organism is more important for predictions. This change in accuracy is translated into theDocket No. 2932719-000286-W01

[0174] Filed: March 27, 2026 organism’s feature importance. Additionally, SHAP analysis provides the opportunity to assess how features contribute to model predictions for individual samples and participants.

[0175] Figure 15 shows ANN model specific parameters, under an embodiment.

[0176] REFERENCES

[0177] (1) Hillier S, Marrazzo J, Holmes K. Bacterial vaginosis. Bacterial vaginosis. In: Holmes KK, Sparling PF, Mardh P-A, et al., editors. 4th ed., New York, NY: McGraw-Hill: 2008, p. 737-68.

[0178] (2) Srinivasan S, Fredricks DN. The human vaginal bacterial biota and bacterial vaginosis.

[0179] Interdiscip Perspect Infect Dis 2008;2008:750479. https: / / doi.org / 10.! 155 / 2008 / 750479. (3) Koumans EH, Sternberg M, Bruce C, McQuillan G, Kendrick J, Sutton M, et al. The Prevalence of Bacterial Vaginosis in the United States, 2001-2004; Associations With Symptoms, Sexual Behaviors, and Reproductive Health. Sexually Transmitted Diseases 2007;34:864-9. https: / / doi.org / 10.1097 / OLQ.0b013e318074e565.

[0180] (4) Kenyon CR, Buyze J, Klebanoff M, Brotman RM. Association between bacterial vaginosis and partner concurrency: a longitudinal study. Sex Transm Infect 2018;94:75-7. https: / / doi.0rg / lO.l 136 / sextrans-2016-052652.

[0181] (5) Srinivasan S, Liu C, Mitchell CM, Fiedler TL, Thomas KK, Agnew KJ, et al. Temporal variability of human vaginal bacteria and relationship with bacterial vaginosis. PLoS One 2010;5:el0197. https: / / doi.org / 10.1371 / joumal.pone.0010197.

[0182] (6) Gajer P, Brotman RM, Bai G, Sakamoto J, Schutte UME, Zhong X, et al. Temporal dynamics of the human vaginal microbiota. Sci Transl Med 2012;4: 132ra52.

[0183] https : / / doi . org / 10.1126 / scitranslmed.3003605.

[0184] (7) Marcos-Zambrano LJ, Karaduzovic-Hadziabdic K, Loncar Turukalo T, Przymus P, Trajkovik V, Aasmets O, et al. Applications of Machine Learning in Human Microbiome Studies: A Review on Feature Selection, Biomarker Identification, Disease Prediction and Treatment. Front Microbiol 2021;12:634511. https: / / doi.org / 10.3389 / fmicb.2021.63451L (8) Ditzler G, Polikar R, Rosen G. Multi-Layer and Recursive Neural Networks for Metagenomic Classification. IEEE Trans Nanobioscience 2015;14:608-16.

[0185] https: / / doi.org / 10.1109 / TNB.2015.2461219.Docket No. 2932719-000286-W01

[0186] Filed: March 27, 2026 (9) Ali A. Artificial Neural Network (ANN). Medium 2019. https: / / medium.com / machine- Iearning-researcher / artificial-neural-network-ann-4481fa33d85a (accessed June 16, 2021). (10) Muzny CA, Elnaggar JH, Sousa LGV, Lima A, Aaron KJ, Eastlund IC, et al. Microbial interactions among Gardnerella , Prevotella and Fannyhessea prior to incident bacterial vaginosis: protocol for a prospective, observational study. BMJ Open 2024;14:e083516. https : / / doi . org / 10.1136 / bmj open-2023 -083516.

[0187] (11) Tettamanti Boshier FA, Srinivasan S, Lopez A, Hoffman NG, Proll S, Fredricks DN, et al.

[0188] Complementing 16S rRNA Gene Amplicon Sequencing with Total Bacterial Load To Infer Absolute Species Concentrations in the Vaginal Microbiome. mSystems 2020;5:e00777-19. https: / / doi.org / 10.! 128 / mSystems.00777-19.

[0189] (12) Amabebe E, Anumba DOC. The Vaginal Microenvironment: The Physiologic Role of Lactobacilli. Front Med (Lausanne) 2018;5 : 181. https: / / doi.org / 10.3389 / fmed.2018.00181. (13) Muzny CA, Blanchard E, Taylor CM, Aaron KJ, Talluri R, Griswold ME, et al.

[0190] Identification of Key Bacteria Involved in the Induction of Incident Bacterial Vaginosis: A Prospective Study. J Infect Dis 2018;218:966-78. https: / / doi.org / 10.1093 / infdis / jiy243. (14) Hutt P, Lapp E, Stsepetova J, Smidt I, Taelma H, Borovkova N, et al. Characterisation of probiotic properties in human vaginal lactobacilli strains. Microb Ecol Health Dis 2016;27: 10.3402 / mehd.v27.30484. https: / / doi.org / 10.3402 / mehd.v27.30484.

[0191] (15) Pan M, Hidalgo-Cantabrana C, Goh YJ, Sanozky -Dawes R, Barrangou R. Comparative Analysis of Lactobacillus gasseri and Lactobacillus crispatus Isolated From Human Urogenital and Gastrointestinal Tracts. Front Microbiol 2020; 10:3146. https: / / doi.org / 10.3389 / fmicb.2019.03146.

[0192] (16) Amsel R, Totten PA, Spiegel CA, Chen KC, Eschenbach D, Holmes KK. Nonspecific vaginitis. Diagnostic criteria and microbial and epidemiologic associations. Am J Med 1983;74:14-22. https: / / doi.org / 10.1016 / 0002-9343(83)91112-9.

[0193] (17) Nugent RP, Krohn MA, Hillier SL. Reliability of diagnosing bacterial vaginosis is improved by a standardized method of gram stain interpretation. J Clin Microbiol 1991;29:297-301. https: / / doi.Org / 10.1128 / JCM.29.2.297-301.1991.

[0194] (18) Van Der Pol WJ, Kumar R, Morrow CD, Blanchard EE, Taylor CM, Martin DH, et al. In Silico and Experimental Evaluation of Primer Sets for Species-Level Resolution of theDocket No. 2932719-000286-W01

[0195] Filed: March 27, 2026 Vaginal Microbiota Using 16S Ribosomal RNA Gene Sequencing. J Infect Dis 2019;219:305-14. https: / / doi.org / 10.1093 / infdis / jiy508.

[0196] (19) Callahan BJ, McMurdie PJ, Rosen MJ, Han AW, Johnson AJA, Holmes SP. DADA2:

[0197] High-resolution sample inference from Illumina amplicon data. Nat Methods 2016; 13 :581 - 3. https: / / doi.org / 10.1038 / nmeth.3869.

[0198] (20) Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ. Basic local alignment search tool.

[0199] J Mol Biol 1990;215:403-10. https: / / doi.org / 10.1016 / S0022-2836(05)80360-2.

[0200] (21) Elnaggar JH, Ardizzone CM, Cerca N, Toh E, Laniewski P, Lillis RA, et al. A novel Gardnerella, Prevotella, and Lactobacillus standard that improves accuracy in quantifying bacterial burden in vaginal microbial communities. Front Cell Infect Microbiol

[0201] 2023 ; 13 : 1198113. https: / / doi.org / 10.3389 / fcimb .2023.1198113.

[0202] (22) Sarle WS. Neural Networks and Statistical Models. Nineteenth Annual SAS Users Group International Conference, Cary, NC, USA: SAS Institute Inc.; 1994.

[0203] (23) TensorFlow Developers. TensorFlow 2021. https: / / doi.org / 10.5281 / ZENODO.4758419. (24) Chollet F, others. Keras 2015.

[0204] (25) Yu H, Samuels DC, Zhao Y, Guo Y. Architectures and accuracy of artificial neural network for disease classification from omics data. BMC Genomics 2019;20:167. https:

[0205] / / doi.org / 10.1186 / sl2864-019-5546-z.

[0206] Bacterial vaginosis (BV) is a dysbiosis of the vaginal microbiota, characterized by the depletion of protective Lactobacillus spp. and an overgrowth of facultative and strict anaerobes. BV is linked to infertility, preterm birth, pelvic inflammatory disease, and increased risk of HIV / STI acquisition. Current BV diagnostics focus on symptomatic women and infection commonly recurs after treatment. We aimed to determine if artificial neural network (ANN) modeling of vaginal microbial communities can allow for early prediction of incident BV (iBV)with a high degree of accuracy.

[0207] Methods

[0208] 16S rRNA gene sequencing was performed to characterize the composition of vaginal microbial communities. This approach is a widely used, high-throughput amplicon-based sequencing method in which conserved regions of the bacterial 16S ribosomal RNA gene are amplified and sequenced to enable taxonomic identification and relative quantification ofDocket No. 2932719-000286-W01

[0209] Filed: March 27, 2026 bacterial taxa within a sample. Sequencing data were processed to generate feature tables representing the relative abundance of taxa. In certain embodiments, abundance data may be transformed using compositional data analysis techniques, including centered log-ratio (CLR) transformation, to account for the compositional nature of microbiome data and reduce spurious correlations.

[0210] Specimens (n= 1,201) included were obtained from two distinct longitudinal cohorts of women (n = 58) with baseline optimal vaginal microbiota who were followed for iBV development (per Nugent score). ANNs were trained using the relative abundance of vaginal taxa to predict whether a specimen was from a participant that developed iBV (pre-iBV) or not. Shapely additive explanations (SHAP) analysis was used to determine feature importance and assess how features contributed to model predictions.

[0211] Results

[0212] ANN modeling using the relative abundance of 20 vaginal bacterial taxa accurately classified >93% of specimens as either pre-iBV or healthy (sensitivity=95%, specificity=92%). Models trained using only the top five most important features achieved accuracy >91%, sensitivity >91%, and specificity >91%. Models trained using specimens from White and Black participants separately achieved high performance only using five features (accuracy >93%, sensitivity >92%, specificity >90%). SHAP analysis indicated that low relative abundance of Lactobacillus species contributed to pre-iBV predictions while low relative abundance of Gardnerella spp. contributed to healthy predictions.

[0213] Our study demonstrates that ANN modeling of few vaginal taxa can predict iBV prior to clinical onset, paving the way for potential interventions to prevent iBV. Modeling in race-specific cohorts underscores the importance of demographic considerations in microbiome-based diagnostics. Future work validating these models in larger cohorts is necessary.

[0214] Bacterial vaginosis (BV) is the most common vaginal infection worldwide and is associated with numerous adverse health outcomes. Changes in the vaginal microbiota occur 7- 14 days prior to the onset of incident BV (iBV). Modeling vaginal microbiome compositional patterns leading up to iBV onset can be used to predict the development of iBV.

[0215] In this study we developed artificial neural network models to predict iBV up to two weeks prior to its development based on the abundance of key vaginal bacterial taxa. The models predicted whether a vaginal specimen was collected from a participant who would develop iBVDocket No. 2932719-000286-W01

[0216] Filed: March 27, 2026 within two weeks of specimen collection or from a participant who did not develop iBV within two weeks of specimen collection with high accuracy (> 91%) by measuring as-few-as as five key vaginal bacterial taxa. Furthermore, our study suggests that personalizing models specific to key demographic features such as race improves modeling classification accuracy (> 93%) using five vaginal bacterial taxa.

[0217] Our models allow for the accurate prediction of iBV by surveying the vaginal microbiome, serving as a valuable tool to determine which patients are at risk of developing iBV in time to intervene before iBV development. Prediction of iBV development could lead to wider adoption of clinical interventions useful in the prevention of iBV such as live biotherapeutics, prophylactic antibiotics, and / or behavioral modifications. Similarly, our findings highlight the value of developing models personalized to specific patient populations, improving accuracy while reducing the number of features required for predictions.

[0218] INTRODUCTION

[0219] Bacterial vaginosis (BV) is the most common vaginal infection, affecting approximately 30% of reproductive-age women worldwide.1 3BV is a vaginal dysbiosis that is associated with multiple adverse health outcomes, including infertility, adverse birth outcomes, increased risk of HIV acquisition and other sexually transmitted infections (STIs), pelvic inflammatory disease, and increased risk of post-gynecologic surgery pelvic infections.1 4’5Although a large body of observational and interventional data suggest that BV is sexually transmitted,9it remains unclear whether BV results from acquisition of a single key pathogen, a polymicrobial consortium of pathogenic bacteria,2 10 12or depletion of protective lactobacilli,13 16which may allow for subsequent colonization by BV-associated bacteria (BVAB).17

[0220] The application of machine learning (ML) techniques to vaginal microbiome data presents an opportunity to advance BV diagnostics by modeling community dynamics prior to incident BV (iBV).18Artificial neural networks (ANNs) are one such supervised ML tool that has been successfully applied to model complex microbial relationships.19,20In our application of ANNs for vaginal microbiome analysis, sequencing-derived quantitative bacterial taxa data serve as ANN inputs.18This approach enables the identification of microbial signatures and patterns that may be critical in understanding and diagnosing BV.

[0221] In this study, we applied ANNs to predict iBV development using data from two prospective iBV pathogenesis studies.21 23While the prior studies tracked changes in vaginalDocket No. 2932719-000286-W01

[0222] Filed: March 27, 2026 microbial communities over time, our specialized ANN models allowed us to accurately predict whether a participant would develop iBV within the next two weeks based on a single vaginal specimen. These highly accurate models assess key vaginal microbial features, support development of personalized diagnostics, and improve characterization of microbial taxa contributing to iBV pathogenesis.

[0223] RESULTS

[0224] Clinical sampling and modeling workflow

[0225] For this study, we analyzed 1,201 self-collected vaginal specimens from a total of 58 women participating in 2 BV pathogenesis studies (32 who developed iBV and 26 healthy controls).21 23Following vaginal bacterial taxonomic assignment,20vaginal microbial taxa were retained for analysis after filtering to remove low-abundance, non-vaginal, or unclassified taxa to ensure biological relevance and model stability. These included Lactobacillus iners, L crispatus, L mulieris, Gardnerella spp., L. gasseri, Megasphaera lornae, Staphylococcus aureus, Finegoldia magna, Fannyhessea vaginae, Sneathia vaginalis, Aerococcus christensenii, Megamonas spp., Limosilactobacillus coleohominis, Prevotella timonensis, Corynebacterium kefirresidentii, Sneathia sanguinegens, Streptococcus anginosus, Candidatus Lachnocurva vaginae (previously known as BVAB 1), Gemella sanguinis, and Anaerococcus prevotii.

[0226] Of these vaginal specimens, 495 were collected 14 days prior to iBV onset (pre-iBV) and 706 were from healthy controls. The final ANN model was trained to predict whether a single vaginal specimen came from a pre-iBV participant or from a healthy participant. The workflow for specimen collection, data generation, and model training is shown in Figure 16A-16B, with cohort composition summarized in Figure 17.

[0227] For both cohorts shown in Figure 17, participants self-collected vaginal samples twice daily. Sampling was performed for up to 60 days or until the onset of incident bacterial vaginosis (iBV), whichever occurred first. Participants who did not develop iBV continued sampling for the full 60-day period.

[0228] For women who developed iBV, vaginal specimens collected within 14 days prior to iBV onset were selected for sequencing and analysis. Control participants were selected from women who did not develop iBV and who maintained optimal vaginal microbiota for the majority of the study period. Each control participant was matched to a specific iBV case participant based on age, race, and contraceptive method. Control specimens were then selected to correspond to theDocket No. 2932719-000286-W01

[0229] Filed: March 27, 2026 matched case specimens by day of menses, thereby aligning control and case samples according to physiologic timing rather than calendar day alone.

[0230] Control samples are not chosen based on a fixed interval. Instead, once a control participant is matched to a case, samples are selected from a comparable time window aligned by menstrual cycle phase so that the timing mirrors the case’s sampling window. For example, if a case participant develops iBV on day 20 of her cycle and samples for the 14 days prior to onset are analyzed, samples from the matched control participant are selected from a similar phase of her cycle, even though no iBV event occurs.

[0231] Although sampling was designed to capture specimens within 14 days prior to iBV onset, additional samples were included in certain cases, including specimens collected on the day of iBV diagnosis and in the immediate days following onset. As a result, some participants contributed more than 14 days of samples (e.g., up to approximately 17-18 days of sampling). Control participants were selected to provide a comparable number of samples, resulting in total sample counts that may exceed those expected from a strict 14-day, twice-daily sampling schedule.

[0232] The number of case and control specimens may differ due to the event-driven nature of iBV onset. For case participants, only samples collected prior to iBV onset are included, such that participants who develop iBV earlier in the collection period contribute fewer pre-iBV samples. In contrast, control participants do not have a corresponding onset event, allowing for selection of samples across the full matched sampling window. As a result, the total number of control specimens may exceed the number of case specimens.

[0233] Figure 16A-16B. Study workflow for training ANNs to predict pre-iBV from vaginal specimens. (A) Daily vaginal specimens were collected from 58 participants and monitored for development of iBV using Nugent scoring. DNA was isolated from vaginal specimens which was used for 16S sequencing. The relative abundance of key vaginal taxa was utilized to train ANN models to classify specimens as pre-iBV or healthy. (B) ANN models were trained using relative abundance of key vaginal taxa from individual specimens (n = 1201) to predict if a single specimen was from a participant who developed iBV within 14 days (pre-iBV) or from a participant who did not develop iBV within the next 14 days (healthy). The model was designed so that it can be used to classify a single vaginal sample as healthy or pre-iBV.

[0234] Figure 17. Incident bacterial vaginosis cohorts included in the study. Vaginal specimens collected from women participating in two longitudinal BV pathogenesis studies were used to trainDocket No. 2932719-000286-W01

[0235] Filed: March 27, 2026 the models. Cohort A consists of 22 African American women who have sex with women and Cohort B consists of 36 women who have sex with men between the ages of 18-40. Specimens in Cohort A were collected once daily during the 14 days leading up to iBV. Specimen in Cohort B were collected twice daily on the 14 days leading up to iBV.

[0236] Modeling for early prediction of iBV

[0237] To assess whether characterization of the vaginal microbiota could aid in early prediction of iBV, we trained ANN models using the relative abundance of all 20 vaginal bacterial taxa. Following parameter and architecture optimizations, the final model achieved a peak accuracy of 92% on training data and 93% on testing data. In addition to attaining high accuracy, the 20- feature model also achieved 95% sensitivity and 92% specificity (Figure 18A-18C, Table 6, Figure 19A-19B). Vaginal specimens from the test set of Cohort A were classified with 94% mean accuracy, while specimens from Cohort B were classified with 93% mean accuracy for specimens in the testing set (Figure 18C, Table 6).

[0238] A+ B A B

[0239] Metric sfraihs|i T rain|::|:s|il:7est:

[0240] Accuracy 0.92±0.03 093±0.06 0.88±0.03 0.94±0.11 0.94±0.03 0.93±0.06

[0241] Sensitivity 0.90±0.05 o 95±c.O8 0.86±0.10 C.88+0.29 0.93±0.06 0.96±0.07

[0242] Specificity 0.94±0.03 092±0.08 0.90±0.09 0.97±0.08 0.95±0.04 0.90±0.10

[0243] F1 Score liodbi lOiOibi):

[0244] AUROC 0.97±0.01 097±0.04 0.94±0.05 0.99±0.02 0.98±0.01 0.97±0.04

[0245] PPV NPV 0.93±0.03 096±0.05 0.84±0.10 0.94±0.17 0.96±0.03 0.97±0.05

[0246] Cohen. Kappa

[0247]

[0248] iiooiBO liOOiii ioii 1 s

[0249] Table 6. Model performance metrics. Classification accuracy on vaginal specimens used in the training set is shown in Train columns, while accuracy of model classification on testing samples is shown in Test columns. Metrics were calculated using data from both cohorts combined and for each separate cohort. Bootstrapping was used to calculate the mean values for each metric and with corresponding 95% confidence intervals, shown in each cell.

[0250] Abbreviations: AUROC = area under receiver operating curve, PPV = positive predictive value, NPV = negative predictive value

[0251] Figure 18A-18C. 20-Feature Model Classification Performance & AUC.

[0252] Summary of ANN modeling performance trained using 16S relative abundance data of 20 common vaginal bacterial taxa (n =1201). (A-B) Confusion Matrices model classifications in relation toDocket No. 2932719-000286-W01

[0253] Filed: March 27, 2026 true classification of (A) training data (n = 941) and (B) testing data (n = 260). Tiles containing correct classifications are shown in grey and tiles containing incorrect classification are shown in orange. Tile labels specify the number of samples in each category and percentage of samples in each category based on their true classification. (C) Bootstrapped classification accuracy is shown for all samples combined, cohort A samples, and cohort B samples.

[0254] Abbreviations: AUROC = area under receiver operating curve

[0255] Figure 19A-19B. 20-feature model training and validation curves. Training and validation curves plotting (A) model loss calculated from binary cross entropy and (B) classification accuracy for 600 epochs of model training, (training samples, n = 941; validation, n = 260).

[0256] To evaluate whether model performance varied based on proximity of specimen collection to iBV onset, we assessed classification accuracy separately for early pre-iBV specimens (collected 6-14 days prior to iBV onset) and late pre-iBV specimens (collected 1-5 days prior). The model correctly classified 97% of early pre-iBV specimens and 77% of late pre-iBV specimens across training and testing sets. This higher accuracy for early pre-iBV specimens may reflect more distinct microbial transitions that occur in the days preceding symptomatic onset, whereas late pre-iBV specimens may represent microbiomes already undergoing dynamic or unstable shifts, reducing model precision. Classification accuracy results are shown in Figure 20, and metrics are listed Table 7

[0257] B-Model W-M odel Metric Train Test Train Test Accuracy 0.90±0.02 0.93±0.06 0.95±0.04 0.94±0.09 Sensitivity 0.90±0.06 0.88±0.17 0.82±0.16 0.98±0.05 Specificity 0.90 ±0.06 0.95±0.06 0.99±0.01 0.85±0.30 Fl Score 0.91±0.04 0.86±0.13 0.89±0.10 0.96±0.06 AUROC 0.96 ±0.02 0.98±0.03 0.98±0.02 0.96±0.08 PPV 0.90 ±0.06 0.88±0.15 0.98 ±0.04 0.94±0.17 NPV 0.91 ±0.05 0.95±0.06 0.94 ±0.04 0.94±0.10

[0258]

[0259] Cohen Kappa 0.81 ±0.09 0.84±0.17 0.87±0.12 0.85±0.26 Table 7. Modeling classification performance metrics for race-specific models. Metrics for classification accuracy for models trained and tested only on samples from participants that selfidentified as black (B-Model) or white (W-Model). Classification metrics of training and testing samples are shown in Train columns and Test columns. Bootstrapping was used to calculate average values and 95% confidence interval shown in each cell.Docket No. 2932719-000286-W01

[0260] Filed: March 27, 2026 Figure 20.20-feature model performance predicting early and late pre-iBV. Accuracy of the 20-feature model was assessed for samples collected 6-14 days prior to iBV (early pre-iBV) and samples collected 1-5 days prior to iBV (late pre-iBV).

[0261] Role of microbial features in pre-iBV classification

[0262] Shapely additive explanations (SHAP) analysis of the 20-feature model assessed how each taxon (feature) contributed to model predictions (Figure 21).24This analysis revealed that several vaginal Lactobacillus spp. ranked highly as important features for model accuracy.25Specifically, high relative abundance of L. crispatus, L. iners, L. mulieris, and L. gasseri was associated with classifying specimens as healthy. A higher abundance of some BVAB such as Gardnerella spp. and Megasphaera lornae were found to be associated with pre-iBV predictions.

[0263] Figure 21. Feature importance and feature use determined from SHAP analysis. (A) Importance of features included in the 20-feature model determined by mean absolute SHAP value (n = 260). (B) Beeswarm plot visualizing how features contribute to model predictions. The relative abundance of each bacterial taxa is reflected by a color gradient: orange indicates a high relative abundance while blue indicates a low relative abundance. X-axis position indicates whether the feature contributes to predicting a vaginal specimen is classified as healthy (left) or pre-iBV (right). Distance from 0 on the x-axis indicates how much each feature contributes to accurately classifying an individual specimen as healthy or pre-iBV.

[0264] Minimum feature analysis

[0265] To evaluate the number of vaginal bacterial taxa required to accurately classify specimens, successive models were constructed with a training strategy using a limited subset of features. Specifically, five subset models were generated though training models using the top 3, 5, 7, 9, and 12 most important features (Figure 22, Table 8, Figure 23). The specific taxa included in each subset model are listed in Figure 24. The model trained using the top three most important features (L. crispatus, Gardenerella, L. iners) had a peak training accuracy of 87.6% and a testing accuracy of 89.6%. Models trained using the top 5, 7, 9, and 12 features achieved peak training accuracy ranging from 90%-93% and peak sensitivity ranging from 91-97%, indicating that the inclusion of additional taxa improves overall performance and classification of pre-iBV specimens. These findings indicate that, while surveying a few key vaginal bacterial taxa is sufficient to classify specimens as pre-iBV or healthy, incorporating additional taxa enhances the classification of pre-iBV specimens. When models were evaluated separately by cohort, bothDocket No. 2932719-000286-W01

[0266] Filed: March 27, 2026 Cohort A and Cohort B specimens were classified with high accuracy, though Cohort A specimens consistently showed higher classification performance across all subset models. (Figure 25).

[0267] Figure 22. N-feature models receiver operating curves and precision-recall curves.

[0268] Performance of models trained on the three, five, seven, nine, and twelve most important features determined by mean absolute SHAP value are visualized as (A-B) receiver operating curves and (C-D) precision recall curves. (A,C) Curves calculated from training data (n = 941). (B,D) Curves calculated from testing data (n = 260).

[0269] Figure 23A-23E. N-feature model training and validation curves. Training and validations curves plotting model loss over 400-600 epochs of training are shown on the left of each panel. Training and validation curves plotting accuracy over the course of training is shown on the right of each panel.

[0270] Figure 24. Taxa used to train race-specific feature subset models. Taxa included in each subset models as features are denoted with a check mark.

[0271] Figure 25A-25E. Subset model performance stratified by cohort. Metrics for each subset model were calculated for the grouped cohorts (A+B), the Cohort A, and Cohort B. Each box plot was generated by calculated accuracies from 1000 bootstrapped resamples of the testing data.

[0272] Top 3 Top 5 Top 7 Top 9 Top 12 Metric

[0273] Accuracy 0.87±0.04 0.89±0.07 0.92±0.03 0.91 ±0.06 0.92±0.03 0.91 ±0.06 0.93±0.03 0.93±0.05 0.90±0.03 0.93±0.05 Sensitivity

[0274] Specificity 0.88±0.05 0.84±0.12 0.91 ±0.04 0.91 ±0.08 0.93±0.03 0.90±0.06 0.93±0.03 0.91 ±0.09 0.92±0.04 0.96±0.05 F1 Score

[0275] AU ROC 0.95±0.02 0.96±0.04 0.97±0.01 0.96±0.04 0.97±0.01 0.97±0.03 0.98±0.01 0.96±0.03 0.96±0.01 0.93±0.08 PPV 088+005 088+010 090-004 088-010 089+005 094-008 NPV 0.89±0.04 0.977±0.04 0.95±0.03 0.94±0.06 0.93±0.03 0.95±0.06 0.94±0.03 0.97±0.04 0.91 ±0.03 0.93±0.06 Cohen Kappa 0.84±0.06 0.82±0.14 0.84±0.06 0.83±0.13 0.85±0.06 0.86±0.11 0.80±0.07 0.87±0.12

[0276]

[0277] Table 8. N-feature model performance metrics.

[0278] Classification accuracy metrics for models trained on the top three, five, seven, nine, and twelve most important features. Classification metrics of training and testing specimens are shown in Train columns and Test columns. Bootstrapping was used to calculate average values and 95% confidence interval, shown in each cell.Docket No. 2932719-000286-W01

[0279] Filed: March 27, 2026 Abbreviations: AUROC = area under receiver Operating curve, PPV = positive predictive value, NPV = negative predictive value

[0280] SHAP analysis was applied to determine how features were utilized by the n-feature models (Figure 26). Across all n-feature models, L. crispatus and Gardnerella spp. consistently emerged as the two most important bacterial taxa, similar to findings from the 20-feature model. Overall, results from SHAP analysis reflect Lactobacillus spp. and Gardnerella spp. are important to predictions across all feature sets tested.

[0281] Figure 26A-26E. SHAP analysis of n-feature models. (A-E) Beeswarm plots indicating feature importance and feature contribution for models trained using the top three, five, seven, nine, and twelve most important taxa for model predictions. Orange indicates high relative abundance, and blue indicates low relative abundance. X-axis position indicates whether the feature contributes to predicting a sample is classified as healthy (left) or pre-iBV (right). Distance from 0 on the X-axis indicates the magnitude of each features contribution to predictions.

[0282] External validation performance

[0283] To further evaluate model performance, models were tested using 421 specimens (Healthy n = 389; PreiBV n = 32) from an external validation cohort consisting of 27 women (Healthy n = 12 ; PreiBV n= 15).26From this analysis, we found that the 20-feature models performed best on external validation, achieving a balanced accuracy of 80% and an AUROC of 83%. While performance varied across subset models, the top performing models included the 5-feature model and 12-feature model, which achieved a balanced accuracy score of 76% and 79%, with corresponding AUROC values of 80% and 83% respectively. Model performance metrics are further listed in Table 9.

[0284] Metric Top 3 Top 5 Top 7 Top 9 Top 12 Top 20

[0285] Balanced Accuracy 0.67 076 0.71 0.63 0.79 0.80

[0286] Accuracy

[0287] Sensitivity 083 074 058 064 077 080

[0288] Specificity

[0289] AUROC 0.78 083 0.82 0.77 0.80 0.83

[0290] Youden’s J i3Sl|i

[0291]

[0292] Table 9. Model performance metrics on external validation cohort. Classification accuracy metrics for models trained on the top three, five, seven, nine, twelve, and twenty features. Abbreviations: AUROC = area under receiver operating curveDocket No. 2932719-000286-W01

[0293] Filed: March 27, 2026 Race-specific models: performance and feature utilization

[0294] Considering there is evidence suggesting that the composition of the vaginal microbiota differs based on race, 27, 28 we trained models separately using specimens from Black (B-Model, n = 811) and White (W-Model, n = 380) participants. After training, the W-Model achieved >94% accuracy on both the training and testing datasets, while the B-Model achieved >90% accuracy (Figure 27, Table 7. The disparity in performance is likely due to the B-Model including data from both Cohort A and Cohort B as training and testing data, while the W-Model only utilized data from Cohort B.

[0295] Figure 27. Classification performance and SHAP analysis of models trained on racespecific data. (A-B) Confusion matrices of classification results of testing data on models trained on specimens from (A) Black participants (n = 811) or (B) White participants (n = 380). Grey tiles indicate correct classifications by the model. Tile labels specify the number and percentage of specimens in each category based on their true classification. (C-D) Beeswarm plots indicate feature importance and how features contribute to model predictions. SHAP analysis was performed on testing data from models trained using specimens from (C) Black participants (n = 190 testing specimens) or (D) White participants (n = 73 testing specimens). The relative abundance of each taxa is reflected by the color gradient; orange indicates high relative abundance and blue indicates low relative abundance for a given taxa in comparison to other samples. X-axis position indicates whether the feature contributes to predicting a specimen is classified as healthy (left) or pre-iBV (right). Distance from 0 on the x-axis indicates how much each feature contributes to accurately classifying an individual specimen.

[0296] Race-specific models were additionally trained using feature subsets to evaluate whether models maintained high classification accuracy when trained using few features. From this analysis we found that each race-specific model maintained high classification accuracy and AUROC when trained on three highly important features (Test Accuracy = 95%, AUROC > 0.99), and (Accuracy = 0.88 , AUROC = 0.96) Peak classification accuracy was observed in the 9- feature B-Model (Accuracy = 92% ) and 7-feature W-model (Test Accuracy = 100%) (Figure 28, Table 10).

[0297] A. B ack-Subset Models

[0298] Top 3 Top 5 Top 7 Top 9 Top 12 Metric Train Test Train Test Train Test Train Test Train Test Accuracy 0.88±0.04 0.92±0.06 0.94±0.03 0.93±0.06 0 89±0.04 0.82±0 10 0.90±004 0.91±0.07 0.96±0.02 0.92±0.07 Sensitivity 0.87±0.07 0.84±0.19 0.95±0.04 0.92±0.13 0 88±0.07 0.84±021 0.91±006 0.82±0.21 0.98±0.02 0.97±0.05

[0299]

[0300] Specificity 0.89±0.06 0.95±0.07 0.93±0.05 0.94±0.07 090±0.06 0.8H0.12 0.88±006 0.94±0.07 0.95±0.04 0.90±0.10Docket No. 2932719-000286-W01

[0301] Filed: March 27, 2026 Fl Score 0.88+0.05 0.85+0.13 0.94+0.03 0.88+0.12 0 89+0.05 0.72+0 14 0.90+004 0.83+0.15 0.96+0.02 0.87+0.11 AUC 0.96+0.02 0.97+0.03 0.98+0.01 0.98+0.02 094+0.03 0.94+006 0.96+002 0.93+0.06 0.99+0.00 0.98+0.02 PPV 0.88+0.06 0.86+0.17 0.93+0.05 0.85+0.16 0 89+0.06 0.63+0 17 0.88+005 0.84+0.16 0.95+0.04 0.79+0.17 NPV 0.87+0.06 0.94+0.06 0.95+0.04 0.97+0.05 0 89+0.06 0.93+008 0.91+005 0.93+0.07 0.98+0.02 0.99+0.02 Cohen Kappa 0,76*0,10 0,80*0,18 0,88*0.07 0.84*0,17 0,79*0,09 0,60*0.22 0,87*0.09 0,77*0.21 0,73*0,05 0,82*0,17 B. White-Subset Models

[0302] Top 3 Top 5 Top 7 Top 9 Top 12 Metric Train Test Train Test Train Test Tram Test Train Test Accuracy 0.95+0.04 0.89+0.13 0.93+0.04 0.94+0.09 095+0.04 0.93+0 10 0.95+004 0.93+0.10 0.95+0.04 0.94+0.09 Sensitivity 0.89+0.14 0.90+0.15 0.78+0.18 0.96+0.09 0 85+0.16 0.96+009 0.85+0 16 0.98+0.05 0.85+0.16 1.00+0.00 Specificity 0.97+0.03 0.85+0.30 0.97+0.03 0.90+0.25 098+0.03 0.85+030 0.98+003 0.79+0.35 0.98+0.03 0.79+0.35 Fl Score 0.90+0.09 0.92+0.10 0.84+0.12 0.96+0.06 090+0.10 0.95+007 0.90+0 10 0.95+0.07 0.90+0.10 0.96+0.06 AUC 0.98+0.01 0.91+0.16 0.96+0.03 0.96+0.10 098+0.01 0.94+0 12 0.98+001 0.93+0.13 0.98+0.01 0.96+0.09 PPV 0.91+0.11 0.94+0.11 0.92+0.12 0.96+0.08 095+0.09 0.94+0 10 0.95+009 0.92+0.11 0.95+0.09 0.93+0.11 NPV 0.96+0.04 0 77+0 30 093+0 05 090+0 22 095+004 0 89+024 0.95+0.04 0 94+0 18 0 95+004 1 00+0 00

[0303]

[0304] Cohen Kappa 0 87+0 12 0 73+0 33 0 80+0 15 0 86+0 24 0 87+0 13 0 82+029 0 87+0 13 0 81+031 0 87+0 13 0 84+0 27 Table 10. Performance of race specific models trained on feature subsets. Metrics for classification accuracy for models trained on three, five, seven, nine, and twelve of the most important features from 20-feature model trained on (A) black participants or (B) white participants. Bootstrapping was used to calculate average values and 95% confidence interval shown in each cell.

[0305] Figure 28. Receiver operating curves of race-specific n-feature models. (A-B) Receiver operating curves of models trained the top three, five, seven, nine, twelve most important features to B-model predictions based on classification on (A) training data (n = 572) and (B) testing data (n = 239). (C-D) Receiver operating curves of models trained the top three, five, seven, nine, and twelve most important features to W-Model predictions based on classification on (C) training data (n = 190) and (D) testing data (n = 73).

[0306] To further evaluate how individual features contributed to model predictions, SHAP analysis was performed on each race-specific subset model (Figure 29, Figure 30). Across both the Black (B-models) and White (W-models) participant models, L. crispatus and Gardnerella spp. consistently ranked as the most important features. A high relative abundance of L. crispatus was associated with healthy classifications, while a high abundance of Gardnerella spp. was associated with pre-iBV classifications, consistent with findings from the 20-feature models. Interestingly, in the B-models, L. iners ranked as the third most important feature but did not show a clear relationship with either healthy or pre-iBV states. These results suggest that race-specific models differ slightly in how individual bacterial taxa influence predictions, reflecting possible biological or compositional differences between groups.

[0307] Figure 29A-28D. SHAP analysis of models trained on top features set in Black participants. (A-E) Beeswarm plots indicating feature importance and feature contribution for models trained on samples from black participants using the top three, five, seven, nine, and twelve most important taxa for model predictions. Orange indicates high relative abundance, and blueDocket No. 2932719-000286-W01

[0308] Filed: March 27, 2026 indicates low relative abundance. X-axis position indicates whether the feature contributes to predicting a sample is classified as healthy (left) or pre-iBV (right). Distance from 0 on the X-axis indicates the magnitude of each features contribution to predictions.

[0309] Figure 30A-30E. SHAP analysis of models trained on top features set in White participants. (A-E) Beeswarm plots indicating feature importance and feature contribution for models trained on samples from white participants using the top three, five, seven, nine, and twelve most important taxa for model predictions. Orange indicates high relative abundance, and blue indicates low relative abundance. X-axis position indicates whether the feature contributes to predicting a sample is classified as healthy (left) or pre-iBV (right). Distance from 0 on the X-axis indicates the magnitude of each features contribution to predictions.

[0310] Early Pre-i BV Samples Late Pre-iE IV Samples Metric Train Test Train Test

[0311]

[0312] Accuracy 0.95±0.05 0.97±0.07 0.89±0.08 0.77±0.22 Table 11. Classification accuracy of early and late pre-iBV samples. Classification accuracy of 20-feature model on late pre-iBV samples (collected <5 days prior to iBV) and early pre-iBV (collected >5 days prior to iBV). Values shown in each cell indicate mean bootstrapped accuracy ± 95% confidence interval.

[0313] DISCUSSION

[0314] In this study, we developed ANN models capable of accurately predicting the onset of iBV up to 14 days in advance from a single vaginal specimen. The models were trained using vaginal specimens from two distinct longitudinal cohorts of women that employed dense specimen collection and clearly defined endpoints of iBV21-22-29Our models achieved high accuracy, sensitivity, and specificity using the relative abundance of vaginal bacterial taxa to make predictions.30 From training models with a variety of taxa subsets, we found that an ANN model trained on 20 vaginal bacterial taxa achieved the highest overall classification performance on external and internal validation specimens, indicating it generalizes well to a variety of specimens. This predictive capability enables early prediction of iBV, allowing for proactive interventions and / or behavioral modifications in an effort to prevent its development.

[0315] Feature analysis provided insight into the microbial patterns correlating with pre-iBV and healthy states. From SHAP analysis, we found that common vaginal Lactobacillus spp. and Gardnerella spp. to be the most important taxa to accurately classify specimens as healthy or pre-iBV. As expected, a higher abundance of Lactobacillus spp. was indicative of stable healthyDocket No. 2932719-000286-W01

[0316] Filed: March 27, 2026 communities, while a higher abundance of Gardnerella spp. was indicative of pre-iBV communities.23,25,27Additionally, our results support the notion that L. crispatus is critical to protection against iBV, while less common Lactobacillus spp., such as L. gasseri and mulieris (previously classified as L. jensenii)^

[0317]

[0318] may vary in their protective capacity due to strain variation and host features.32,33Additionally, we found that taxa in the genus Megamonas ranked high as the fifth most important feature in subset models. While this taxon has not previously been implicated in iBV, these findings suggest Megamonas may contribute to the microbial shifts preceding iBV onset. Further studies will be required to further assess the relationship between Megamonas and iBV pathogenesis.

[0319] Consistent with current hypothetical models of iBV pathogenesis, our findings support the hypothesis that colonization by Gardnerella spp. and depletion of protective Lactobacillus species occur as early events preceding iBV onset.10,34,35The strong predictive importance of Gardnerella spp. in our models, coupled with the reduced relative abundance of L. crispatus in pre-iBV specimens, reinforces the concept that Gardnerella spp. proliferation marks a transitional stage from a Lactobacillus-dominant to a dysbiotic vaginal environment. These results align with prior longitudinal studies showing that Gardnerella spp. expansion precede the development of symptomatic BV and contributes to biofilm formation and ecological instability in the vaginal microbiota.17,27

[0320] Minimum feature analysis revealed that only a few taxa were necessary for highly accurate predictions. Models trained on the five most important features (L. crispatus, Gardnerella spp., L. iners, L. mulieris, and Megamonas) achieved >90% accuracy. While the inclusion of additional vaginal bacterial taxa slightly improved accuracy, the strong performance of minimal models suggests that simplified approaches to characterizing the vaginal microbiota may be sufficient for specific research and diagnostic purposes. For example, using targeted qPCR panels of five to seven organisms could be feasible alternatives to complex sequencing-based profiling for early iBV prediction.

[0321] Models developed from race-specific subgroups demonstrated improved accuracy compared to models trained on the complete cohort. These findings are consistent with prior studies showing that longitudinal changes in vaginal bacterial community composition differ by race.27,28While the overall predictive features were similar between groups, L. crispatus and Gardnerella spp. consistently ranked as the strongest predictors. In contrast, the contribution of L.Docket No. 2932719-000286-W01

[0322] Filed: March 27, 2026 iners differed notably. In models trained on specimens from Black participants, L. iners was variably associated with both healthy and pre-iBV states, suggesting its dual role as a transitional species within less stable microbial communities. In contrast, in models derived from White participants, L. iners was more often associated with healthy classifications, consistent with greater community stability. These findings underscore that race-specific microbial ecology influences how key bacterial taxa relate to vaginal health and dysbiosis, highlighting the potential value of personalized diagnostic models for early iBV prediction.

[0323] Beyond the core predictive taxa incorporated into our final models, several additional bacterial features emerged as highly ranked during feature selection, despite not being retained in the optimized ANN models. These taxa, including F. vaginae, P. bivia, M. lornae, and Sneathia spp. have been implicated in iBV pathogenesis and recurrence through their roles in biofilm formation, metabolic cross-feeding, and host immune modulation.10,36The inclusion of these taxa in our top performing classifiers supports current models of BV as a polymicrobial infection characterized by synergistic interactions among multiple BVAB rather than dominance by a single bacterial pathogen. Moreover, the identification of these taxa reinforces the biological relevance of our modeling approach by independently highlighting microbial players known to contribute to a vaginal dysbiosis. While these organisms were ultimately excluded to prevent model overfitting, their prominence underscores how machine learning can uncover secondary but biologically meaningful contributors to iBV development, providing both validation of known microbial mechanisms and a framework for refining future diagnostic and therapeutic strategies.

[0324] Early prediction of iBV could have significant clinical implications, allowing for early clinical interventions such as live biotherapeutics, prophylactic antibiotics, and / or behavioral modifications in an attempt to prevent development of iBV. Furthermore, high accuracy observed on race-specific models underscores the importance of considering demographic and biological diversity in the development of diagnostic models trained on the vaginal microbiome.

[0325] Our inclusion of race-stratified models provides additional insight into how vaginal microbiome composition and predictive features may differ across populations. Although the top predictive taxa, L. crispatus, Gardnerella spp., and L. iners, were consistent between groups, the relative contribution of each varied by race, suggesting that the microbial context surrounding shared taxa differs across populations. These findings highlight the importance of considering host demographic and microbial diversity in model development, as aggregating data acrossDocket No. 2932719-000286-W01

[0326] Filed: March 27, 2026 heterogeneous groups may obscure population-specific microbial patterns relevant to iBV pathogenesis and diagnostic accuracy. Incorporating such demographic factors into future models could improve their generalizability and clinical applicability across diverse patient populations.

[0327] While we have rigorously conducted our analysis and added valuable insight into the early prediction of iBV, there are several limitations to our study. Although our longitudinal datasets are rich in specimens, the number of women in the two cohorts are relatively small. Additionally, the data used to train our ANN models combined two cohorts which have sampling biases. Cohort A exclusively included African American women, while Cohort B reflects the racial demographics of the Birmingham, AL metropolitan area, as it was primarily composed of participants that identified as Black or White. These sampling biases limit our ability to identify specific racial / ethnic relationships between pre-iBV and the vaginal microbiome for other groups such as Hispanics, Asians, and Native Americans.37 39Future studies externally validating our models with larger cohorts from more racially diverse cohorts will be required to verify model validity. Additionally, while the external validation cohort utilized in this study provided evidence that the models performed well using additional specimens, participants included in this cohort were not screened for recurrent BV prior to sampling.27Furthermore, since SHAP analysis identifies only correlations between model predictions and microbial taxa, future studies are necessary to validate the causal link between vaginal microbiota and pre-iBV classifications. Similarly, future studies are required to identify the underlying biological mechanisms mediating the microbial relationships found through SHAP analysis.40Although our model was trained on a relatively small cohort with specific enrollment criteria, the high accuracy achieved reflects the strong biological patterns underlying the pre-iBV state.

[0328] Our models offer an innovative opportunity to enhance future clinical diagnostic platforms for BV by enabling early prediction of BV in asymptomatic or pre-symptomatic women. This represents a significant advancement beyond current FDA-cleared BV molecular diagnostics which are exclusively authorized for use in symptomatic women (Aptima BV Assay, Hologic, Marlborough, MA; Xpert Xpress MVP, Cepheid, Berkeley, CA; BD MAX Vaginal Panel, Geneohm, Quebec City, QC)41By utilizing our models to recognize early microbial shifts predictive of BV, clinicians could identify patients at risk before the onset of symptoms, allowing for earlier intervention and possibly reducing the incidence of complications such as preterm birth or recurrent BV in certain populations.42Our models, trained on comprehensive vaginalDocket No. 2932719-000286-W01

[0329] Filed: March 27, 2026 microbiome datasets, could detect subclinical dysbiosis patterns and predict BV onset or recurrence, helping to fill a major gap in current diagnostic capabilities.

[0330] Incorporating models, such as the ones described in this study, into the diagnostic development pipeline also holds promise for advancing personalized medicine in the context of microbiome-associated and polymicrobial diseases. Additionally, the use of these models, trained on population-specific data, can support the rapid identification of clinically relevant microbial targets to inform the design of diagnostic panels. This is particularly important given that existing FDA-cleared BV molecular diagnostics differ significantly in the microbial targets they include, with little overlap across platforms.41Integrating ANN approaches could standardize and optimize target selection, enabling the development of more inclusive and effective diagnostics that better reflect the diverse microbial populations associated with BV.

[0331] In summary, our study demonstrates that machine learning approaches can effectively predict iBV prior to its onset, highlighting the potential for preemptive diagnostics for iBV based on a single self-collected vaginal specimen. Our findings emphasize the role that Lactobacillus spp. play in regulating vaginal microbiome dynamics as well as race-specific differences in microbial predictors of pre-iBV states. Furthermore, these results support the hypothesis that iBV is mediated by a depletion of Lactobacillus spp. and enrichment of Gardnerella spp. prior to iBV onset. Early prediction of iBV could lead to wider adoption of clinical interventions useful in the prevention of iBV, such as live biotherapeutics, prophylactic antibiotics, and / or behavioral modifications.43,44Integrating personalized microbiome-based diagnostics capable of early prediction of vaginal disorders could revolutionize the management of BV and other important vaginal health conditions.

[0332] METHODS

[0333] Clinical enrollment

[0334] The training data were derived from two prospective longitudinal cohorts designed to study the sequence of microbiological events prior to iBV.21,29The two cohorts were active between 2014-2017 (Cohort A)29and 2020-2024 (Cohort B)21, enrolling English-speaking women aged 18-45 in the Birmingham, AL metropolitan area.

[0335] Exclusion criteria for both cohorts included oral or intravaginal antibiotics within the past 14 days, self-reported HIV infection, or current pregnancy. Participants had to be asymptomatic at enrollment, with no Amsel criteria and normal Nugent scores of 0-3 with no Gardnerella spp.Docket No. 2932719-000286-W01

[0336] Filed: March 27, 2026 morphotypes on baseline vaginal Gram stain45,46At enrollment, participants in both cohorts were screened for Trichomonas vaginalis, Chlamydia trachomatis, Neisseria gonorrhoeae by highly sensitive nucleic acid amplification tests (NAATs). Participants who tested positive for one or more of these infections were treated per standard of care and dropped from the study.47A total of 58 participants were included in this ANN analysis: Cohort A included 22 participants, while Cohort B included 36 participants. Cohort A exclusively enrolled African American / Black women who have sex with women (WSW), while Cohort B enrolled women who have sex with men (WSM) of all races and ethnicities. Detailed enrollment criteria are available in prior publications.21 23

[0337] Participants in Cohort A self-collected vaginal specimens once daily for 90 days or until the onset of iBV, while participants in Cohort B self-collected vaginal specimens twice daily for 60 days or until the onset of iBV. Additional details regarding specimen collection, delivery to the study site, storage, and preparation for 16S rRNA gene sequencing are available in prior publications.21,23These cohorts were specifically selected for this analysis due to their highly similar study designs, and strict inclusion and exclusion criteria, which allowed for reliable detection of iBV cases with matching controls. Controls were participants that did not develop iBV matched to specific cases based on age, race, and contraceptive method.

[0338] Selection of vaginal specimens for vaginal microbiome characterization

[0339] Participants were monitored for iBV throughout the course of both studies. iBV was defined as a Nugent score of 7-10 on >2 consecutive days in Cohort A and on >4 consecutive vaginal specimens / 2 days in Cohort B and A. For women who developed iBV in both cohorts, specimens collected up to 14 days prior to iBV were selected for 16S sequencing.46 Controls from both cohorts were selected if they maintained optimal vaginal microbiota for the majority of the study (e.g., a Nugent score of 0-3 for at least 85% of study days). Control specimens were matched to iBV case specimens by day of menses, as previously described.22,23All specimens were shipped to the Microbial Genomics Resource Group (MGRG) at the Louisiana State University Health Sciences Center (LSUHSC) in New Orleans, LA for DNA isolation, 16S rRNA sequencing, and bioinformatics analysis.Docket No. 2932719-000286-W01

[0340] Filed: March 27, 2026 Molecular methods

[0341] DNA isolated from stored vaginal specimens from women in both cohorts was used for 16S rRNA gene sequencing from women who developed iBV and matched healthy controls to determine changes in the relative abundance of vaginal bacteria over time.

[0342] DNA was isolated from the vaginal specimens, as described previously, using a modified QIAamp DNA Mini Kit (QIAGEN).23,48PCR amplicon libraries spanning the taxonomically informative fourth hypervariable (V4) region of the 16S rRNA were generated and sequenced on the Illumina MiSeq platform.48Sequencing data were analyzed in R v4.2.1 using DADA2 vl.16 to identify sequence variants,49with taxonomic classification performed using SpeciatelT and SpeciateDB.50

[0343] External validation cohort

[0344] To ensure that the ANN models generalized beyond the original study populations and were not overfit to cohort-specific characteristics, we included an external validation cohort in this study to independently assess model performance on a distinct dataset. Validation data were obtained from PRJNA46331 (SRR17141071-SRR17141708).27This external validation cohort consists of longitudinal vaginal specimens collected twice-weekly from 32 women living in the Baltimore, Maryland metropolitan area between 2008- 2011, with accompanying Nugent scores and 16S rRNA sequencing data. For the purposes of our study, cases were defined as participants with a single Nugent score >7 (n=12) and controls were defined as participants without a Nugent score >7 at any point during the study (n=15). Vaginal specimens (n=32) collected within 14 days prior the first Nugent score >7 were considered pre-iBV specimens; 389 specimens were used from the control population. Five participants were excluded due to missing Nugent scores or Nugent scores >7 at the first available timepoint. Amplicon sequence variants (ASV) were identified using the DADA2 workflow, and taxa were mapped to ASVs using SpeciatelT.50

[0345] Modeling approach and parameters

[0346] Using the sequencing data obtained from cohort A and B vaginal specimens, ANNs were built and compared using TensorFlow (v2.6.1) and Keras (v3.0).51,52 Specimens were grouped into “pre-iBV” (collected <14 days prior to iBV onset) or “healthy” (no iBV). Data were split by participant: 80% for training with the labels “pre-iBV” or “healthy,” and 20% for validation without labels. All specimens from a participant were assigned to either training or testing sets to avoid data leakage.Docket No. 2932719-000286-W01

[0347] Filed: March 27, 2026 Each ANN was trained 200-600 iterations (epochs) with early stopping to prevent overfitting. Several architectures and combinations of hidden layers were compared to achieve optimal parameters. This range of input structures and model architectures is known to be optimal for similar omics datasets.53Stochastic gradient descent was used as the optimization algorithm, and neurons used a rectified linear unit (ReLU) activation function. (Under an alternative embodiment, ADAM may be used as the optimization algorithm). A Gaussian noise layer was included before the first dense layer to limit overfitting and improve generalization. A grid search approach was used to determine optimal number of hidden layers, node combinations, dropout percentages, Gaussian noise, batch size, and regularization. Alternative optimizers and activation functions were also tested. Final parameters were selected based on accuracy and loss of training and validation data.

[0348] Following training and optimization of the 20-feature model, SHAP analysis was used to determine feature importance. Subsequent models were trained using subsets of top-ranked features. ANNs were implemented in Python3 in Visual Studio Code (vl.96.2). Visualizations were created using matplotlib (V3.7.1)?4

[0349] Using sequencing data obtained from Cohort A and Cohort B vaginal specimens, artificial neural networks (ANNs) were constructed and compared using TensorFlow (v2.6.1) and Keras (v3.0).

[0350] Specimens were grouped into “pre-iBV” (collected <14 days prior to iBV onset) or “healthy” (no iBV). Data were split at the participant level, with 80% assigned to training and 20% to validation, ensuring that all samples from a given participant were included exclusively in one set to prevent data leakage.

[0351] The final ANN architecture consisted of a fully connected feed-forward network with an input layer corresponding to the number of taxa features (n = 20), followed by one or more hidden dense layers with rectified linear unit (ReLU) activation functions, and a single-node output layer with sigmoid activation for binary classification.

[0352] Each ANN was trained for 200-600 epochs with early stopping to prevent overfitting. Stochastic gradient descent was used as the optimization algorithm. A Gaussian noise layer was applied to the input layer during training, in which small random values sampled from a normal distribution (mean = 0, standard deviation optimized via grid search) were added to input features. This regularization approach improves generalization and reduces overfitting.Docket No. 2932719-000286-W01

[0353] Filed: March 27, 2026 A grid search approach was used to determine optimal model parameters, including number of hidden layers, node combinations, dropout rates, Gaussian noise magnitude, batch size, and L1 / L2 regularization. Alternative optimizers and activation functions were also evaluated. Final model parameters were selected based on training and validation accuracy and loss.

[0354] Following training and optimization of the 20-feature model, SHAP analysis was used to determine feature importance. Subsequent models were trained using subsets of top-ranked features. ANNs were implemented in Python 3 in Visual Studio Code (vl .96.2), and visualizations were generated using matplotlib (v3.7.1).

[0355] The ANN described here is the same feed-forward neural network architecture described above, with additional implementation details provided for reproducibility.

[0356] A Gaussian noise layer was applied to the input layer during training, in which small random values sampled from a normal distribution (mean = 0, standard deviation optimized via grid search) were added to input features. This approach serves as a regularization technique, improving model generalization by reducing sensitivity to small fluctuations in feature values and limiting overfitting.

[0357] As shown in Figure 31, the artificial neural network (ANN) comprises a feed-forward fully connected architecture. The input layer consists of features corresponding to the relative abundance of selected vaginal microbial taxa. Prior to entering the first hidden layer, a Gaussian noise layer is applied, in which small random values sampled from a normal distribution (mean = 0, standard deviation « 0.003) are added to input features during training to improve generalization and reduce overfitting.

[0358] The network includes two hidden dense layers, with node counts varying depending on the feature set (e.g., 38 and 34 nodes for the 20-feature model). Each hidden layer uses a rectified linear unit (ReLU) activation function. Optional regularization techniques, including dropout and L1 / L2 weight penalties, are applied as indicated in Figure 31, where L1 / L2 values represent equal coefficients applied to both LI and L2 regularization terms.

[0359] The output layer consists of a single neuron with a sigmoid activation function, producing a probability score corresponding to classification of a specimen as pre-iBV or healthy.

[0360] Model training is performed using stochastic gradient descent with a learning rate of approximately 0.001 and batch sizes as specified in Figure 31. The various parameter combinationsDocket No. 2932719-000286-W01

[0361] Filed: March 27, 2026 shown in Figure 31 represent configurations evaluated during model optimization, with final models selected based on training and validation performance.

[0362] Figure 31 shows ANN model specific parameters, under an embodiment.

[0363] Determination of key BV-associated bacteria in the development of iBV

[0364] Once the ANN models were trained and validated, we deconstructed them using the SHAP (v0.46.0) Kernel Explainer method to determine the “relative importance” of features included in each model.24The data used to train the models was included as the background, while testing data was used to compute SHAP values for individual predictions. Training specimens were summarized into 50 weighted mean samples using the shap.kmeans function. Mean absolute SHAP values were calculated for each feature by averaging SHAP values across all participants to assess feature importance for each model. Additionally, SHAP analysis was used to assess how features contributed to model predictions for individual specimens. Mean absolute SHAP values were used to select features for each model trained on a subset of features.

[0365] Sample size estimation

[0366] Sample size of the cohorts was determined based on convenience sampling. All vaginal specimen data available at the time of analysis were included in the study.

[0367] Statistical analysis and Metric Calculations

[0368] A bootstrapping approach was used to assess model performance metrics. 1000 bootstrapped resamples were performed for each model. Bootstrap resamples were used to calculate 95% confidence intervals and average model performance. Models were evaluated on training and testing sets using accuracy, Area Under Receiver Operating Curve (AUROC), Negative Predictive Value (NPV), Positive Predictive Value (PPV), sensitivity, specificity, Fl Score, and Kappa Cohen. Considering the external validation set was highly imbalanced (< 20% of specimens were pre-iBV) NPV, PPV, Fl score, and Kappa Cohen were not used to evaluate the models’ classification performance on external validation data. Balanced accuracy was included as an additional metric for evaluating external validation data classification performance since it equally weighed classification accuracy of healthy and pre-iBV specimens. Similarly, Youden’s J Index was included as an additional metric to summarize model performance on external validation specimens since it is insensitive to class imbalances. Metrics were calculated using sci-kit learn (vl.6.1).Docket No. 2932719-000286-W01

[0369] Filed: March 27, 2026 Treatment: Tn certain embodiments, the predictive model is used to identify individuals at risk of developing bacterial vaginosis prior to clinical onset. Upon identification of a high-risk or pre-iBV state, a therapeutic intervention may be initiated prophylactically or early in the disease course.

[0370] Treatment options may include administration of antibiotics commonly used for bacterial vaginosis, including but not limited to metronidazole (oral or intravaginal), clindamycin (oral or intravaginal), or tinidazole. In some embodiments, treatment may further include adjunctive or alternative therapies, such as probiotics, vaginal microbiome-modulating agents, or other interventions aimed at restoring a Lactobacillus-dominant microbiota.

[0371] In certain implementations, treatment selection, timing, and duration may be guided by the output of the predictive model, including probability thresholds or risk stratification categories.

[0372] Clinical deployment: In certain embodiments, the predictive model is deployed within a clinical or healthcare-associated system for real-time or near real-time assessment of patient microbiome data. Biological samples, including vaginal swabs, may be collected in a clinical setting or by the patient in an at-home setting and processed using sequencing or other microbial profiling techniques.

[0373] Resulting data may be input into the predictive model, which generates a risk score or classification indicating likelihood of impending bacterial vaginosis. This output may be communicated to a healthcare provider through an electronic health record (EHR) system, clinical decision support tool, or dedicated software interface.

[0374] In some embodiments, the system may generate automated alerts or recommendations for intervention when a predefined risk threshold is exceeded. In other embodiments, results may be delivered directly to patients via a mobile application or digital health platform, optionally with guidance regarding follow-up care or treatment.

[0375] The system may further include periodic or longitudinal monitoring, enabling repeated sampling and dynamic risk assessment over time.

[0376] References

[0377] 1. Hillier S, Marrazzo J, Holmes K. Bacterial vaginosis. In: Bacterial vaginosis In: Holmes KK, Sparling PF, Mardh P-A, et al, editors. 4th ed. New York, NY: McGraw-Hill; 2008. p. 737-68.Docket No. 2932719-000286-W01

[0378] Filed: March 27, 2026 2. Srinivasan S, Fredricks DN. The human vaginal bacterial biota and bacterial vaginosis.

[0379] Interdiscip Perspect Infect Dis. 2008;2008:750479.

[0380] 3. Koumans EH, Sternberg M, Bruce C, McQuillan G, Kendrick J, Sutton M, et al. The Prevalence of Bacterial Vaginosis in the United States, 2001-2004; Associations With Symptoms, Sexual Behaviors, and Reproductive Health. Sex Transm Dis. 2007 Nov;34(l l):864-9.

[0381] 4. Cohen CR, Lingappa JR, Baeten JM, Ngayo MO, Spiegel CA, Hong T, et al. Bacterial vaginosis associated with increased risk of female-to-male HIV-1 transmission: a prospective cohort analysis among African couples. PLoS Med. 2012;9(6):el001251. 5. Eschenbach DA. Bacterial vaginosis and anaerobes in obstetric-gynecologic infection. Clin Infect Dis Off Publ Infect Dis Soc Am. 1993 June;16 Suppl 4:S282-287.

[0382] 6. Muzny CA, Lensing SY, Aaron KJ, Schwebke JR. Incubation period and risk factors support sexual transmission of bacterial vaginosis in women who have sex with women. Sex Transm Infect. 2019 Nov;95(7):511-5.

[0383] 7. Bradshaw CS, Walker SM, Vodstrcil LA, Bilardi JE, Law M, Hocking JS, et al. The influence of behaviors and relationships on the vaginal microbiota of women and their female partners: the WOW Health Study. J Infect Dis. 2014 May 15;209(10): 1562-72. 8. Vodstrcil LA, Plummer EL, Fairley CK, Hocking JS, Law MG, Petoumenos K, et al. Male- Partner Treatment to Prevent Recurrence of Bacterial Vaginosis. N Engl J Med. 2025 Mar 6;392(10):947-57.

[0384] 9. Bradshaw CS, Plummer EL, Muzny CA, Mitchell CM, Fredricks DN, Herbst-Kralovetz MM, et al. Bacterial vaginosis. Nat Rev Dis Primer. 2025 June 19; 11(1):43.

[0385] 10. Muzny CA, Taylor CM, Swords WE, Tamhane A, Chattopadhyay D, Cerca N, et al. An Updated Conceptual Model on the Pathogenesis of Bacterial Vaginosis. J Infect Dis. 2019 Sept 26;220(9): 1399-405.

[0386] 11. Nelson DE, Van Der Pol B, Dong Q, Revanna KV, Fan B, Easwaran S, et al. Characteristic male urine microbiomes associate with asymptomatic sexually transmitted infection. PloS One. 2010 Nov 24;5(1 l):el4116.

[0387] 12. Schwebke JR, Muzny CA, Josey WE. Role of Gardnerella vaginalis in the pathogenesis of bacterial vaginosis: a conceptual model. J Infect Dis. 2014 Aug 1 ;210(3):338- 43.Docket No. 2932719-000286-W01

[0388] Filed: March 27, 2026 Pavlova ST, Kilic AO, Mou SM , Tao L. Phage infection in vaginal lactobacilli: an in vitro study. Infect Ob stet Gynecol. 1997;5(l):36-44.

[0389] Blackwell AL. Vaginal bacterial phaginosis? Sex Transm Infect. 1999 Oct;75(5):352-3. Tao L, Pavlova SI, Mou SM, Ma WG, Kilic AO. Analysis of lactobacillus products for phages and bacteriocins that inhibit vaginal lactobacilli. Infect Obstet Gynecol.

[0390] 1997;5(3):244-51.

[0391] Pavlova SI, Tao L. Induction of vaginal Lactobacillus phages by the cigarette smoke chemical benzo[a]pyrene diol epoxide. MutatRes. 2000 Mar 3;466(l):57-62.

[0392] Lambert JA, John S, Sobel JD, Akins RA. Longitudinal Analysis of Vaginal Microbiome Dynamics in Women with Recurrent Bacterial Vaginosis: Recognition of the Conversion Process. Schlievert PM, editor. PLoS ONE. 2013 Dec 20;8(12):e82599.

[0393] Marcos-Zambrano LJ, Karaduzovic-Hadziabdic K, Loncar Turukalo T, Przymus P, Trajkovik V, Aasmets O, et al. Applications of Machine Learning in Human Microbiome Studies: A Review on Feature Selection, Biomarker Identification, Disease Prediction and Treatment. Front Microbiol. 2021 Feb 19;12:634511.

[0394] Ditzler G, Polikar R, Rosen G. Multi-Layer and Recursive Neural Networks for Metagenomic Classification. IEEE Trans Nanobioscience. 2015 Sept;14(6):608-16. Ali A. Artificial Neural Network (ANN) [Internet]. Medium. 2019 [cited 2021 June 16], Available from: https: / / medium.com / machine-learning-researcher / artificial-neural-network-ann-4481 fa33 d85 a

[0395] Muzny CA, Elnaggar JH, Sousa LGV, Lima A, Aaron KJ, Eastlund IC, et al. Microbial interactions among Gardnerella , Prevotella and Fannyhessea prior to incident bacterial vaginosis: protocol for a prospective, observational study. BMJ Open. 2024 Feb;14(2):e083516.

[0396] George SD, Amerson-Brown MH, Sousa LGV, Carter TM, Rinehart AH, Riegler AN, et al. Investigating Bacterial Vaginosis Pathogenesis Using Peptide Nucleic Acid-Fluorescence In Situ Hybridization With a Focus on the Roles of Gardnerella Species, Prevotella bivia , and Fannyhessea vaginae. Open Forum Infect Dis. 2025 Aug 29;12(9):ofaf556.Docket No. 2932719-000286-W01

[0397] Filed: March 27, 2026 23. Muzny CA, Blanchard E, Taylor CM, Aaron KJ, Talluri R, Griswold ME, et al.

[0398] Identification of Key Bacteria Involved in the Induction of Incident Bacterial Vaginosis: A Prospective Study. J Infect Dis. 2018 Aug 14;218(6):966-78.

[0399] 24. Lundberg S, Lee SI. A Unified Approach to Interpreting Model Predictions [Internet], arXiv; 2017 [cited 2025 Mar 23], Available from: https: / / arxiv.org / abs / 1705.07874 25. Amabebe E, Anumba DOC. The Vaginal Microenvironment: The Physiologic Role of Lactobacilli. Front Med. 2018 June 13;5: 181.

[0400] 26. Gajer P, Brotman RM, Bai G, Sakamoto J, Schutte UME, Zhong X, et al. Temporal dynamics of the human vaginal microbiota. Sci Transl Med. 2012 May 2;4(132):132ra52.

[0401] 27. Gajer P, Brotman RM, Bai G, Sakamoto J, Schutte UME, Zhong X, et al. Temporal dynamics of the human vaginal microbiota. Sci Transl Med. 2012 May 2;4(132):132ra52.

[0402] 28. Ravel J, Gajer P, Abdo Z, Schneider GM, Koenig SSK, McCulle SL, et al. Vaginal microbiome of reproductive-age women. Proc Natl Acad Sci U S A. 2011 Mar 15; 108 Suppl 1:4680-7.

[0403] 29. Muzny CA, Blanchard E, Taylor CM, Aaron KJ, Talluri R, Griswold ME, et al.

[0404] Identification of Key Bacteria Involved in the Induction of Incident Bacterial Vaginosis: A Prospective Study. J Infect Dis. 2018 Aug 14;218(6):966— 78.

[0405] 30. Tettamanti Boshier FA, Srinivasan S, Lopez A, Hoffman NG, Proll S, Fredricks DN, et al.

[0406] Complementing 16S rRNA Gene Amplicon Sequencing with Total Bacterial Load To Infer Absolute Species Concentrations in the Vaginal Microbiome. Caporaso JG, editor. mSystems. 2020 Apr 28;5(2):e00777-19.

[0407] 31. Zheng J, Wittouck S, Salvetti E, Franz CMAP, Harris HMB, Mattarelli P, et al. A taxonomic note on the genus Lactobacillus: Description of 23 novel genera, emended description of the genus Lactobacillus Beijerinck 1901, and union of Lactobacillaceae and Leuconostocaceae. Int J SystEvol Microbiol. 2020 Apr l;70(4):2782-858.

[0408] 32. Hiitt P, Lapp E, Stsepetova J, Smidt I, Taelma H, Borovkova N, et al. Characterisation of probiotic properties in human vaginal lactobacilli strains. Microb Ecol Health Dis. 2016 Aug 12;27: 10.3402 / mehd.v27.30484.

[0409] 33. Pan M, Hidalgo-Cantabrana C, Goh YJ, Sanozky-Dawes R, Barrangou R. Comparative Analysis of Lactobacillus gasseri and Lactobacillus crispatus Isolated From Human Urogenital and Gastrointestinal Tracts. Front Microbiol. 2020 Jan 22; 10:3146.Docket No. 2932719-000286-W01

[0410] Filed: March 27, 2026 34. Machado A, Cerca N. Influence of Biofilm Formation by Gardnerella vaginalis and Other Anaerobes on Bacterial Vaginosis. J Infect Dis. 2015 Dec 15;212(12): 1856-61.

[0411] 35. Castro J, Machado D, Cerca N. Unveiling the role of Gardnerella vaginalis in polymicrobial Bacterial Vaginosis biofilms: the impact of other vaginal pathogens living as neighbors. ISME J. 2019 May;13(5):1306-17.

[0412] 36. Muzny CA, Laniewski P, Schwebke JR, Herbst-Kralovetz MM. Host-vaginal microbiota interactions in the pathogenesis of bacterial vaginosis: Curr Opin Infect Dis. 2020 Feb;33(l):59-65.

[0413] 37. Condori-Catachura S, Ahannach S, Ticlla M, Kenfack J, Livo E, Anukam KC, et al.

[0414] Diversity in women and their vaginal microbiota. Trends Microbiol. 2025 Jan 20;S0966- 842X(24)00328-7.

[0415] 38. Mancilla V, Jimenez NR, Bishop NS, Flores M, Herbst-Kralovetz MM. The Vaginal Microbiota, Human Papillomavirus Infection, and Cervical Carcinogenesis: A Systematic Review in the Latina Population. J Epidemiol Glob Health. 2024 June;14(2):480-97. 39. Laniewski P, Joe TR, Jimenez NR, Eddie TL, Bordeaux SJ, Quiroz V, et al. Viewing Native American Cervical Cancer Disparities through the Lens of the Vaginal Microbiome: A Pilot Study. Cancer Prev Res Phila Pa. 2024 Nov 4; 17( 11): 525— 38.

[0416] 40. Laniewski P, Herbst-Kralovetz MM. Bacterial vaginosis and health-associated bacteria modulate the immunometabolic landscape in 3D model of human cervix. NPJ Biofilms Microbiomes. 2021 Dec 13;7(1):88.

[0417] 41. Muzny CA, Cerca N, Elnaggar JH, Taylor CM, Sobel JD, Van Der Pol B. State of the Art for Diagnosis of Bacterial Vaginosis. J Clin Microbiol. 2023 Aug 23;61(8):e0083722. 42. Hillier SL, Nugent RP, Eschenbach DA, Krohn MA, Gibbs RS, Martin DH, et al.

[0418] Association between bacterial vaginosis and preterm delivery of a low-birth-weight infant. The Vaginal Infections and Prematurity Study Group. N Engl J Med. 1995 Dec 28;333(26): 1737-42.

[0419] 43. Cohen CR, Wierzbicki MR, French AL, Morris S, Newmann S, Reno H, et al. Randomized Trial of Lactin-V to Prevent Recurrence of Bacterial Vaginosis. N Engl J Med. 2020 May 14;382(20): 1906-15.

[0420] 44. Bosma EF, Mortensen B, DeLong K, Ropke MA, Juel HB, Rich R, et al. Antibiotic-free vaginal microbiota transplantation (VMT) changes vaginal microbiota and immune profileDocket No. 2932719-000286-W01

[0421] Filed: March 27, 2026 in women with asymptomatic dysbiosis - reporting of a randomized, placebo-controlled trial [Internet], Obstetrics and Gynecology; 2024 [cited 2025 Apr 15], Available from: http: / / medrxiv.Org / lookup / doi / 10.l 101 / 2024.06.25.24309408

[0422] 45. Amsel R, Totten PA, Spiegel CA, Chen KC, Eschenbach D, Holmes KK. Nonspecific vaginitis. Diagnostic criteria and microbial and epidemiologic associations. Am J Med.

[0423] 1983 Jan;74(l): 14-22.

[0424] 46. Nugent RP, Krohn MA, Hillier SL. Reliability of diagnosing bacterial vaginosis is improved by a standardized method of gram stain interpretation. J Clin Microbiol. 1991 Feb; 29(2): 297-301.

[0425] 47. Workowski KA, Bachmann LH, Chan PA, Johnston CM, Muzny CA, Park I, et al. Sexually Transmitted Infections Treatment Guidelines, 2021. MMWR Recomm Rep. 2021 July 23;70(4): 1-187.

[0426] 48. Van Der Pol WJ, Kumar R, Morrow CD, Blanchard EE, Taylor CM, Martin DH, et al. In Silico and Experimental Evaluation of Primer Sets for Species-Level Resolution of the Vaginal Microbiota Using 16S Ribosomal RNA Gene Sequencing. J Infect Dis. 2019 Jan 7;219(2):305— 14.

[0427] 49. Callahan BJ, McMurdie PJ, Rosen MJ, Han AW, Johnson AJA, Holmes SP. DADA2:

[0428] High-resolution sample inference from Illumina amplicon data. Nat Methods. 2016 July;13(7):581-3.

[0429] 50. Quast C, Pruesse E, Yilmaz P, Gerken J, Schweer T, Yarza P, et al. The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Res. 2013 Jan;41 (Database issue):D590-596.

[0430] 51. TensorFlow Developers. TensorFlow [Internet], Zenodo; 2021 [cited 2021 June 13], Available from: https: / / zenodo.org / record / 4758419

[0431] 52. Chollet F, others. Keras [Internet], GitHub; 2015. Available from:

[0432] https: / / github.com / fchollet / keras

[0433] 53. Yu H, Samuels DC, Zhao Y yong, Guo Y. Architectures and accuracy of artificial neural network for disease classification from omics data. BMC Genomics. 2019 Mar 4;20(l):167.

[0434] 54. Hunter JD. Matplotlib: A 2D Graphics Environment. Comput Sci Eng. 2007;9(3):90-5.

Claims

Docket No. 2932719-000286-W01Filed: March 27, 2026 CLAIMSWhat is claimed is:

1. A method comprising,receiving a vaginal sample of a subject;using the vaginal sample of the subject to detect levels of a plurality or organisms in the subject using a sequencing analysis;applying an artificial neural network model to determine a probability that the subject develops bacterial vaginosis over a first period of time using information of the subject’s detected levels;treating the subject for bacterial vaginosis with an antibiotic when the probability is greater than a threshold value.

2. The method of claim 1, wherein the plurality of organisms comprises at least three of L. iners, L. crispatus, L. jensenii, Gardnerella, L. gasseri, Prevotella, Lactobacillus, Streptococcus, Aerococcus, Sneathia, Dialister, BVAB1, Megasphaera, Fannyhessea, Fastidiosipila, Mobiluncus, Parvimonas, Limosilactobacillus, L. intestinalis, and Ligilactobacillus.

3. The method of claim 2, wherein the sequencing analysis comprises applying 16S rRNA sequencing to the vaginal sample.

4. The method of claim 4, wherein the detected levels in the subject comprise relative abundance of each organism of the plurality of organisms in the vaginal sample.

5. The method of claim 1, comprising training the artificial neural network model.

6. The method of claim 5, wherein the training comprises receiving vaginal samples from a plurality of subjects, wherein a first population of the plurality of subjects develops bacterial vaginosis over a second period of time, wherein a second population of the plurality of subjects does not develop bacterial vaginosis over a third period of time.Docket No. 2932719-000286-W01Filed: March 27, 2026 7. The method of claim 6, wherein the training comprises using the vaginal samples of the plurality of subjects to detect levels of the plurality of organisms in the first population and in the second population using the sequencing analysis, wherein the detected levels in the first population and the second population comprise relative abundance of each organism of the plurality of organisms in the vaginal samples.

8. The method of claim 7, wherein the training comprises using the detected levels of the plurality of organisms in the first population and the second population and corresponding information of the disease state to train an artificial neural network for predicting future onset of bacterial vaginosis.

9. The method of claim 8, wherein the artificial neural network model comprises a stochastic gradient descent optimization algorithm.

10. The method of claim 9, wherein the artificial neural network model comprises a cross entropy loss function.

11. The method of claim 10, wherein the artificial neural network model comprises a Rectified Linear Unit (ReLU) activation function.

12. The method of claim 11, wherein the artificial neural network model comprises thirtyeight (38) layer one nodes.

13. The method of claim 12, wherein the artificial neural network model comprises thirty-four (34) layer two nodes.

14. The method of claim 13, wherein the artificial neural network model comprises a drop out rate of zero (0).

15. The method of claim 14, wherein the artificial neural network model comprises a batch size of thirty-two (32).Docket No. 2932719-000286-W01Filed: March 27, 202616. The method of claim 15, wherein the artificial neural network model comprises an L1 / L2 activity regularization of zero (0).

17. The method of claim 16, wherein the artificial neural network model comprises a learning rate of .001.

18. The method of claim 17, wherein the artificial neural network model further comprises a Gaussian noise layer with a standard deviation of approximately 0.003 applied to the input features.

19. The method of claim 18, wherein the antibiotics comprises at least one of metronidazole, clindamycin, and tinidazole.