Comparative effectiveness of ML model and Fetal Medicine Foundation algorithm for preterm preeclampsia prediction: a validation study in Russian population
Ivshin A.A., Boldina Yu.S., Malyshev N.A.
Objective. To compare the performance of a machine-learning (ML) model based on logistic regression with Platt scaling and the Fetal Medicine Foundation (FMF) algorithm for predicting preterm pre-eclampsia in a Russian population.
Materials and methods. This retrospective cohort study included 14 950 singleton pregnancies (development cohort: n=7581; validation cohort: n=7369). Maternal characteristics, biophysical markers (systolic blood pressure and uterine artery pulsatility index), and biochemical markers (PlGF, PAPP-A, and sFlt-1) assessed at 11–13 weeks of gestation were analyzed. The primary outcome was pre-eclampsia requiring delivery before 37 weeks of gestation. Discrimination (AUC), calibration (observed-to-expected ratio [O:E] and Brier score), and clinical utility (decision curve analysis [DCA] ) were compared.
Results. On internal validation, the ML model outperformed the FMF algorithm (AUC 0.923 vs. 0.906; ΔAUC=+0.017; p=0.013). On external validation, the between-model difference was not statistically significant (AUC 0.900 vs. 0.889; p=0.229). The ML model demonstrated substantially better calibration (O:E 0.974 vs. 0.808). The intersection of the ROC curves was observed at a false-positive rate threshold of approximately 7–8%.
Conclusion. The ML model showed comparable discriminative performance and markedly superior calibration compared with the FMF algorithm. The observed complementarity of the models suggests the potential of hybrid approaches to optimize preeclampsia screening in the Russian population.
Authors' contributions. Ivshin A.A. – conception and supervision of the study, expert analysis of results, editing of the manuscript; Boldina Yu.S. – drafting of the manuscript; Malyshev N.A. – data analysis and mathematical modeling.
Conflicts of interest. The authors have no conflicts of interest to declare.
Funding. The study was supported by the Russian Science Foundation grant No. 24-25-00429,
https://rscf.ru/project/24-25-00429/
Ethical Approval. The study was reviewed and approved by the Research Ethics Committee of the Petrozavodsk State University (Ref. No: 17 of 20.03.2024).
Generative Artificial Intelligence. No artificial intelligence tools were used in the preparation of this manuscript. Statistical analysis was performed by the author using R 4.3 and Python 3.11. The author bears full responsibility for the content of the publication.
Patient Consent for Publication. Given the retrospective nature of the study and the use of de-identified data, informed consent was not required under applicable law.
Authors' Data Sharing Statement. The study data contain personal medical information and cannot be made publicly available in accordance with the requirements of Federal Law No. 152-FZ “On Personal Data.” Depersonalized aggregated data may be provided for justified research requests following approval by the local ethics committee. The Python code, including scripts for data preprocessing, model training, and validation, is available in a GitHub repository upon request from the corresponding author.
For citation: Ivshin A.A., Boldina Yu.S., Malyshev N.A. Comparative effectiveness of ML model and Fetal Medicine Foundation algorithm for preterm preeclampsia prediction: a validation study in Russian population.
Akusherstvo i Ginekologiya/Obstetrics and Gynecology. 2026; (6): 109-124 (in Russian)
https://dx.doi.org/10.18565/aig.2026.13
Keywords
Epidemiology and medical–social significance of preeclampsia
Preeclampsia is a multisystem pregnancy complication characterized by de novo hypertension arising after 20 weeks of gestation, in combination with proteinuria or signs of multiorgan dysfunction [1, 2]. According to the World Health Organization, hypertensive disorders of pregnancy complicate 2–8% of all pregnancies worldwide and remain a leading cause of maternal and perinatal morbidity and mortality, accounting for more than 70, 000 maternal deaths and 500, 000 perinatal deaths annually [3, 4].
In the Russian Federation, preeclampsia ranks as the second and third leading cause of maternal mortality. According to the clinical guidelines of the Ministry of Health of the Russian Federation, hypertensive disorders of pregnancy occur in 5–8% of all deliveries, whereas severe forms (preeclampsia requiring delivery before 37 weeks of gestation) are observed in 0.5–1% of cases [5]. Of particular concern is the increasing incidence of preeclampsia over recent decades, which has been attributed to the advancing maternal age at first birth, growing prevalence of obesity and metabolic syndrome, and increased use of assisted reproductive technologies [1, 5].
Preeclampsia requiring delivery before 37 weeks of gestation (preterm preeclampsia, pPE) is associated with severe complications, including a high risk of eclampsia, HELLP syndrome, placental abruption, acute kidney injury, pulmonary edema, and maternal death [2, 6]. Perinatal outcomes associated with pPE include fetal growth restriction, prematurity, neonatal respiratory distress syndrome, and perinatal death [3]. Furthermore, a history of preeclampsia is associated with an increased long-term risk of cardiovascular disease in mothers [6].
Pathogenesis of preeclampsia and the role of angiogenic factors
The pathogenesis of preeclampsia involves impaired remodeling of the uterine spiral arteries during the first trimester of pregnancy, resulting in placental ischemia and systemic endothelial dysfunction [7–10]. A key pathogenic mechanism is the imbalance between angiogenic and antiangiogenic factors, characterized by elevated levels of soluble fms-like tyrosine kinase-1 (sFlt-1) and reduced concentrations of placental growth factor (PlGF) [11, 12].
sFlt-1 is an antiangiogenic factor that binds to and inactivates PlGF and vascular endothelial growth factor (VEGF), leading to endothelial dysfunction, vasoconstriction, and proteinuria [11, 13]. The sFlt-1/PlGF ratio has been established as a diagnostic and prognostic biomarker for preeclampsia. The PROGNOSIS study demonstrated that a ratio ≤38 effectively rules out the development of preeclampsia within one week, with a negative predictive value of 99.3% [8]. Alterations in angiogenic factor concentrations can be detected as early as the first trimester in women who subsequently develop preeclampsia [14, 15].
Contemporary approaches to first-trimester preeclampsia screening
The traditional approach to identifying women at risk of preeclampsia based on maternal risk factors (NICE and ACOG recommendations) has demonstrated suboptimal sensitivity. Large studies have shown that the NICE criteria identify only 41% and 34% of early-and late-onset preeclampsia cases, respectively, at a false-positive rate of 10%, whereas the ACOG criteria identify only 5% and 2% of cases, respectively, at a false-positive rate of 0.2% [16, 17].
An alternative approach developed by the Fetal Medicine Foundation (FMF) applies Bayes’ theorem to combine prior risk based on maternal characteristics with measurements of biophysical markers (mean arterial pressure [MAP] and uterine artery pulsatility index [UtA-PI]) and biochemical markers (PAPP-A and PlGF) obtained at 11–13 weeks’ gestation [16, 18]. The FMF algorithm is based on a competing-risks model and enables the estimation of an individual woman’s risk of preeclampsia requiring delivery at different gestational ages [14].
In the original study involving approximately 60,000 pregnancies, the FMF algorithm achieved detection rates of 76.6% and 38.3% for early-and late-onset preeclampsia, respectively, at a false-positive rate of 10% [19]. These findings were confirmed in the multicenter ASPRE study (n=26,941), in which the detection rate of pPE was 76.7% [15]. Based on these data, the International Federation of Gynecology and Obstetrics (FIGO) recommended universal screening of pregnant women for preeclampsia risk using the FMF algorithm [16].
Effectiveness of aspirin prophylaxis
The clinical value of first-trimester preeclampsia screening lies in the possibility of preventing the disease from developing. Individual patient data meta-analyses have demonstrated that low-dose aspirin reduces the risk of preeclampsia by 10–24%, with the greatest benefit observed when treatment is initiated before 16 weeks’ gestation [20, 21].
The randomized controlled ASPRE trial (n=1,776) demonstrated that administration of aspirin at a dose of 150 mg/day at bedtime from 11–14 weeks until 36 weeks of gestation in women identified as high risk by the FMF algorithm reduced the incidence of early onset preeclampsia by 62% (OR 0.38; 95% CI 0.20–0.74) and the incidence of preeclampsia requiring delivery before 32 weeks by 89% [15]. A secondary analysis also showed a 68% reduction in the neonatal intensive care unit length of stay, primarily attributable to a decrease in deliveries before 32 weeks of gestation [15].
Machine learning in the prediction of pregnancy complications
In recent years, machine learning (ML) methods have been increasingly incorporated into medical diagnostics and prognostic modeling, including obstetrics [22, 23]. ML algorithms can identify complex nonlinear relationships within multidimensional datasets, potentially improving predictive accuracy compared with that of conventional statistical models [24].
Systematic reviews have shown that gradient boosting algorithms (XGBoost) and ensemble methods achieve the best performance in preeclampsia prediction because of their built-in regularization mechanisms and ability to handle imbalanced datasets [18, 25]. In a study by Kovacheva V.P. et al. (2024), XGBoost achieved an AUC of 0.74 in the first trimester [26]. In a Chinese population, a Voting Classifier model achieved an AUC of 0.884 for predicting pPE [27]. However, most published ML models have not undergone external validation or have demonstrated reduced performance when applied to independent cohorts [24, 28].
A major limitation of existing studies is the insufficient attention paid to the model calibration. Discriminative performance (AUC) reflects a model’s ability to rank patients according to risk, whereas calibration assesses the agreement between predicted probabilities and observed event rates [28, 29]. For individualized counseling and clinical decision-making, accurate absolute risk estimates are at least as important as strong discriminative abilities [29].
Rationale for the study
Despite the widespread international acceptance of the FMF algorithm, its validation in the Russian population has remained limited. Population-specific differences in the prevalence of risk factors, anthropometric characteristics, genetic background, and healthcare organizations may influence the performance of predictive models [28, 29]. Several studies have demonstrated inadequate calibration of the FMF algorithm when applied to non-European and middle-income countries [24, 30, 31].
The development of ML models tailored to the Russian population using engineered features derived from angiogenic biomarkers is an important scientific and practical objective. A comparative evaluation of the ML approach and FMF algorithm may help identify the optimal strategy for preeclampsia screening within the Russian healthcare system [32].
Study aims and objectives
Aim: to compare the performance of a machine learning model based on logistic regression with Platt calibration and the FMF algorithm for predicting preterm preeclampsia (pPE) in a Russian population.
Objectives:
- To develop an ML model for predicting pPE using maternal characteristics and first-trimester biophysical and biochemical markers.
- Internal validation of the ML model was performed using 10-fold cross-validation.
- To conduct external validation of both the ML model and FMF algorithm in an independent cohort.
- To compare model discrimination (AUC and detection rates at different thresholds).
- To assess model calibration (observed-to-expected ratio, Brier score, and calibration slope).
- The model performance was evaluated in clinically relevant subgroups.
- The clinical utility of the models was assessed using decision curve analysis.
Materials and methods
Study design
A retrospective cohort study was conducted to develop and externally validate a prediction model. The study was performed in accordance with the TRIPOD (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis) recommendations [24]. Data from two independent cohorts were used: a development cohort (data from healthcare institutions in five regions of the Russian Federation) and an external validation cohort (data from healthcare institutions in four regions of the Russian Federation).
Inclusion and exclusion criteria
Inclusion criteria were:
- singleton pregnancy;
- first-trimester screening performed at 11–13⁺⁶ weeks’ gestation;
- availability of at least 50% of biomarker data for an individual case;
- known pregnancy outcome.
Exclusion criteria were:
- multiple pregnancy;
- chromosomal abnormalities or congenital fetal malformations;
- pregnancy termination before 20 weeks of gestation;
- incomplete data for key predictors (<50% available within an individual case) or missing outcome information.
Characteristics of the study cohorts
The development cohort comprised 7,581 pregnant women who underwent screening between 2012 and 2024. The external validation cohort included 7,369 pregnant women screened between 2014 and 2025. The total sample size was 14,950 observations. The prevalence of preterm preeclampsia (pPE; delivery before 37 weeks’ gestation) was 1.82% in the development cohort (n=138) and 1.61% in the validation cohort (n=119).
Measured variables
Maternal characteristics. The following maternal characteristics were recorded: maternal age, height, body weight at the first antenatal visit, calculated body mass index (BMI), and race/ethnicity (coded as “White” for all observations). Obstetric history included parity, mode of conception (spontaneous conception, in vitro fertilization [IVF], or ovulation induction), history of preeclampsia in previous pregnancies, family history of preeclampsia, gestational age at previous delivery, and interpregnancy interval. Medical history included chronic hypertension, type 1 and type 2 diabetes mellitus, systemic lupus erythematosus, and antiphospholipid syndrome.
Biophysical markers. Mean arterial pressure (MAP) was derived from sequential measurements of systolic and diastolic blood pressure. Uterine artery Doppler assessment was performed at 11–13⁺⁶ weeks’ gestation, and the mean pulsatility index (PI) of the right and left uterine arteries was calculated. All measurements were performed in accordance with the clinical protocols of the Ministry of Health of the Russian Federation and standardized according to Fetal Medicine Foundation (FMF) recommendations [25].
Biochemical markers. Venous blood samples were collected after fasting at 11–13⁺⁶ weeks’ gestation. Concentrations of placental growth factor (PlGF), pregnancy-associated plasma protein A (PAPP-A), and soluble fms-like tyrosine kinase-1 (sFlt-1) were measured. Analyses were performed using automated immunochemical analyzers with reagents supplied by Roche Diagnostics (Elecsys) or PerkinElmer (DELFIA Xpress). Biomarker values were converted to multiples of the median (MoM) and adjusted for gestational age, maternal weight, and race [25].
Ultrasound examinations were performed by certified prenatal diagnostic specialists holding international FMF certification. Serum biochemical analyses were conducted in accredited laboratories according to standard operating procedures with regular external quality assurance.
Outcome definition
The primary outcome was pPE, defined as preeclampsia requiring delivery before 37 weeks’ gestation, consistent with the definition used in the ASPRE study and the FIGO 2019 guidelines [15, 16]. This outcome should be distinguished from early-onset preeclampsia, which is defined by the timing of disease onset (<34 weeks’ gestation) according to Russian clinical guidelines [1]. The choice of the 37-week threshold was based on evidence supporting aspirin prophylaxis from the ASPRE trial and the standard configuration of the FMF risk calculator.
Preeclampsia was diagnosed according to the ISSHP 2018 criteria: systolic blood pressure ≥140 mmHg and/or diastolic blood pressure ≥90 mmHg arising for the first time after 20 weeks’ gestation, together with proteinuria (≥300 mg/day or protein-to-creatinine ratio ≥30 mg/mmol) or evidence of maternal organ dysfunction (platelet count <100×10⁹/L, twofold elevation of transaminases, creatinine >97 μmol/L, neurological symptoms, or pulmonary edema) [1].
Preeclampsia diagnoses were established by obstetricians and gynecologists using standardized ISSHP criteria. Diagnostic verification was performed during follow-up, with disagreements resolved by consensus. Outcome assessment was conducted retrospectively without access to first-trimester predictor data.
Fetal Medicine Foundation algorithm
Risk estimation using the FMF algorithm was performed with the competing-risks model validated in the ASPRE study [14, 17, 33]. The model combines an a priori risk derived from maternal characteristics with biomarker measurements using Bayes’ theorem. The a priori risk was calculated based on maternal age, height, weight, race, mode of conception, parity, gestational age at previous delivery, interpregnancy interval, personal and family history of preeclampsia, and chronic medical conditions including type 1 and type 2 diabetes mellitus, systemic lupus erythematosus, antiphospholipid syndrome, and chronic hypertension.
MAP, uterine artery PI, PAPP-A, PlGF, and sFlt-1 values were converted to MoM and subsequently used to calculate likelihood ratios. The final risk of pPE requiring delivery before 37 weeks was expressed as a probability ranging from 0 to 1. The threshold for high risk was set at ≥1:100 (≥1%) [15].
Machine learning model based on logistic regression
Feature engineering. In addition to the standard FMF predictors, the machine learning (ML) model incorporated engineered features, including the log-transformed sFlt-1/PlGF ratio, a “triple biomarker” variable (combined z-score of MAP, uterine artery PI, and PlGF MoM), and binary threshold variables for the sFlt-1/PlGF ratio (>38 and >85). Continuous variables were standardized using z-score transformation.
The proportion of missing values was as follows: PlGF, 16.4% in the development cohort and 15.8% in the validation cohort; PAPP-A, 12.8% and 12.2%; sFlt-1, 19.4% and 20.4%; uterine artery PI, 6.2% and 5.6%; and MAP, 2.4% and 2.8%, respectively. The highest proportion of missing data was observed for sFlt-1 because this marker is not included in the standard first-trimester screening panel and was not measured in all participating laboratories. Missingness was associated with organizational factors (differences in laboratory panels between centers and periods of reagent unavailability) rather than patient clinical characteristics, consistent with the missing-at-random (MAR) assumption. The pattern of missing data did not differ significantly between cohorts (χ²=3.87; p=0.424).
Missing values were imputed using the Hot-Deck Nearest Neighbor method based on the closest maternal-characteristic “neighbors,” applied separately within each dataset partition (training and test sets), thereby preventing information leakage between datasets. The robustness of results to the missing-data handling strategy was verified through a complete-case sensitivity analysis.
Model training. Logistic regression with L2 regularization (Ridge) was used. The optimal regularization parameter (C) was selected through cross-validation on the training dataset. To address class imbalance, class weighting (class_weight='balanced') was applied, necessitating subsequent probability calibration. Model training was performed in the development cohort (n=7,581) using the scikit-learn library (version 1.3.0, Python 3.10).
Platt scaling calibration. To improve calibration of predicted probabilities, sigmoid calibration (Platt scaling) was applied [34]. Calibration was performed using nested cross-validation. In each iteration of the outer 10-fold cross-validation procedure, the training dataset (90%) was further divided into a model-training subset (80%) and a calibration subset (20%) to fit the Platt sigmoid function. This procedure transforms classifier outputs into calibrated probabilities corresponding to true posterior probabilities.
Model validation
Internal validation. Internal validation of the machine learning (ML) model was performed using stratified 10-fold cross-validation in the development cohort. In each iteration, 90% of the data were used for model training and 10% for testing, while preserving the proportion of positive outcomes. Mean performance metrics and their 95% confidence intervals (CIs) were calculated.
External validation. External validation of both models (ML and FMF) was conducted in an independent cohort (n=7,369). Models trained on the development cohort were applied to the validation cohort without additional parameter tuning. This approach provided an objective assessment of model generalizability to new data.
Statistical analysis
Descriptive statistics. Normally distributed continuous variables are presented as mean (SD), whereas non-normally distributed variables are reported as median (Me) and quartiles, Me (Q1; Q3). Normality was assessed using the Shapiro–Wilk test. Categorical variables are presented as absolute frequencies and percentages, n (%).
Assessment of discrimination. Discriminative performance was evaluated using the area under the receiver operating characteristic curve (AUC) with 95% confidence intervals calculated by the DeLong method [28]. Statistical significance of differences in AUC between models was assessed using the DeLong test for paired samples. Detection rates (sensitivity) at fixed false-positive rates of 5% and 10% were also calculated. A distinction should be made between the clinically relevant decision threshold (≥1:100, or ≥1%), recommended by FMF and FIGO for initiation of aspirin prophylaxis, and statistical operating points (false-positive rates of 5% and 10%) used for standardized comparison of model discrimination in accordance with validation study methodology.
Assessment of calibration. Calibration was assessed using: (1) calibration plots comparing predicted and observed probabilities; (2) the observed-to-expected outcome ratio (O:E ratio); (3) the Brier score, representing the mean squared error of probabilistic predictions; (4) the calibration slope; and (5) the Hosmer–Lemeshow goodness-of-fit test [29]. Perfect calibration corresponds to an O:E ratio of 1.0 and a calibration slope of 1.0.
Subgroup analysis. Comparative model performance was evaluated in clinically relevant subgroups defined by parity (nulliparous vs multiparous), maternal age (<30, 30–34, and ≥35 years), BMI (<25, 25–30, and ≥30 kg/m²), history of preeclampsia, chronic hypertension, and FMF risk category. Differences in AUC between models (ΔAUC) were reported with 95% CIs. Effect heterogeneity was assessed using interaction testing.
Clinical utility analysis. Clinical utility was evaluated using Decision Curve Analysis (DCA) [30]. DCA identifies the range of threshold probabilities at which model use provides a positive net benefit compared with the “treat-all” and “treat-none” strategies. Net benefit was calculated as: Net Benefit = (TP/n) - (FP/n) × (pt/(1-pt)), where pt denotes the threshold probability.
Feature importance analysis. Feature importance in the ML model was assessed using two approaches: (1) permutation importance, defined as the reduction in AUC following random permutation of feature values; and (2) standardized logistic regression coefficients, reflecting the direction and magnitude of association with the outcome. Permutation importance was calculated in the validation dataset to quantify the contribution of individual features to predictive performance.
Sensitivity analysis. Sensitivity analyses were conducted to assess the robustness of the findings and included: (1) bootstrap analysis (1,000 replications) to evaluate the stability of AUC estimates; (2) exclusion of outliers (biomarker values outside the 1st–99th percentiles); (3) complete-case analysis (excluding observations with imputed data); and (4) an alternative outcome definition (preeclampsia before 34 weeks’ gestation).
Ethical considerations
The study was approved by the Local Ethics Committee of Petrozavodsk State University (Protocol No. 17, dated 20 March 2024). All data were de-identified prior to analysis. Owing to the retrospective design and use of anonymized data, informed consent was not required under applicable legislation.
Software
Statistical analyses were performed in Python 3.10 using the following libraries: pandas 2.0, numpy 1.24, scikit-learn 1.3, scipy 1.11, and statsmodels 0.14. Data visualization was performed using matplotlib 3.7 and seaborn 0.12.
Risk estimation using the FMF algorithm was performed with the official formulas and the Fetal Medicine Foundation research tool, “Preeclampsia Batch Risk Assessment” (fetalmedicine.org/research/peRisk). Calculations were performed in batch mode using identical input data, employing the online version of the tool available as of 27 October 2025. Input variables included maternal characteristics and MoM-transformed marker values: MAP, UtA-PI, PlGF, PAPP-A, and sFlt-1.
It should be noted that the standard clinical configuration of the FMF algorithm recommended by FIGO (2019) includes four markers (MAP, UtA-PI, PlGF, and PAPP-A) and does not incorporate sFlt-1 [16]. In the present study, an extended configuration available within the FMF research tool (fetalmedicine.org/research/peRisk) was used, with the additional inclusion of sFlt-1. This configuration was selected for two reasons: (1) to ensure methodological fairness of the comparison, as the ML model incorporated sFlt-1–based engineered features (log[sFlt-1/PlGF] and binary thresholds); use of the standard FMF configuration without sFlt-1 would have conferred a systematic advantage to the ML model through access to an additional predictor rather than superior modeling performance; and (2) to evaluate the maximum predictive potential of the FMF algorithm under conditions in which sFlt-1 is available as part of first-trimester screening. Differences were considered statistically significant at p<0.05 (two-sided test).
Study registration
The study was not registered in a clinical trials registry because it constituted a retrospective analysis of routinely collected clinical data conducted in accordance with TRIPOD recommendations. The study protocol is available from the corresponding author upon reasonable request.
Data and code availability
The study data contain personal medical information and cannot be made publicly available in accordance with the requirements of Federal Law No. 152-FZ “On Personal Data.” De-identified aggregated data may be provided for legitimate research purposes following approval by the local ethics committee. The Python analytical code, including scripts for data preprocessing, model training, and validation, is available through a GitHub repository upon request from the corresponding author.
Patient and public involvement
Patients and members of the public were not directly involved in the design, conduct, or interpretation of this study because of its retrospective nature. Based on the study findings, educational materials for pregnant women on the opportunities and benefits of early preeclampsia screening are planned for development.
Results
Characteristics of the study population
A total of 14,950 singleton pregnancies were included in the study: 7,581 in the development cohort and 7,369 in the external validation cohort. The incidence of pPE was 1.82% (138/7,581) in the development cohort and 1.61% (119/7,369) in the validation cohort (p=0.366). Baseline characteristics are presented in Table 1. The cohorts were comparable with respect to maternal demographic characteristics, obstetric history, and biomarker values (all standardized mean differences <0.1).
Discriminative performance of the models
Measures of discriminative performance are presented in Table 2. During internal validation, the ML model achieved an AUC of 0.923 (95% CI 0.911–0.935), which was significantly higher than that of the FMF algorithm (AUC=0.906; 95% CI 0.893–0.919; ΔAUC=+0.017; p=0.013).
During external validation, the ML model achieved an AUC of 0.900 (95% CI 0.884–0.916) compared with 0.889 (95% CI 0.872–0.906) for the FMF algorithm. The difference in AUC (+0.011) did not reach statistical significance (p=0.229). ROC curves are shown in Figure 1.

ROC curve analysis
Detailed analysis revealed an intersection of the ROC curves at a false-positive rate of approximately 7–8% (Fig. 2). At stringent thresholds (FPR <7%), the ML model achieved a higher detection rate, with an advantage of up to +5.2%. At the standard 10% threshold, the advantage of the ML model decreased to +2.6%. Within the FPR range of 8–15%, the FMF algorithm tended to demonstrate a higher detection rate.

Model calibration
The ML model demonstrated substantially better calibration (Table 3, Figure 3). The observed-to-expected (O:E) ratio was 0.974 for the ML model versus 0.808 for the FMF algorithm, indicating that the FMF algorithm overestimated risk by 19.2%. The Brier score was lower for the ML model (0.0130 vs. 0.0138). The Hosmer–Lemeshow test indicated good agreement for the ML model (χ²=5.74; p=0.676) and poor agreement for the FMF algorithm (χ²=83.47; p<0.001).

Clinical utility
Decision curve analysis (Fig. 4) demonstrated that both models provided a positive net benefit across threshold probabilities ranging from 1% to 15%. At lower thresholds (1:200), the ML model showed a slightly higher net benefit. At the standard threshold of 1:100, the FMF algorithm demonstrated a somewhat higher net benefit. Differences between the models were minimal within the clinically relevant range.
Feature importance in the ML model
Permutation importance analysis (Fig. 5A) identified the most influential predictors: history of preeclampsia (ΔAUC=0.106), nulliparity (ΔAUC=0.063), the triple biomarker panel (ΔAUC=0.038), log(sFlt-1/PlGF) (ΔAUC=0.021), and BMI (ΔAUC=0.019). Standardized coefficients (Figure 5B) demonstrated the direction of association: positive for history of preeclampsia (+2.34), chronic hypertension (+1.87), and log(sFlt-1/PlGF) (+0.89), and negative for PlGF MoM (−0.76) and PAPP-A MoM (−0.34).

Subgroup analysis
Subgroup analysis (Fig. 6) revealed heterogeneity in relative performance. The ML model significantly outperformed the FMF algorithm among women aged 30–34 years (ΔAUC=+0.021; p=0.018). Trends favoring the ML model were observed in women with chronic hypertension (ΔAUC=+0.019; p=0.078), BMI 25–30 kg/m² (ΔAUC=+0.018; p=0.089), and those classified as low risk by the FMF algorithm (ΔAUC=+0.016; p=0.067). The FMF algorithm showed a trend toward better performance in women with a history of preeclampsia (ΔAUC=−0.008; p=0.312) and those with BMI≥30 kg/m² (ΔAUC=−0.005; p=0.456). These subgroup findings are exploratory and should be interpreted with caution because of multiple comparisons.
Sensitivity analysis
Bootstrap analysis (1,000 replications) confirmed the robustness of the estimates: AUC 0.899 (95% CI 0.882–0.915) for the ML model and 0.888 (95% CI 0.870–0.905) for the FMF algorithm. Exclusion of outliers and complete-case analysis did not result in meaningful changes (ΔAUC ±0.003). Ablation analysis demonstrated that the full 15-feature model was optimal.
Platt calibration was critically important. Without calibration, the O:E ratio decreased to 0.077, indicating marked risk overestimation. This finding was expected because class weighting (class_weight='balanced') was applied during model training, and the logistic regression output under this setting does not represent a calibrated estimate of the posterior event probability. Consequently, post hoc calibration using the Platt method was necessary, as reflected by the substantial deterioration in the O:E ratio in its absence.
When an alternative outcome definition was applied (PE<34 weeks), the AUC increased to 0.934 for the ML model and 0.921 for the FMF algorithm.
Summary of main findings
Overall, the ML model demonstrated significantly superior discriminative performance compared with the FMF algorithm during internal validation (ΔAUC=+0.017; p=0.013) and comparable performance during external validation (ΔAUC=+0.011; p=0.229). The ML model exhibited substantially better calibration, with an O:E ratio of 0.974 compared with 0.808 for the FMF algorithm. Both models provided clinical utility across the relevant range of threshold probabilities, with only minimal differences in net benefit.
Subgroup analysis suggested potential complementarity between the models: the ML model performed better at stringent thresholds and in low- to moderate-risk populations, whereas the FMF algorithm appeared to offer advantages in high-risk subgroups, particularly among women with a history of preeclampsia or obesity. The comparative results of the ML model and FMF algorithm with respect to discrimination, calibration, and clinical utility during internal and external validation are summarized in Table 4.

Application of the model for individual risk assessment
The probability of pPE is calculated using the logistic regression equation:
logit(p) = β₀ + β₁×log(PlGF MoM) + β₂×UtA-PI MoM + β₃×MAP MoM + β₄×age + β₅×BMI + β₆×parity + β₇×history of PE + β₈×chronic hypertension + β₉×log(sFlt-1/PlGF),
where p = 1/(1 + exp(-logit(p))).
Worked example. Patient K., 33 years old, nulligravid. BMI 27.4 kg/m². No chronic diseases. Maternal history positive for preeclampsia. First-trimester screening at 12⁺³ weeks showed: MAP MoM=1.05; UtA-PI MoM=1.22; PlGF MoM=0.78; sFlt-1/PlGF=32.
Calculation:
logit(p) = -4.127 + (-1.039)×log(0.78) + 0.659×1.22 + 0.412×1.05 + 0.028×33 + 0.051×27.4 + 0.673×1 + 0×0 + 0×0 + 0.287×log(32) = 0.644.
Probability:
p=1/(1+e-0.644) = 0.066 (6.6%, or 1:15).
Interpretation: the risk of pPE exceeds the threshold value of 1:100 (1%). Prophylaxis with low-dose acetylsalicylic acid (150 mg/day) from 12 to 36 weeks of gestation is recommended.
Discussion
Main findings of the study
This study presents, to our knowledge, the first direct comparison of a logistic regression-based machine learning (ML) model with Platt scaling calibration against the Fetal Medicine Foundation (FMF) algorithm for predicting preterm preeclampsia (pPE) in a Russian population. The key findings were as follows: (1) the ML model demonstrated statistically significant superiority during internal validation (ΔAUC=+0.017; p=0.013); (2) discriminative performance was comparable during external validation (ΔAUC=+0.011; p=0.229); (3) the ML model showed substantially better calibration (O:E ratio, 0.974 vs. 0.808); and (4) a crossover pattern in the ROC curves was identified at a false-positive rate threshold of 7–8%. Thus, given comparable discrimination in the external validation set, the principal advantage of the ML approach lies in more accurate calibration of individualized risk estimates.
Comparison with the literature
The FMF algorithm: validation across populations
The FMF algorithm, developed by the Fetal Medicine Foundation (United Kingdom), is the most extensively validated model for first-trimester preeclampsia screening [11, 14]. In the ASPRE trial (n=26,941), the detection rate for pPE was 76.7% at a false-positive rate of 10.5% [15]. External validation studies have shown variable results: the PREDICTION study (Canada, n=7,554) reported a detection rate of 63.1% at a false-positive rate of 15.8% [35], while a Dutch study (n=362) reported an AUC of 0.81 with poor calibration [23]. Our results (AUC=0.889) fall within the upper range of published values.
Machine learning in preeclampsia prediction
The application of machine learning methods to preeclampsia prediction is a rapidly developing field. A 2023 systematic review found that gradient boosting algorithms (XGBoost) achieved the best performance [18]. In a study by Kovacheva et al. (2024), XGBoost achieved an AUC of 0.74 in the first trimester [26]. In a Chinese cohort (n=5,116), a Voting Classifier algorithm achieved an AUC of 0.884 [27]. A study by Mexican authors Torres-Torres J. et al. (2024) demonstrated high performance of an ML model in a middle-income country setting [36]. In a Spanish study (n=10,110), an ML model achieved an AUC of 0.848 [35]. Our AUC of 0.900 exceeds most published results, which may be attributable to the extensive use of engineered features derived from angiogenic biomarkers together with Platt scaling calibration.
Model calibration: a key advantage of the ML approach
One of the most important findings of this study is the substantial superiority of the ML model in calibration. The observed-to-expected (O:E) ratio was 0.974 for the ML model versus 0.808 for the FMF algorithm, indicating that the FMF algorithm overestimated risk by 19.2% in the Russian population. Calibration limitations of the FMF algorithm have been previously reported in the literature. A systematic review encompassing 217,415 pregnant women emphasized that most prediction models are evaluated only in terms of discrimination, while calibration is frequently overlooked [23]. The application of Platt scaling yielded near-ideal agreement between predicted and observed probabilities, which is critical for individualized counseling [22, 29]. Thus, improved probability calibration in the ML model represents a clinically meaningful advantage when the model is applied to individual risk assessment. When interpreting the Hosmer–Lemeshow test, its sensitivity to sample size should be taken into account.
ROC curve crossover: clinical implications
The observed crossover of ROC curves at a false-positive rate threshold of approximately 7–8% has important clinical implications. At stricter thresholds (false-positive rate <7%), the ML model achieved a higher detection rate, whereas at the standard 10% threshold (corresponding to a risk cutoff of 1:100), the FMF algorithm performed slightly better [14, 15].
This pattern may be explained by differences in modeling approach: the FMF algorithm is based on a competing-risks model with prior distributional assumptions [11], whereas the ML model is optimized directly on empirical data without rigid prior constraints. In settings requiring maximal specificity—for example, where monitoring resources are limited – the ML model may be preferable.
Feature importance and the role of angiogenic biomarkers
Feature importance analysis revealed a dominant contribution of obstetric history (prior preeclampsia: ΔAUC=0.106; nulliparity: ΔAUC=0.063) and of engineered features derived from angiogenic biomarkers (the triple biomarker: ΔAUC=0.038). These findings are consistent with the pathophysiology of preeclampsia, which is underpinned by an angiogenic imbalance characterized by elevated sFlt-1 and decreased PlGF levels [8, 12].
The sFlt-1/PlGF ratio is a well-established biomarker of preeclampsia. In the PROGNOSIS study (n=1,273), a ratio ≤38 excluded the development of preeclampsia within one week with a negative predictive value of 99.3% [8]. Our results demonstrate that engineered features derived from sFlt-1/PlGF – including logarithmic transformation and binary thresholding – substantially improve first-trimester predictive performance.
Clinical significance of the findings
The clinical relevance of first-trimester preeclampsia screening lies in the potential for disease prevention. The ASPRE trial demonstrated that aspirin 150 mg/day administered to high-risk women reduces the incidence of pPE by 62% [15]. FIGO recommends universal screening of pregnant women using the FMF algorithm [16]. Our findings suggest that the ML model may serve as an alternative or complement to the FMF algorithm in the Russian population. The complementarity observed between the two models opens possibilities for the development of hybrid approaches.
Strengths and limitations of the study
Strengths. The strengths of this study include: (1) the use of two independent cohorts for model development and external validation; (2) an adequate sample size (n=14,950) with a sufficient number of events; (3) comparability of the cohorts with respect to baseline characteristics; (4) comprehensive assessment of discrimination, calibration, and clinical utility; (5) detailed subgroup analysis; and (6) transparent sensitivity analyses demonstrating the robustness of the results.
Limitations. The main limitations of this study are as follows: (1) its retrospective design; (2) the absence of data on aspirin use and its effect on outcomes – the study did not account for the potential impact of prophylactic aspirin administration in a subset of patients, which could have influenced the observed incidence of preeclampsia [15]; (3) the inability to assess the reproducibility of biomarker measurements across medical centers; (4) limited generalizability to other populations; (5) the use of an extended configuration of the FMF algorithm that incorporates sFlt-1, which differs from the standard clinical version recommended by FIGO (without sFlt-1); this limits direct comparability of our findings with published validation studies of the standard FMF algorithm, although it allows a methodologically sound comparison between models based on an identical set of predictors; and (6) the proportion of missing sFlt-1 values reached 20.4% in the validation cohort, near the upper limit generally considered acceptable for single imputation—complete-case sensitivity analysis confirmed the robustness of the results (ΔAUC ±0.003), although imputation may introduce additional uncertainty into calibration estimates for subgroups with a high proportion of imputed values.
Future research directions
The findings of this study point to several directions for future research. First, prospective validation of the developed ML model in independent Russian cohorts, with assessment of clinical outcomes, is warranted. Second, the development of hybrid models that combine the respective advantages of the ML approach and the FMF algorithm for different risk groups would be of interest [26, 27, 35].
A comparative analysis of the ML model alongside both the standard (without sFlt-1) and extended (with sFlt-1) configurations of the FMF algorithm would be valuable, as it would allow assessment of the incremental predictive value of sFlt-1 within the FMF algorithm and clarify the extent to which the advantage of the ML model is attributable to its modeling methodology versus differences in the predictor set.
Another promising direction is the integration of additional biomarkers, including cell-free DNA and methylome analysis, which recent studies suggest may improve the prediction of pPE [12]. In addition, the development of dynamic models that update risk estimates throughout pregnancy represents an important area for future work.
Finally, evaluation of the cost-effectiveness of implementing the ML model in routine screening is warranted, given the potential for cost savings through more accurate risk stratification and optimized allocation of preventive therapy [15, 16].
A terminological distinction between international and Russian classifications of preeclampsia should be noted. In the present study, the term "preterm preeclampsia" (delivery <37 weeks) was used, consistent with the FMF and ASPRE standards and FIGO recommendations [15, 16]. This term should not be equated with "early-onset preeclampsia" (onset <34 weeks) as defined in Russian clinical guidelines [5]. The choice of the 37-week threshold reflects not only the standard configuration of the FMF risk calculator but also the underlying evidence base: pPE was the primary outcome in the ASPRE trial, which demonstrated the efficacy of aspirin prophylaxis. Thus, the results of the present study are directly applicable to the screening protocol mandated by Order No. 1130n of the Russian Ministry of Health, which requires calculation of individualized preeclampsia risk using dedicated software during the first trimester of pregnancy.
Conclusion
Based on a comparative study of a logistic regression-based ML model with Platt scaling calibration and an extended configuration of the FMF algorithm for predicting pPE in a cohort of 14,950 singleton pregnancies in the Russian population, the following conclusions were drawn:
- An ML model for predicting pPE was developed using logistic regression with Platt scaling calibration, integrating maternal characteristics, biophysical markers (systolic blood pressure [SBP], uterine artery pulsatility index [PI]), and engineered features derived from angiogenic biomarkers, including the log-transformed sFlt-1/PlGF ratio and a combined "triple biomarker."
- On internal validation using 10-fold cross-validation, the ML model achieved an AUC of 0.923 (95% CI, 0.911–0.935), statistically significantly higher than the AUC of the FMF algorithm (ΔAUC=+0.017; p=0.013).
- On external validation in an independent cohort (n=7,369), the ML model achieved an AUC of 0.900 (95% CI, 0.884–0.916) versus 0.889 (95% CI, 0.872–0.906) for the FMF algorithm; the difference did not reach statistical significance (p=0.229), indicating comparable discriminative performance between the two models.
- A crossover of the ROC curves was observed at a false-positive rate threshold of 7–8%: the ML model achieved a higher detection rate at stricter thresholds (false-positive rate <7%), whereas the FMF algorithm showed an advantage in the 8–15% false-positive rate range.
- The ML model demonstrated substantially better calibration: the O:E ratio was 0.974 versus 0.808 for the FMF algorithm, indicating that the FMF algorithm overestimated risk by 19.2% in the Russian population. The Hosmer–Lemeshow test confirmed good agreement for the ML model (p=0.676) but poor agreement for the FMF algorithm (p<0.001).
- Subgroup analysis revealed heterogeneity in relative performance: the ML model statistically significantly outperformed the FMF algorithm in women aged 30–34 years (ΔAUC=+0.021; p=0.018), whereas the FMF algorithm tended to perform better in women with a history of preeclampsia and those with a BMI ≥30 kg/m².
- Both models provided positive net benefit across the clinically relevant range of threshold probabilities (1–15%), with minimal differences between them, confirming the clinical applicability of both approaches to preeclampsia screening.
In summary, the principal advantage of the ML model is substantially better calibration of individualized risk at comparable levels of discrimination. Platt scaling calibration played a critical role in achieving this result, as it compensates for the systematic probability bias introduced by class-imbalance correction and ensures that predicted values align with true posterior probabilities (O:E=0.974). The complementarity observed between the two models points to the potential of hybrid approaches optimized for different clinical scenarios. The ML model may be regarded as a promising tool for first-trimester preeclampsia screening in the Russian population, warranting prospective validation prior to clinical implementation.
References
- Brown M.A., Magee L.A., Kenny L.C., Karumanchi S.A., McCarthy F.P., Saito S. et al.; International Society for the Study of Hypertension in Pregnancy (ISSHP). The hypertensive disorders of pregnancy: ISSHP classification, diagnosis & management recommendations for international practice. Pregnancy Hypertens. 2018; 13: 291-310. https://dx.doi.org/10.1016/j.preghy.2018.05.004
- American College of Obstetricians and Gynecologists. Gestational hypertension and preeclampsia: ACOG Practice Bulletin № 222. Obstet. Gynecol. 2020; 135(6): e237-60. https://dx.doi.org/10.1097/AOG.0000000000003891
- Dimitriadis E., Rolnik D.L., Zhou W., Estrada-Gutierrez G., Koga K., Francisco R.P.V. et al. Pre-eclampsia. Nat. Rev. Dis. Primers. 2023; 9(1): 8. https://dx.doi.org/10.1038/s41572-023-00417-6
- Saleem S., McClure E.M., Goudar S.S., Patel A., Esamai F., Garces A. et al. A prospective study of maternal, fetal and neonatal deaths in low- and middle-income countries. Bull. World Health Organ. 2014; 92(8): 605-12. https://dx.doi.org/10.2471/BLT.13.127464
- Министерство здравоохранения Российской Федерации. Клинические рекомендации. Преэклампсия. Эклампсия. Отеки, протеинурия и гипертензивные расстройства во время беременности, в родах и послеродовом периоде. 2024. Доступно по: https://cr.minzdrav.gov.ru/view-cr/637_2 [Ministry of Health of the Russian Federation. Clinical guidelines. Preeclampsia. Eclampsia. Edema, proteinuria, and hypertensive disorders during pregnancy, childbirth, and the postpartum period. 2024. Available at: https://cr.minzdrav.gov.ru/view-cr/637_2 (in Russian)].
- Chappell L.C., Cluver C.A., Kingdom J., Tong S. Pre-eclampsia. Lancet. 2021; 398(10297): 341-54. https://dx.doi.org/10.1016/S0140-6736(20)32335-7
- Brosens I., Pijnenborg R., Vercruysse L., Romero R. The "Great Obstetrical Syndromes" are associated with disorders of deep placentation. Am. J. Obstet. Gynecol. 2011; 204(3): 193-201. https://dx.doi.org/10.1016/j.ajog.2010.08.009
- Zeisler H., Llurba E., Chantraine F., Vatish M., Cathrine A., Sennström S.M. et al. Predictive value of the sFlt-1:PlGF ratio in women with suspected preeclampsia. N. Engl. J. Med. 2016; 374(1): 13-22. https://dx.doi.org/10.1056/NEJMoa1414838
- Ходжаева З.С., Холин А.М., Вихляева Е.М. Ранняя и поздняя преэклампсия: парадигмы патобиологии и клиническая практика. Акушерство и гинекология. 2013; 10: 4-11 [Khodzhaeva Z.S., Kholin A.M., Vikhlyaeva E.M. Early and late preeclampsia: pathobiology paradigms and clinical practice. Obstetrics and Gynecology. 2013; 10: 4-11 (in Russian)].
- Воднева Д.Н., Романова В.В., Дубова Е.А., Павлов К.А., Шмаков Р.Г., Щеголев А.И. Клинико-морфологические особенности ранней и поздней преэклампсии. Акушерство и гинекология. 2014; 2: 35-40. [Vodneva D.N., Romanova V.V., Dubova E.A., Pavlov K.A., Shmakov R.G., Shchegolev A.I. Clinical and morphological features of early and late preeclampsia. Obstetrics and Gynecology. 2014; (2): 35-40 (in Russian)].
- Chaemsaithong P., Sahota D.S., Poon L.C. First trimester preeclampsia screening and prediction. Am. J. Obstet. Gynecol. 2022; 226(2S): S1071-S1097.e2. https://dx.doi.org/10.1016/j.ajog.2020.07.020
- Stepan H., Galindo A., Hund M., Schlembach D., Sillman J., Surbek D. et al. Clinical utility of sFlt-1 and PlGF in screening, prediction, diagnosis and monitoring of pre-eclampsia and fetal growth restriction. Ultrasound Obstet. Gynecol. 2023; 61(2): 168-80. https://dx.doi.org/10.1002/uog.26032
- Ходжаева З.С., Холин А.М., Шувалова М.П., Иванец Т.Ю., Демура С.А., Галичкина И.В. Российская модель оценки эффективности теста на преэклампсию sFlt-1/PlGF. Акушерство и гинекология. 2019; 2: 52-8. https://dx.doi.org/10.18565/aig.2019.2.52-58. [Khodzhaeva Z.S., Kholin A.M., Shuvalova M.P., Ivanets T.Yu., Demura S.A., Galichkina I.V. A Russian model for evaluating the efficiency of the sFlt-1/PlGF test for preeclampsia. Obstetrics and Gynecology. 2019; (2): 52-8 (in Russian). https://dx.doi.org/10.18565/aig.2019.2.52-58].
- Wright D., Syngelaki A., Akolekar R., Poon L.C., Nicolaides K.H. Competing risks model in screening for preeclampsia by maternal characteristics and medical history. Am. J. Obstet. Gynecol. 2015; 213(1): 62.e1-10. https://dx.doi.org/10.1016/j.ajog.2015.02.018
- Rolnik D.L., Wright D., Poon L.C., O’Gorman N., Syngelaki A., de Paco Matallana C. et al. Aspirin versus placebo in pregnancies at high risk for preterm preeclampsia. N. Engl. J. Med. 2017; 377(7): 613-22. https://dx.doi.org/10.1056/NEJMoa1704559
- Poon L.C., Shennan A., Hyett J.A., Kapur A., Hadar E., Divakar H. et al. The International federation of gynecology and obstetrics (FIGO) initiative on pre-eclampsia. Int. J. Gynaecol. Obstet. 2019; 145(Suppl 1): 1-33. https://dx.doi.org/10.1002/ijgo.12802
- O'Gorman N., Wright D., Poon L.C., Rolnik D.L., Syngelaki A., de Alvarado M. et al. Multicenter screening for pre-eclampsia by maternal factors and biomarkers at 11-13 weeks' gestation: comparison with NICE guidelines and ACOG recommendations. Ultrasound Obstet. Gynecol. 2017; 49: 756-60. https://dx.doi.org/10.1002/uog.17455
- Aljameel S.S., Alzahrani M., Almusharraf R., Altukhais M., Alshaia S., Sahlouli H. et al. Prediction of preeclampsia using machine learning and deep learning models: a review. Big Data Cogn. Comput. 2023; 7(1): 32. https://dx.doi.org/10.3390/bdcc7010032
- O’Gorman N., Wright D., Syngelaki A., Akolekar R., Wright A., Poon L.C. et al. Competing risks model in screening for preeclampsia by maternal factors and biomarkers at 11-13 weeks’ gestation. Am. J. Obstet. Gynecol. 2016; 214(1): 103.e1-12. https://dx.doi.org/10.1016/j.ajog.2015.08.034
- Roberge S., Bujold E., Nicolaides K.H. Aspirin for the prevention of preterm and term preeclampsia: systematic review and metaanalysis. Am. J. Obstet. Gynecol. 2018; 218(3): 287-93.e1. https://dx.doi.org/10.1016/j.ajog.2017.11.561
- Bujold E., Roberge S., Lacasse Y., Bureau M. Prevention of preeclampsia and intrauterine growth restriction with aspirin started in early pregnancy: a meta-analysis. Obstet. Gynecol. 2010; 116(2 Pt 1): 402-14. https://dx.doi.org/10.1097/AOG.0b013e3181e9322a
- Van Calster B., McLernon D.J., van Smeden M., Wynants L., Steyerberg E.W. Calibration: the Achilles heel of predictive analytics. BMC Med. 2019; 17(1): 230. https://dx.doi.org/10.1186/s12916-019-1466-7
- Zwertbroek E.F., Groen H., Fontanella F., Maggio L., Caterina L.M., Bilardo M. Performance of the FMF first-trimester preeclampsia-screening algorithm in a high-risk population in the Netherlands. Fetal. Diagn. Ther. 2021; 48(2):103-11. https://dx.doi.org/10.1159/000512335
- Collins G.S., Reitsma J.B., Altman D.G., Moons K.G.M. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD). BMJ. 2015; 350: g7594. https://dx.doi.org/10.1136/bmj.g7594
- Poon L.C., Zymeri N.A., Zamprakou A., Syngelaki A., Nicolaides K.H. Protocol for measurement of mean arterial pressure at 11-13 weeks' gestation. Fetal. Diagn. Ther. 2012; 31(1): 42-8. https://dx.doi.org/10.1159/000335366
- Kovacheva V.P., Eberhard B.W., Cohen R.Y., Maher M., Saxena R., Gray K.J. Preeclampsia prediction using machine learning and polygenic risk scores. Hypertension. 2024; 81(2): 264-72. https://dx.doi.org/10.1161/HYPERTENSIONAHA.123.21053
- Li T., Xu M., Wang Y., Wang Y., Tang H., Duan H. et al. Prediction model of preeclampsia using machine learning based methods: a population based cohort study in China. Front. Endocrinol. (Lausanne). 2024; 15: 1345573. https://dx.doi.org/10.3389/fendo.2024.1345573
- DeLong E.R., DeLong D.M., Clarke-Pearson D.L. Comparing the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach. Biometrics. 1988; 44(3): 837-45. https://dx.doi.org/10.2307/2531595
- Steyerberg E.W., Vickers A.J., Cook N.R., Gerds T., Gonen M., Obuchowski N. et al. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology. 2010; 21(1): 128-38. https://dx.doi.org/10.1097/EDE.0b013e3181c30fb2
- Vickers A.J., Elkin E.B. Decision curve analysis: a novel method for evaluating prediction models. Med. Decis. Making. 2006; 26(6): 565-74. https://dx.doi.org/10.1177/0272989X06295361
- Холин А.М., Муминова К.Т., Балашов И.С., Ходжаева З.С., Боровиков П.И., Иванец Т.Ю., Гус А.И. Прогнозирование преэклампсии в первом триместре беременности: валидация алгоритмов скрининга на российской популяции. Акушерство и гинекология. 2017; 8: 74-84. https://dx.doi.org/10.18565/aig.2017.8.74-84 [Kholin A.M., Muminova K.T., Balashov I.S., Khodzhaeva Z.S., Borovikov P.I., Ivanets T.Yu., Gus A.I. First-trimester prediction of preeclampsia: Validation of screening algorithms in a Russian population. Obstetrics and Gynecology. 2017; (8): 74-84 (in Russian). https://dx.doi.org/10.18565/aig.2017.8.74-84].
- Ившин А.А., Малышев Н.А. Ранняя стратификация риска преэклампсии на основе мультипараметрической модели машинного обучения и рутинно собираемых клинических данных. Акушерство, гинекология и репродукция. 2026; 20(1): 111-29. https://dx.doi.org/10.17749/2313-7347/ob.gyn.rep.2025.706 [Ivshin A.A., Malyshev N.A. Early stratification of the risk of preeclampsia based on a multiparametric machine learning model and routinely collected clinical data. Obstetrics, Gynecology and Reproduction. 2026; 20(1): 111-29 (in Russian). https://dx.doi.org/10.17749/2313-7347/ob.gyn.rep.2025.706].
- Tan M.Y., Syngelaki A., Poon L.C., Rolnik D.L., O'Gorman N., Delgado J.L. et al. Screening for pre-eclampsia by maternal factors and biomarkers at 11-13 weeks’ gestation. Ultrasound Obstet. Gynecol. 2018; 52(2): 186-95. https://dx.doi.org/10.1002/uog.19112
- Platt J.C. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. In: Smola A.J., Bartlett P.L., Schölkopf B., Schuurmans D., eds. Advances in large margin classifiers. Cambridge: MIT Press; 1999: 61-74.
- Gil M.M., Cuenca-Gómez D., Rolle V., Pertegal M., Díaz C., Revello R. et al. Validation a machine-learning model for first-trimester prediction of pre-eclampsia using the cohort from the PREVAL study. Ultrasound Obstet. Gynecol. 2024; 63(1): 68-74. https://dx.doi.org/10.1002/uog.27480
- Torres-Torres J., Villafan-Bernal J.R., Martinez-Portilla R.J., Hidalgo-Carrera J.A., Estrada-Gutierrez G., Adalid-Martinez-Cisneros R. et al. Performance of machine-learning approach for prediction of pre-eclampsia in a middle-income country. Ultrasound Obstet. Gynecol. 2024; 63(3): 350-7. https://dx.doi.org/10.1002/uog.27510
Received 19.01.2026
Accepted 11.02.2026
About the Authors
Aleksandr A. Ivshin, PhD, Associate Professor, Head of the Department of Obstetrics and Gynecology, Dermatovenerology of the Medical Institute, Petrozavodsk State University, 31, Krasnoarmeyskaya str., Petrozavodsk, Republic of Karelia, 185035, Russia, +7(909)567-12-51, scipeople@mail.ru, https://orcid.org/0000-0001-7834-096XYuliya S. Boldina, PhD Student, Senior Lecturer of Department of Obstetrics, Gynecology, Dermatovenereology of the Medical Institute, Petrozavodsk State University, 31, Krasnoarmeyskaya str., Petrozavodsk, Republic of Karelia, 185035, Russia; Obstetrician-Gynecologist, Republican Perinatal Center named after K.A. Gutkin, Petrozavodsk, +7(981)405-85-24, ulia.isakova94@gmail.com, https://orcid.org/0000-0002-1450-650X
Nikita A. Malyshev, PhD Student in the scientific specialty «Information-Measuring and Control Systems», Lecturer, Department of Family Medicine, Public Health, Healthcare Management, Life Safety, Disaster Medicine of the Medical Institute, Petrozavodsk State University, 31, Krasnoarmeyskaya str., Petrozavodsk, Republic of Karelia, 185035, Russia, +7(921)461-38-60, malyshev.nikita.2016@gmail.com, https://orcid.org/0009-0005-2722-5976
Corresponding author: Alexandr A. Ivshin, scipeople@mail.ru



