AutoKaggle ML Competition Pipeline

Iterative feature engineering, hypothesis validation, and LightGBM model tuning on the Titanic dataset.

0.865 Best 5-Fold CV Score
5 Rounds Autonomous Loop Iterations
LightGBM Target Model Architecture
0.789 Public Leaderboard Submission
5-Fold Cross-Validation Score Iteration Curve
0.78 0.81 0.84 0.87 R1: Baseline R2: Title Feature R3: Family Size R4: Overfitting Noise R5: Deck & Fare Bin 0.814 0.838 0.849 0.799 0.865
5-Fold CV Score Regression Rejected Kaggle Submission (0.789)
Autonomous Iteration Logs
Round 1 — Baseline Model
CV: 0.814
Trained baseline LightGBM model on raw numeric features. Established 5-fold cross-validation pipeline.
Round 2 — Feature Engineering (Title Extraction)
CV: 0.838 (+0.024)
Extracted honorific titles (Mr, Mrs, Miss, Master) from passenger names. CV score improved by +0.024.
Round 3 — Feature Engineering (Family Grouping)
CV: 0.849 (+0.011)
Combined SibSp and Parch into FamilySize and IsAlone flags. Hypothesis validated and committed.
Round 4 — High-Cardinality Ticket Frequency (Regressed)
CV: 0.799 (-0.050)
Attempted ticket prefix frequency encoding. CV score dropped by -0.050 due to noise overfitting. Automatically rolled back.
Round 5 — Hyperparameter Tuning & Cabin Deck Encoding
CV: 0.865 (+0.016)
Parsed Cabin deck letters and tuned num_leaves to 15. Achieved highest 5-fold CV score of 0.865. Submission file generated.