M.TECH · DISSERTATION · FINAL VIVA BITS PILANI WILP · 2025–2026
Final Presentation ● May 2026

Road Quality Assessment
Using Computer Vision
and GPS-Based Mapping

Crowdsourced, dashcam-based road quality assessment with quality-aware route planning across 808 OSM segments in Kerala.

Student
Anish A.2024TM93051 · M.Tech Software Engg.
Supervisor
Dr. Kavya ManoharAdalat AI
Additional Examiner
Mr. Ashik SalahudeenAuxmoney GmbH
01 / 10
02 · The Problem Road Quality Assessment
02 ● The Problem

Navigation today optimises distance and time —
not the quality of the road you actually drive on.

01 · The Gap
Maps ignore surface quality.
Every major navigation system today optimises distance and ETA. Road comfort, vehicle wear, and surface safety are invisible.
02 · The Indian Context
Western models do not transfer.
High visual clutter, monsoon damage, varied surfaces (asphalt, concrete, gravel, dirt) — pothole detectors trained on Japan / US / EU datasets fail on Indian roads.
03 · The Data Gap
Municipal surveys are manual & infrequent.
Surveys happen once or twice a year, are not published openly, and cannot inform driver-facing routing decisions in real time.
Research
Question
Can dashcam footage and computer vision automatically classify road quality at segment level, map it to a navigable graph, and route drivers around the worst of it?
02 / 10 Anish A. · 2024TM93051
03 · System Pipeline Six stages · dashcam → routing
03 ● End-to-End Pipeline

From dashcam video to quality-aware routes.

01
Dashcam
Ingestion
1080p footage from Kerala roads; FFmpeg frame extraction @ 1 fps.
~34k frames
per camera
›
02
OCR + GPS
Filtering
Tesseract OCR on GPS overlays; static / night / out-of-bounds frames dropped.
~26k frames
annotatable
›
03
Consensus
Annotation
Gamified platform · 2-annotator agreement with 3rd-annotator tiebreaker.
3,216 imgs
53 annotators
›
04
Model
Training
Architecture sweep + 72-trial HPO. Swin-Small fine-tuned at lr=3e-5.
15 archs
72 HPO trials
›
05
GPS
Snapping
osmnx + UTM projection; nearest-edge within 35 m; per-edge aggregation.
808 OSM
segments
›
06
Routing &
Map Server
Flask + Leaflet UI; shortest vs. smoothest path via quality-penalised Dijkstra.
REST API
+ live map
3,216
Consensus images
808
OSM segments mapped
79.6%
Validation acc · 3-class
75%
Macro F1 · 3-class
03 / 10 Anish A. · 2024TM93051
04 · Dataset & Annotation 3,216 consensus-validated images
04 ● Dataset & Annotation

A crowdsourced dataset built with consensus and active learning.

Class Distribution  ·  3-class (production)
3,216 images
Good73.9%2,376
Bad15.9%511
Invalid10.2%329
5-Class Breakdown (pre-merge)
33.2%
40.6%
12.4%
Excellent 1,069 Good 1,307 Fair 399 Poor 112 Invalid 329

Active learning moved class ratio from 9.2:1 → 4.6:1 (Good:Bad).

Annotation Protocol
Two annotators  →  agreement, or third tiebreaker.
Gamified platform with leaderboard and consensus-accuracy feedback kept contributors engaged across multiple weeks.
53
Contributors
3,216
Confirmed labels
4
Dashcam sources
Active Learning Loop
Frames where max(softmax) < 0.7 are pushed up the queue.
Targets minority & ambiguous samples. The Bad class grew from 4.1% → 15.9% across iterations — directly addressing imbalance. Group-aware train/val split on video identity prevents near-duplicate leakage.
04 / 10 Anish A. · 2024TM93051
05 · Architecture Sweep 13 successful · Swin-Small wins
05 ● Model Selection

A 13-architecture sweep, ranked by macro F1.

5-class · fixed lr=1e-4 · weighted loss · ~1.6 h
#ArchitectureParamsAccMacro F1
01swin_small ✓50M53.8%47.0%
02swin_tiny28M51.5%46.1%
03resnet3421M51.1%45.4%
04convnext_tiny28M50.4%44.8%
05resnet1811M49.7%44.4%
06vit_small22M49.8%44.0%
07efficientnet_b18M49.4%41.4%
8–13efficientnet_b0/b2 · resnet50 · convnext_s/base · vit_base (efficientnetv2_s/m failed)5–89M45–50%39–42%
Why Swin Wins
Hierarchical, shifted-window attention fits road texture.
Local 7×7 windows match pothole-scale features; shifted windows preserve cross-boundary context; multi-scale stages capture both crack-pattern texture and segment-level structure.
Why ViT Loses  ·  Inverse Scaling
Capacity-data mismatch  97% train · 59% val
Isotropic transformers memorise rather than learn on a 3k-image domain dataset. The same pattern repeats across families: ConvNeXt-Base (89M) ranked 13th vs. ConvNeXt-Tiny (28M) at 4th; ViT-Base (86M) at 11th vs. ViT-Small at 6th. Data-limited, not capacity-limited.
05 / 10 Anish A. · 2024TM93051
06 · Hyperparameter Tuning 72 trials · ~14 hours
06 ● Hyperparameter Tuning

72 trials surfaced a counter-intuitive recipe.

Learning Rate
3e-551.9%
1e-451.7%
3e-451.7%
Backbone Freeze
0 epochs52.1%
5 epochs51.7%
10 epochs51.4%
Label Smoothing
none51.9%
0.151.6%
Class Weighting
unweighted52.0%
weighted51.5%
Final Production Configuration
modelswin_small
learning rate · scheduler3e-5 · CosineAnnealingLR
backbone freezenone — full fine-tune
loss · smoothingunweighted CE · 0.0
optimiser · batch · stopAdam · 32 · val-loss patience 30
Surprise · 01 · Backbone freezing hurts
ImageNet → Indian dashcam is a wide domain gap. Early layers need to adapt. Two-phase training underperformed end-to-end fine-tuning at every duration.
Surprise · 02 · Smoothing & weights add nothing
Swin already learns reasonably balanced representations at this scale. The bottleneck is the class boundaries themselves, not optimisation.
06 / 10 Anish A. · 2024TM93051
07 · Class Consolidation Data-driven 5 → 3
07 ● The Class Consolidation Decision

Five classes were the bottleneck — not the model.

Confusion matrices across all 13 successful architectures showed persistent confusion at the Excellent ↔ Good and Fair ↔ Poor boundaries. Annotator-disagreement data corroborated this — humans split on those exact pairs. The ambiguity was real, not a modelling failure.

5-class · Excellent / Good / Fair / Poor / Invalid
53.0% acc
57.0% macro F1
Production swin_small on the 5-class consensus dataset. Macro F1 was respectable but minority-class recall remained poor.
›
3-class · Good / Bad / Invalid  ·  ADOPTED
79.6% acc
75.0% macro F1
Excellent + Good → Good; Fair + Poor → Bad. Directly answers the navigation question: is this segment safe to drive?
Insight The biggest single performance gain came from listening to the data — not from a deeper network or longer training. 53.0% → 79.6% acc  ·  57.0% → 75.0% macro F1.
07 / 10 Anish A. · 2024TM93051
08 · Final Model Performance swin_small · 3-class · 941 val samples
08 ● Final Model Performance

Strong on Good roads — conservative on Bad.

Confusion Matrix · 941 validation samples · Invalid excluded
Pred · Bad
Pred · Good
True
Bad
182
true negative · 53%
160
false positive · 47%
True
Good
38
false negative · 6%
561
true positive · 94%

Pessimistic edge aggregation at the segment level partially mitigates the 47% Bad miss-rate — a segment is only labelled Good if all observations agree.

Per-Class Classification Report
ClassPrecisionRecallF1Support
Bad83%53%65%342
Good78%94%85%599
Macro avg80%73%75%941
79.6%
Overall accuracy
75%
Macro F1
808
OSM segments mapped
Known Limitation
53% recall on Bad — the model is conservative.
Residual imbalance & visual ambiguity at driving speed. Mitigated at segment level via pessimistic edge aggregation in routing.
08 / 10 Anish A. · 2024TM93051
09 · Live Demonstration Four-part walkthrough
09 ● Live Demonstration

From annotation through to navigation — end to end.

01
Annotation Platform
flask · jinja · 53 contributors · 3,216 confirmed labels
02
Road Quality Maps
folium · osmnx · 808 segments · 35 m snapping
03
Quality-Aware Route Planning
dijkstra · pessimistic / majority modes · unrated = 2.5×1.2
04
Navigation Server
flask · leaflet.js · REST + GeoJSON · port 5050
05
Project Home roads.anishsheela.com  ↗
Reports · source code · live demos · maps · slides · dataset notes
09 / 10 Anish A. · 2024TM93051
10 · Conclusion & Future Work Thank you
10 ● Conclusion & Future Work

A complete, deployable pipeline —
dashcam to driver, in one system.

Project Home
roads.anishsheela.com
Reports · source · live demos · maps · slides  ↗
Contributions
Empirical case for classification over detection.

17 YOLO variants benchmarked; best mAP50 of 20.3%.

Consensus-annotation methodology.

3,216 labels across 53 contributors · gamified + active learning.

Data-driven class consolidation.

5 → 3 merge drove val accuracy 53% → 79.6%.

End-to-end geospatial integration.

UTM snapping, dual aggregation, Flask + Leaflet nav server.

Future Work
01 · Open Release
Anonymise & release dataset and models
PII-stripped frames + trained weights for the research community.
02 · Deployment
Mobile navigation application
Phone-first routing app for in-field driver use.
03 · Coverage
Expanded fleet & beyond Kerala
More cameras, post-monsoon seasons, other Indian states.
Student
Anish A.2024TM93051 · M.Tech Software Engg.
Supervisor
Dr. Kavya ManoharAdalat AI
Additional Examiner
Mr. Ashik SalahudeenAuxmoney GmbH
10 / 10 Thank you.