Ankan ghosh

Create your portfolio with ProoVCreate your portfolio with ProoV
ProoV
DE Verified work portfolio

ProoV Portfolio

Ankan ghosh

Computer science and engineering · Lovely professional university

View the ProoV leaderboard

Projects

Certified

Self-directed project

Business & Sponsorships — The FC Barcelona Case

Data Science · July 2026

92/ 100

Built an end-to-end sponsorship case for FC Barcelona: analyzed the revenue mix, segmented fans by region, selected an awareness-focused target audience for ElectraX, assembled a budget-fit asset package, and evaluated campaign performance against KPIs. The final recommendation was explicit and commercially grounded, with clear caveats about the synthetic dataset and a defined threshold for revisiting the deal.

Graded against

  • Business reasoning & decision quality35%
  • Correct use of data25%
  • Frameworks applied20%
  • Communication & structure20%

Passed · pass mark 55/100

What stood out6
  • Correctly identified commercial revenue as the controllable growth stream and computed the 12% uplift target as €392m.
  • Chose Asia-Pacific for ElectraX using a clear awareness logic: largest segment and highest engagement, even though its per-fan value was lowest.
  • Built a relevant awareness-led package and kept spend within budget at €2.9m of €3.2m.
  • Reported KPI performance with actual-vs-target percentages and clearly isolated content reach as the lagging metric.
  • Leakage-Free Commercial Segmentation
  • Evidence-Based Sponsorship Pricing
The work I submitted10 tasks

Revenue Mix

My solution
{"checkpointId":"revenue_mix","code":"import pandas as pd\nimport matplotlib.pyplot as plt\n\n# Synthetic club revenue (illustrative, in € millions)\ndf = pd.read_csv('fcb_commercial.csv')\n\n# Each stream as a share of total revenue\ntotal = df['amount_eur_m'].sum()\ndf['share_%'] = (df['amount_eur_m'] / total * 100).round(1)\n\n# Sort biggest -> smallest so the table reads top-down\ndf = df.sort_values('amount_eur_m', ascending=False)\nprint(df.to_string(index=False))\nprint(f\"\\nTotal revenue: €{total:.0f}m\")\n\n# Make a simple bar chart of each stream's share of revenue\nplt.figure(figsize=(6, 3.5))\nplt.bar(df['stream'], df['share_%'], color='#a50044')\nplt.title('Revenue mix, share of total (%)')\nplt.ylabel('Share of revenue (%)')\nfor i, pct in enumerate(df['share_%']):\n    plt.text(i, pct + 1, f\"{pct}%\", ha='center', fontsize=9)\nplt.tight_layout()\nplt.show()","stdout":"      stream  amount_eur_m  share_%\nBroadcasting           360     45.0\n  Commercial           300     37.5\n    Matchday           140     17.5\n\nTotal revenue: €800m\n","stderr":"<proov-cell>:24: UserWarning: FigureCanvasAgg is non-interactive, and thus cannot be shown\n","error":null,"imageCount":1,"durationMs":889,"ranAt":"2026-07-05T04:37:19.032Z"}

Revenue Mix Compute

My solution
{"prompt":"Your team's target is to grow commercial revenue by 12% next season. Work out the new commercial figure, and name the one assumption your number depends on.","value":"392","assumption":"We land at least one new global partner or renew existing sponsors at last year's rate, with no drop-off in the category."}

Fan Segmentation

My solution
{"checkpointId":"fan_segmentation","code":"import pandas as pd\n\n# Synthetic fan sample (illustrative)\nfans = pd.read_csv('fcb_fans.csv')\n\n# Group by region: how many fans, and their average value to a sponsor\nsummary = (fans.groupby('region')\n                .agg(fans=('fan_id', 'count'),\n                     avg_value_eur=('est_value_eur', 'mean'),\n                     avg_engagement=('engagement_score', 'mean'))\n                .round(1)\n                .sort_values('fans', ascending=False))\n\nprint(summary.to_string())","stdout":"              fans  avg_value_eur  avg_engagement\nregion                                           \nAsia-Pacific    10           10.3            80.2\nAmericas         7           21.6            60.1\nEurope           7           20.1            60.3\nAfrica           6            8.2            71.7\nMENA             6           13.5            63.2\n","stderr":"","error":null,"imageCount":0,"durationMs":33,"ranAt":"2026-07-05T04:42:16.542Z"}

Fan Segmentation Note

My solution
{"prompt":"Which segment would you target for the sponsor, and why? Reference the size, value, or engagement numbers you just saw.","answer":"Asia-Pacific — largest segment (10 fans), highest engagement (80.2), and this is exactly what an emerging EV brand chasing global awareness needs: scale and attention, not high per-fan value. Their average value (10.3) is actually the lowest of all regions, but that doesn't matter for an awareness play — you want reach and engaged eyeballs, not high-spending fans. Value-driven targeting (Americas/Europe, ~€20-21 avg value) would matter more for a commercial-delivery brand, not an awareness one."}

Package Builder

My solution
{"selectedItems":["training-kit","led-tier1","social-tier1","branded-content"],"selectedNames":"Training Kit, LED Boards - Tier 1, Social Media Pack A, Branded Content Series","total":2.9,"aiScore":100,"aiFeedback":[{"title":"Budget put to work","text":"You committed €2.9M of the €3.2M, idle budget is wasted reach.","type":"success"},{"title":"Reach-led mix","text":"You leaned on broad-reach assets, the right instinct for a brand chasing global awareness.","type":"success"}]}

Negotiation Simulator

My solution
{"finalSentiment":50,"dealSecured":false,"endedReason":"connection","transcript":"[Maria]: Thank you for coming. I'll be direct, we've spoken to Real Madrid and PSG as well. I need to understand why Barça is different. Not emotionally. Commercially.\n[Player]: I hear you on budget, we can drop the price and keep the full package the same if that gets us to yes today.\n[Maria]: Sorry, you cut out for a second there. Could you say that again?\nFinal Sentiment: 50 (ended: connection)"}

Inbox Responses

My solution
{"responses":{"0":"Hi Maria,\n\nGreat news, we can make this happen. Let's align on the match with your CEO's exact travel dates so we lock in front-of-camera LED rotation during peak broadcast windows, not just any slot. I'll coordinate directly with matchday operations to guarantee visibility and can also arrange a hospitality touchpoint for him on the day. Send me the dates and I'll confirm within 24 hours.","1":"Understood, but ElectraX's contract guarantees first-half visibility, so this creates a shortfall we need to fix before their CEO's visit next month. Can we offset with additional LED slots in match 4, or find inventory in another zone? I'll inform Maria proactively with a make-good so this doesn't derail the renewal conversation.","2":"Understood, but ElectraX's contract guarantees first-half visibility, so this creates a shortfall we need to fix before their CEO's visit next month. Can we offset with additional LED slots in match 4, or find inventory in another zone? I'll inform Maria proactively with a make-good so this doesn't derail the renewal conversation."},"aiScore":92,"aiFeedback":[{"title":"The CEO visit","text":"Good, you took ownership of the visit and committed to coordinating it rather than just confirming the basics.","type":"success"},{"title":"The LED cut","text":"Strong instinct, leading with transparency and a make-good is exactly how you keep trust when an asset slips.","type":"success"},{"title":"The renewal signal","text":"An unhappy partner eight weeks out needs a proactive face-to-face, not a wait-and-see reply.","type":"warning"}]}

KPI Scorecard

My solution
{"checkpointId":"kpi_scorecard","code":"import pandas as pd\n\n# Synthetic mid-season campaign results (illustrative)\nkpis = pd.DataFrame([\n    {\"kpi\": \"Impressions\",    \"actual\": 847, \"target\": 600},\n    {\"kpi\": \"Brand awareness\",\"actual\": 12,  \"target\": 8},\n    {\"kpi\": \"Web traffic\",    \"actual\": 340, \"target\": 200},\n    {\"kpi\": \"Newsletter signups\", \"actual\": 4800, \"target\": 5000},\n    {\"kpi\": \"Content reach\",  \"actual\": 1.2, \"target\": 2.0},\n])\n\n# % achievement = actual / target\nkpis['achievement_%'] = (kpis['actual'] / kpis['target'] * 100).round(0)\nprint(kpis.to_string(index=False))","stdout":"               kpi  actual  target  achievement_%\n       Impressions   847.0   600.0          141.0\n   Brand awareness    12.0     8.0          150.0\n       Web traffic   340.0   200.0          170.0\nNewsletter signups  4800.0  5000.0           96.0\n     Content reach     1.2     2.0           60.0\n","stderr":"","error":null,"imageCount":0,"durationMs":24,"ranAt":"2026-07-05T04:51:10.946Z"}

Mid Season Performance Summary

My solution
{"summary":"The partnership has delivered strong momentum for the board. Web traffic reached 340k clicks, beating our 200k target by 70%, while brand awareness rose 12%, well above the +8% goal, and visibility hit 847M impressions across broadcast, social, and digital. One area fell short: content reach landed at 1.2M against a 2.0M target. This gap stems from underdelivered branded content volume in Q2. To close it before season end, we're adding two additional content drops per month and reallocating budget from lower-performing channels toward high-engagement formats. Overall, the partnership is outperforming plan on the metrics that matter most for ElectraX's awareness goals, and we have a clear path to close the one lagging metric.","aiScore":95,"aiFeedback":[{"title":"Lead with the wins","text":"Open with the impressions, awareness lift and web traffic that beat target.","type":"success"},{"title":"You owned the miss","text":"Naming the content-reach shortfall is exactly what builds trust with a board.","type":"success"},{"title":"Concrete fix proposed","text":"Closing with a specific recovery action completes the sandwich.","type":"success"}]}

Final Recommendation

My solution
{"memo":"Recommendation: Sign ElectraX — 3-year deal, front-loaded on awareness assets, ~€2.9-3.2M/year within the negotiated ZOPA.\nEvidence:\nSegmentation data showed Asia-Pacific as the strongest-fit fan segment for ElectraX's mission: largest base (10 fans in-sample), highest engagement (80.2), even though lowest average value (€10.3). That's the correct trade-off for an awareness-stage brand chasing reach over revenue-per-fan.\nThe asset package built for ElectraX (Training Kit, LED Boards Tier 1, Social Media Pack A, Branded Content Series — €2.9M of €3.2M budget) matches this awareness mission directly. We deliberately excluded VIP hospitality, player appearances, and naming rights, since those serve repositioning or commercial-sales briefs, not ElectraX's.\nBudget fit sits inside our negotiated ZOPA (€2.6M–€3.2M), landed using BATNA logic — the club's alternative (selling the category to a rival bidder) kept our floor firm without needing to discount assets.\nMid-season KPIs support renewal: impressions, brand awareness, and web traffic all beat target (141%, 150%, 170% of plan respectively). Content reach underperformed (60% of target) but is a fixable execution gap, not a strategic misfit, and a remediation plan is already in motion.\nCaveats: All revenue, fan, and KPI figures in this case are synthetic/illustrative, not real Barça financials. Real due diligence would need audited fan data and a live competing-bid comparison, not an assumed BATNA.\nWhat would change my call: If content reach stayed under 70% of target after the fix window, or if a rival brand's real offer materially exceeded our ZOPA ceiling, I'd revisit the recommendation to sign at this price and structure.","wordCount":250,"selfCheck":{"rec":true,"evidence":true,"frameworks":true,"caveats":true,"change":true}}
Verified certificateTamper-proof · issued by ProoV
Certified

Self-directed project

Machine Learning for Automotive Safety

Perception Validation Engineer (ADAS Safety Sign-off) · AI / ML · June 2026

92/ 100

A verified, self-directed Perception Validation Engineer (ADAS Safety Sign-off) project — completed and passed a real industry rubric at 92/100 on ProoV.

Graded against

  • Detection metrics correctness (precision, recall, VRU false-negative rate)30%
  • Failure-mode analysis & per-condition rigour25%
  • SOTIF safety case quality (ISO 21448 reasoning)25%
  • Ship verdict — evidence-based & internally consistent20%

Passed · pass mark 60/100

What stood out6
  • Correctly established IoU ≥ 0.5 as the detection match threshold and computed the worked example as a miss when IoU = 0.111
  • Computed and interpreted aggregate precision/recall and the VRU false-negative rate correctly from the synthetic table
  • Identified the dangerous slices clearly: night, occluded, and night+occluded VRU conditions with the worst recall at 0.27
  • Wrote concrete SOTIF mitigations tied to specific failure modes, including daylight-only restriction and sensor-fusion fallback
  • Data-Driven Decision Making
  • Safety-Critical Reasoning
The work I submitted12 tasks

Sensor Match

My solution
{"prompt":"Match each driving condition to its primary-strength sensor and the weak-link sensor that fails there","map":[{"condition":"Pitch-dark rural road","primaryStrength":"Lidar","weakLink":"Camera","correct":true},{"condition":"Dense fog on the motorway","primaryStrength":"Radar","weakLink":"Camera","correct":true},{"condition":"Child behind a parked van","primaryStrength":"Lidar","weakLink":"Occlusion","correct":true},{"condition":"Clear sunny highway","primaryStrength":"Camera","weakLink":"Radar","correct":true}],"passed":true,"attempt":2}

Task Definition Code

My solution
{"checkpointId":"ttcs_iou_classify_run","code":"# Run IoU on this pair, read the number, then classify it (Section 1).\n# Boxes are [x1, y1, x2, y2] in pixels. The object is REAL, there is a ground-truth box.\n\ngt   = [60, 50, 200, 150]   # ground truth: a real object is here\npred = [150, 96, 280, 196]   # the detector's box for this frame\n\ndef iou(a, b):\n    ix1, iy1 = max(a[0], b[0]), max(a[1], b[1])\n    ix2, iy2 = min(a[2], b[2]), min(a[3], b[3])\n    inter = max(0, ix2 - ix1) * max(0, iy2 - iy1)\n    union = (a[2]-a[0])*(a[3]-a[1]) + (b[2]-b[0])*(b[3]-b[1]) - inter\n    return inter / union\n\nscore = iou(gt, pred)\nprint(f\"IoU = {score:.3f}\")\nprint(f\"threshold = 0.5  ->  {'matches' if score >= 0.5 else 'does NOT match'} a ground-truth box\")\nprint(\"So: a real object, no matching box. Your call on the next panel.\")","stdout":"IoU = 0.111\nthreshold = 0.5  ->  does NOT match a ground-truth box\nSo: a real object, no matching box. Your call on the next panel.\n","stderr":"","error":null,"imageCount":0,"durationMs":2,"ranAt":"2026-06-26T03:46:55.158Z"}

Task Definition

My solution
{"prompt":"Confirm the IoU threshold at 0.5, compute IoU for the borderline pair, and classify it as hit / miss / phantom; define a false negative as an unmatched real object","iouThreshold":0.5,"pairIoU":0.111,"classification":"miss","falseNegativeDefinition":"a real (ground-truth) object with no predicted box matched at IoU >= 0.5","cellRun":true,"passed":true,"attempt":1}

Scored Metrics Run

My solution
{"checkpointId":"ttcs_scored_metrics","code":"# Section 2: score the Aurora-7 detector on its detections table.\n# SYNTHETIC, illustrative detections-vs-ground-truth (no raw imagery, no model runs here;\n# proportions grounded in the published KITTI/nuScenes benchmarks).\n# Each row = one scored object, tagged with class + condition, with an outcome:\n#   \"TP\" caught,  \"FN\" a real object MISSED (false negative),  \"FP\" a phantom box.\nimport pandas as pd\n\n# Read the SAME scored detections table Act 3 slices by condition, so Section 2 and\n# Section 3 can never disagree. (Normalise two column names so the scoring below reads\n# cleanly: class -> obj_class, flag -> outcome.)\ndf = pd.read_csv(\"aurora7_detections_kitti.csv\").rename(\n    columns={\"class\": \"obj_class\", \"flag\": \"outcome\"})\nprint(f\"scored objects in table: {len(df)}\")\n\ndef score(d, label):\n    TP = (d.outcome == \"TP\").sum()\n    FP = (d.outcome == \"FP\").sum()\n    FN = (d.outcome == \"FN\").sum()\n    precision = TP / (TP + FP)\n    recall    = TP / (TP + FN)\n    print(f\"{label:18s} TP={TP:4d} FP={FP:3d} FN={FN:3d}  \"\n          f\"precision={precision:.2f}  recall={recall:.2f}\")\n    return recall\n\n# 1) The whole table, the headline (aggregate) numbers.\nagg_recall = score(df, \"ALL (aggregate)\")\n\n# 2) Vulnerable road users only = pedestrians + cyclists. THIS is the safety subset.\nvru = df[df.obj_class.isin([\"Pedestrian\", \"Cyclist\"])]\nvru_recall = score(vru, \"VRU (Ped+Cyc)\")\n\n# 3) The safety number: VRU false-negative RATE = share of real VRUs MISSED.\nvru_fn_rate = 1 - vru_recall\nprint(\"-\" * 60)\nprint(f\"AGGREGATE recall (all classes) : {agg_recall:.2f}   <- looks reassuring\")\nprint(f\"VRU false-negative rate        : {vru_fn_rate:.2f}   <- the safety number\")","stdout":"scored objects in table: 991\nALL (aggregate)    TP= 822 FP= 41 FN=128  precision=0.95  recall=0.87\nVRU (Ped+Cyc)      TP= 392 FP= 35 FN=122  precision=0.92  recall=0.76\n------------------------------------------------------------\nAGGREGATE recall (all classes) : 0.87   <- looks reassuring\nVRU false-negative rate        : 0.24   <- the safety number\n","stderr":"","error":null,"imageCount":0,"durationMs":23,"ranAt":"2026-06-26T03:49:41.322Z"}

Failure Triage

My solution
{"prompt":"Slice VRU recall per condition, then triage the failure modes by (condition likelihood x miss impact); do not bury a rare-but-catastrophic VRU miss in green","aggregateVruRecall":0.76,"perConditionRecall":[{"condition":"Daylight","recall":0.85},{"condition":"Clear","recall":0.83},{"condition":"Common class","recall":0.79},{"condition":"Rare class (cyclist)","recall":0.66},{"condition":"Occluded","recall":0.45},{"condition":"Night","recall":0.39},{"condition":"Night + occluded","recall":0.27}],"worstCondition":"Night + occluded","worstConditionRecall":0.27,"triage":[{"failureMode":"Night-time pedestrian miss","recallSlice":"night 0.39","likelihood":"Common","impact":"Catastrophic","severityBand":"red"},{"failureMode":"Occluded child at night","recallSlice":"night+occluded 0.27","likelihood":"Occasional","impact":"Catastrophic","severityBand":"red"},{"failureMode":"Rare-class cyclist miss","recallSlice":"rare 0.66","likelihood":"Common","impact":"Catastrophic","severityBand":"red"},{"failureMode":"Occluded pedestrian (daytime)","recallSlice":"occluded 0.45","likelihood":"Common","impact":"Catastrophic","severityBand":"red"},{"failureMode":"Night glare phantom brake","recallSlice":"precision dip","likelihood":"Common","impact":"Serious","severityBand":"red"},{"failureMode":"Clear-daylight car miss","recallSlice":"day 0.99","likelihood":"Common","impact":"Serious","severityBand":"red"},{"failureMode":"Dusk / low-sun pedestrian","recallSlice":"transition","likelihood":"Common","impact":"Catastrophic","severityBand":"red"}],"topBandFailureModes":["Night-time pedestrian miss","Occluded child at night","Rare-class cyclist miss","Occluded pedestrian (daytime)","Night glare phantom brake","Clear-daylight car miss","Dusk / low-sun pedestrian"],"passed":false,"attempt":2}

Sotif Case

My solution
{"prompt":"For each top failure mode write one ISO 21448 SOTIF chain (triggering condition -> functional insufficiency -> hazardous behaviour -> concrete mitigation); a vague mitigation or a SOTIF gap framed as a fault fails","rows":[{"failureMode":"Occluded child at night","triggeringCondition":"A small child steps out from behind a parked vehicle after dark. The occlusion removes any pre-emergence warning, and night conditions eliminate the texture and contrast the camera needs to detect a low-height figure.","functionalInsufficiency":"The camera cannot resolve a small, partially occluded figure in low light. The compound condition — night plus occlusion — drops VRU recall to ~0.27, the worst slice in the dataset. The sensor is working as designed; the design is simply not enough here.","hazardousBehaviour":"No braking response, or a critically late one, on a child already in the vehicle's path at close range with near-zero stopping distance remaining.","mitigation":"Mandate lidar-camera fusion confirmation before suppressing a pedestrian brake signal in any scenario where a parked vehicle occlusion is detected. If lidar is unavailable or degraded, fall back to driver-attention monitoring with mandatory takeover request. Do not permit autonomous low-speed urban operation after dark until the night+occluded recall clears an agreed threshold (suggest ≥0.75).","mitigationIsConcrete":true,"framedAsFault":false},{"failureMode":"Rare-class cyclist miss","triggeringCondition":"A cyclist appears in the vehicle's path — any lighting condition, any road type. The detector was trained on data where cyclists were severely underrepresented, creating a systematic class-level blind spot that persists even in clear daylight.","functionalInsufficiency":"The model lacks sufficient training exposure to cyclist morphology, movement patterns, and edge appearances. Daytime recall is ~0.66 — the detector misses one in three cyclists even under ideal conditions. This is not a sensor failure; it is a training distribution gap.","hazardousBehaviour":"Absent or late collision avoidance response when a cyclist is directly in the vehicle's path — the car behaves as though the lane is clear.","mitigation":"Immediately augment the training set with targeted cyclist data (minimum class-balanced representation across lighting, angle, and clothing variation) and retrain until cyclist recall clears ≥0.85 in daylight and ≥0.70 at night. Until that bar is cleared, restrict autonomous operation on roads with designated cycle lanes or high cyclist exposure, and enforce driver-attention monitoring as a mandatory fallback whenever a cyclist is detected by any secondary sensor the camera has not yet confirmed.","mitigationIsConcrete":true,"framedAsFault":false}],"mitigations":["Mandate lidar-camera fusion confirmation before suppressing a pedestrian brake signal in any scenario where a parked vehicle occlusion is detected. If lidar is unavailable or degraded, fall back to driver-attention monitoring with mandatory takeover request. Do not permit autonomous low-speed urban operation after dark until the night+occluded recall clears an agreed threshold (suggest ≥0.75).","Immediately augment the training set with targeted cyclist data (minimum class-balanced representation across lighting, angle, and clothing variation) and retrain until cyclist recall clears ≥0.85 in daylight and ≥0.70 at night. Until that bar is cleared, restrict autonomous operation on roads with designated cycle lanes or high cyclist exposure, and enforce driver-attention monitoring as a mandatory fallback whenever a cyclist is detected by any secondary sensor the camera has not yet confirmed."],"passed":true,"attempt":1}

Ship Verdict

My solution
{"prompt":"Write the GO / GO-WITH-CONDITIONS / NO-GO ship verdict, each of five fields citing its own dossier section number, internally consistent (no unconditional GO over an unmitigated night miss; no NO-GO if a daylight-restricted release is viable)","verdict":"CONDITIONS","fields":{"definition":"A valid detection requires IoU ≥ 0.5 between the predicted and ground-truth bounding box. Any prediction below this threshold is counted as a false negative. This threshold was set in Section 1 as the minimum overlap needed to confirm a VRU is meaningfully localised, not just approximately present.","metrics":"Overall VRU recall is 0.76, meaning the detector misses 1 in 4 vulnerable road users across all conditions. The false-negative rate is 24%. This is not acceptable for unrestricted deployment but is acceptable within a bounded ODD where the worst conditions are excluded.","worst":"The detector breaks hardest at night with occlusion. The night+occluded recall is 0.27 — the system misses nearly 3 in 4 occluded pedestrians after dark. The next worst slice is night-only at 0.39. These two slices alone disqualify an unconditional GO. A 73% miss rate on a child stepping into the road is not a residual risk — it is a primary hazard.","mitigation":"Lidar-camera fusion confirmation is mandated before any pedestrian brake signal is suppressed in occluded scenarios. After dark, if lidar is unavailable, the system falls back to driver-attention monitoring with mandatory takeover request. Autonomous pedestrian braking is capped to daylight operation until night+occluded recall clears ≥ 0.75. Cyclist training data augmentation is required before ODD expansion to roads with cycle infrastructure, with a minimum exit threshold of 0.85 daytime recall.","ask":"I am asking the board to approve Aurora-7 for daylight-only autonomous pedestrian braking within a defined ODD: good weather, dry roads, posted speed ≤ 60 km/h, no mandatory cycle lane exposure.\nI am asking the board to reject unrestricted night deployment in any form until the night+occluded recall clears 0.75 and the rare-class cyclist recall clears 0.85. These are not stretch targets — they are the minimum thresholds below which a miss is more likely than a detection in the worst slice.\nThe 24% overall FN rate is tolerable inside the daylight ODD where recall is approximately 0.99 on the dominant class. It is not tolerable after dark where the FN rate on VRUs exceeds 60%.\nThis is a conditional go. It is not a green light for night operation. Any expansion of the ODD requires the night recall bar to be cleared and re-signed."},"citedSections":["Detection Definition","Metrics","Worst Failure Mode","Mitigation","The Ask"],"internallyConsistent":true,"passed":true,"attempt":1}

Realcase Analysis

My solution
{"prompt":"In your own words, write what YOUR dossier would have caught about the real Uber ATG Tempe 2018 crash: which of your own rows, which real NTSB figure, and the mechanism that links them.","analysis":"My Tab 3 night-time pedestrian miss row — rated Common likelihood, Catastrophic impact, recall 0.39 — maps directly onto what happened in Tempe. The NTSB record shows the system detected an object ~5.6 seconds out but never resolved a stable class, cycling between vehicle, other, and bicycle. That is exactly the mechanism my triage named: in darkness, the camera loses the texture it needs to confirm a person, recall collapses, and a stable hit never forms. The 0.39 night recall figure in my dossier is the number that was missing in the real system — not a random fault, a predictable insufficiency in the dark. My Tab 4 SOTIF chain then named the hazardous behaviour directly: a late or absent brake on a real person already in the crossing. The Tempe system suppressed emergency braking by design approximately 1.3 seconds before impact. My mitigation — cap autonomous pedestrian braking to daylight operation until night recall clears an agreed bar — would have flagged this exact deployment decision as outside the safe ODD before the vehicle was on the road. The Tab 5 verdict makes it explicit: no unconditional GO over a night collapse you haven't closed. Tempe shipped that night collapse unclosed.","signalsDetected":["row","figure","mechanism"],"passed":true,"attempt":1}

Teachback

My solution
{"prompt":"Teach it back in plain words: why is recall on vulnerable road users, not accuracy, the metric that governs safety for a pedestrian detector? Anchor it to a real number you have seen.","teachback":"Accuracy counts every prediction the model gets right, divided by everything it touched. That sounds reasonable until you think about what the two types of mistakes actually cost.\nA false positive means the car braked for something that wasn't there — a phantom, a reflection, a shadow. Annoying, maybe a rear-end risk in traffic, but the car stopped. Nobody got hit.\nA false negative means the car saw nothing when there was a person. The car didn't brake. That person gets hit.\nThose two errors are not symmetric. One is an inconvenience. The other is a fatality. Accuracy treats them as equal — a correct phantom brake and a correct real detection cancel out the same way in the numerator. That is exactly wrong for safety.\nRecall measures only the false negative side: of every real person who was actually there, how many did the detector find? My Aurora-7 dossier showed overall VRU recall of 0.76 — meaning 1 in 4 vulnerable road users was missed across all conditions. At night that collapsed to 0.39, meaning the detector missed more pedestrians than it found in the dark. Accuracy would have looked fine because cars in daylight, the dominant class, were detected at 0.99. The pedestrian misses drowned in a sea of correct car detections.\nThat is why recall governs the safety case and accuracy doesn't. Accuracy can hide a catastrophic miss rate behind a comfortable headline number. Recall forces you to look at exactly the failure that kills people — the person the car never saw.","signalsDetected":["mechanism","number"],"passed":true,"attempt":1}

Example Iou Demo

My solution
{"checkpointId":"example_iou_demo","code":"# WORKED EXAMPLE: Intersection-over-Union for one clean box pair.\n# Boxes are [x1, y1, x2, y2] in pixels. (Next slide: you do this on a tougher pair.)\n\ngt   = [60, 50, 200, 150]   # ground-truth box: where the object really is\npred = [68, 56, 206, 156]   # the detector's predicted box (a near-perfect fit)\n\ndef iou(a, b):\n    # 1) the SHARED rectangle (intersection): the overlap of the two boxes\n    ix1, iy1 = max(a[0], b[0]), max(a[1], b[1])\n    ix2, iy2 = min(a[2], b[2]), min(a[3], b[3])\n    inter = max(0, ix2 - ix1) * max(0, iy2 - iy1)\n\n    # 2) the COMBINED area (union): both boxes together, counted once\n    area_a = (a[2] - a[0]) * (a[3] - a[1])\n    area_b = (b[2] - b[0]) * (b[3] - b[1])\n    union = area_a + area_b - inter\n\n    return inter / union, inter, union\n\nscore, inter, union = iou(gt, pred)\nTHRESHOLD = 0.5\nverdict = \"HIT  (counts as a detection)\" if score >= THRESHOLD else \"MISS (no match)\"\n\nprint(f\"shared area    = {inter:6.0f} px\")\nprint(f\"combined area  = {union:6.0f} px\")\nprint(f\"IoU            = {score:6.3f}   (shared / combined)\")\nprint(f\"threshold      = {THRESHOLD}\")\nprint(f\"VERDICT        = {verdict}\")","stdout":"shared area    =  12408 px\ncombined area  =  15392 px\nIoU            =  0.806   (shared / combined)\nthreshold      = 0.5\nVERDICT        = HIT  (counts as a detection)\n","stderr":"","error":null,"imageCount":0,"durationMs":1,"ranAt":"2026-06-26T04:11:45.426Z"}

Example Metrics Demo

My solution
{"checkpointId":"example_metrics_demo","code":"# WORKED EXAMPLE: score a tiny 8-row detections-vs-ground-truth split.\n# (Next slide: you do this on the detector's full table.)\nimport pandas as pd\n\n# Each row = one scored object. outcome is:\n#   \"TP\" a real object the detector caught,\n#   \"FN\" a real object it MISSED (a false negative),\n#   \"FP\" a phantom box over empty road (a false positive).\ndemo = pd.DataFrame([\n    (\"Pedestrian\", \"TP\"),   # 3 real pedestrians, 2 caught, 1 missed\n    (\"Pedestrian\", \"TP\"),\n    (\"Pedestrian\", \"FN\"),\n    (\"Cyclist\",    \"TP\"),   # 2 real cyclists, 1 caught, 1 missed\n    (\"Cyclist\",    \"FN\"),\n    (\"Car\",        \"TP\"),   # 2 real cars, both caught\n    (\"Car\",        \"TP\"),\n    (\"Car\",        \"FP\"),   # 1 phantom car (braked for nothing)\n], columns=[\"obj_class\", \"outcome\"])\n\n# 1) Confusion counts over the whole split.\nTP = (demo.outcome == \"TP\").sum()\nFP = (demo.outcome == \"FP\").sum()\nFN = (demo.outcome == \"FN\").sum()\nprint(f\"TP = {TP}   FP = {FP}   FN = {FN}\")\n\n# 2) Precision = of everything flagged, how much was real.\nprecision = TP / (TP + FP)\n# 3) Recall = of everything really there, how much we caught.\nrecall = TP / (TP + FN)\nprint(f\"precision = TP/(TP+FP) = {TP}/{TP+FP} = {precision:.2f}\")\nprint(f\"recall    = TP/(TP+FN) = {TP}/{TP+FN} = {recall:.2f}\")\n\n# 4) The safety number: recall on vulnerable road users (pedestrians + cyclists),\n#    then the false-negative RATE = the share of real VRUs we MISSED.\nvru = demo[demo.obj_class.isin([\"Pedestrian\", \"Cyclist\"])]\nvru_TP = (vru.outcome == \"TP\").sum()\nvru_FN = (vru.outcome == \"FN\").sum()\nvru_recall  = vru_TP / (vru_TP + vru_FN)\nvru_fn_rate = 1 - vru_recall\nprint(f\"VRU recall  = {vru_TP}/{vru_TP+vru_FN} = {vru_recall:.2f}\")\nprint(f\"VRU FN-rate = 1 - VRU recall = {vru_fn_rate:.2f}   <- the safety number\")","stdout":"TP = 5   FP = 1   FN = 2\nprecision = TP/(TP+FP) = 5/6 = 0.83\nrecall    = TP/(TP+FN) = 5/7 = 0.71\nVRU recall  = 3/5 = 0.60\nVRU FN-rate = 1 - VRU recall = 0.40   <- the safety number\n","stderr":"","error":null,"imageCount":0,"durationMs":5,"ranAt":"2026-06-26T04:12:27.450Z"}

Scored Metrics

My solution
{"prompt":"Score the detector: read off precision, aggregate recall, and the VRU false-negative rate (the safety number)","aggregatePrecision":0.95,"aggregateRecall":0.87,"vruRecall":0.76,"vruFalseNegativeRate":0.236,"safetyNumber":"VRU false-negative rate","readOffBand":"20–30%","readOffCorrect":true,"scoringCellRan":true,"passed":true,"attempt":1}
Verified certificateTamper-proof · issued by ProoV
2Projects completed
2Verified certificates
92Average score
Create your portfolio with ProoV