From traffic video to safety evidence

My published vision-language and generative-model review organizes research on interpreting traffic videos and examines grounding, temporal consistency, and verification.

The aim is to connect semantic descriptions with observable traffic evidence. Model-generated explanations and scenarios require checks against the underlying events and physical constraints.

Datasets and timing-sensitive evaluation

Two separate SAVeD research outputs address complementary questions:

They are distinct works, rather than a replacement title for the same record.

Occluded pedestrian crossing

The submitted multimodal risk-recognition study addresses unsignalized, occluded pedestrian crossing scenarios and graded prevention. It connects this research area with vulnerable-road-user safety and proactive risk assessment.

Verification and research resources

Benchmark and protocol design must distinguish observed events from generated scenarios, and warning evaluation must consider event timing. Related work includes the SAVeD dataset, warning protocol, ALARM benchmark, and MSC-MTSim simulation platform.

Methods: multimodal learning, vision-language models, structured scene representations, temporal event analysis, and evidence-based evaluation.