THESIS

Group Activity Recognition System Base on Spatio-Temporal Deep Learning in the Malioboro Pedestrian Area

 

ABSTRACT

Pedestrian zones such as Malioboro exhibit high human activity dynamics with unstructured interaction patterns, posing significant challenges for automated group activity monitoring using video-based systems. This study aims to develop a group activity recognition system based on spatio-temporal deep learning to analyze activity dynamics in crowded and dynamic pedestrian environments. The proposed system integrates YOLOv8 for individual detection and tracking, 3D ResNet-18 for human activity recognition, and a 1D Convolutional Neural Network (1D CNN) with an attention mechanism for group activity classification. Detected individual actions are converted into activity ratios and structured as time series to represent group dynamics without explicitly modeling inter-individual relations. Evaluation results show that the individual action recognition model achieves an accuracy of 96% with an average F1- score of 0.96. For group activity recognition, the proposed model achieves 98.17% accuracy with an average F1-score of 0.96. In a comparison using the same dataset, the proposed model significantly outperforms the Actor Relation Graph (ARG method, which demonstrates limited generalization across its two training stages. The proposed model yields group and individual accuracies of 98.17% and 96% respectively, which far exceeds the ARG method that only reaches 78.72% and 77.73% in Stage 1 and drops drastically to 37.72% and 29.66% in Stage 2. Evaluation on Malioboro pedestrian videos proves the system’s ability to generate stable predictions across various density levels. These findings indicate that individual activity ratio representation is more effective and adaptive for real-world environments compared to complex explicit relation-based approaches.

 

PUBLICATION
  • End-to-End Deep Learning System for Group Activity Recognition in Pedestrian Areas: A Case Study of Malioboro Street Yogyakarta (link)