STFT and Gaussian Filter-Based Soundscape classification via 2D Convolutional Neural Networks
DOI:
https://doi.org/10.51485/ajss.v11i2.320Keywords:
STFT, CNN, Sound ClassificationAbstract
A novel approach is presented for soundscape classification, relying on STFT-based spectral characteristics combined with Gaussian filtering to generate enhanced time–frequency representations. Data augmentation is achieved through pitch shifting and time stretching to increase variability in the audio sequences. The signals are transformed using the Short-Time Fourier Transform and subsequently converted to the decibel scale. A normalized Gaussian filter is then applied to yield the Gaussian-STFT feature maps. For the classification stage, a two-dimensional convolutional neural network composed of three convolutional blocks followed by dense layers is employed. Experimental evaluation on the ESC-10 dataset achieves an accuracy of 97%, a precision of 97.34%, a recall of 97.08%, and an F1-score of 97.11%. Evaluation on the Urban Sound dataset yields an accuracy of 90%, a precision of 90.59%, a recall of 89.96%, and an F1-score of 89.97%. Evaluation on the ESC-50 dataset results in an accuracy of 90%, a precision of 90.71%, a recall of 90.00%, and an F1-score of 89.94%.
Downloads
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Razika SOUADEK

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.

