Please use this identifier to cite or link to this item:
http://dspace.aiub.edu:8080/jspui/handle/123456789/111
Title: | Comparative Analysis of Three Improved Deep Learning Architectures for Music Genre Classification |
Authors: | Rafi, Quazi Ghulam Noman, Mohammed Prodhan, Sadia Zahin Alam, Sabrina Nandi, Dip |
Keywords: | Music information retrieval, music genre classification, deep learning, Convolutional Neural Network, Recurrent Neural Network, Convolutional - Recurrent Neural Network |
Issue Date: | 13-Sep-2020 |
Publisher: | I.J. Information Technology and Computer Science |
Citation: | Quazi Ghulam Rafi, Mohammed Noman, Sadia Zahin Prodhan, Sabrina Alam, Dip Nandi, "Comparative Analysis of Three Improved Deep Learning Architectures for Music Genre Classification", International Journal of Information Technology and Computer Science(IJITCS), Vol.13, No.2, pp.1-14, 2021. DOI: 10.5815/ijitcs.2021.02.01 |
Abstract: | Among the many music information retrieval (MIR) tasks, music genre classification is noteworthy. The categorization of music into different groups that came to existence through a complex interplay of cultures, musicians, and various market forces to characterize similarities between compositions and organize collections is known as a music genre. The past researchers extracted various hand-crafted features and developed classifiers based on them. But the major drawback of this approach was the requirement of field expertise. However, in recent times researchers, because of the remarkable classification accuracy of deep learning models, have used similar models for MIR tasks. Convolutional Neural Net- work (CNN), Recurrent Neural Network (RNN), and the hybrid model, Convolutional - Recurrent Neural Network (CRNN), are such prominently used deep learning models for music genre classification along with other MIR tasks and various architectures of these models have achieved state-of-the-art results. In this study, we review and discuss three such architectures of deep learning models, already used for music genre classification of music tracks of length of 29-30 seconds. In particular, we analyze improved CNN, RNN, and CRNN architectures named Bottom-up Broadcast Neural Network (BBNN) [1], Independent Recurrent Neural Network (IndRNN) [2] and CRNN in Time and Frequency dimensions (CRNN- TF) [3] respectively, almost all of the architectures achieved the highest classification accuracy among the variants of their base deep learning model. Hence, this study holds a comparative analysis of the three most impressive architectural variants of the main deep learning models that are prominently used to classify music genre and presents the three architecture, hence the models (CNN, RNN, and CRNN) in one study. We also propose two ways that can improve the performances of the RNN (IndRNN) and CRNN (CRNN-TF) architectures. |
URI: | http://dspace.aiub.edu:8080/jspui/handle/123456789/111 |
ISSN: | 2074-9015 |
Appears in Collections: | Publications: Journals |
Files in This Item:
File | Description | Size | Format | |
---|---|---|---|---|
Draft_DSpace_Publication_Info_Dip_4.pdf | journal article | 165.63 kB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.