Applying data augmentation technique on blast-induced overbreak prediction: Resolving the problem of data shortage and data imbalance

He, Biao and Armaghani, Danial Jahed and Lai, Sai Hin and Samui, Pijush and Mohamad, Edy Tonnizam (2024) Applying data augmentation technique on blast-induced overbreak prediction: Resolving the problem of data shortage and data imbalance. Expert Systems with Applications, 237 (C). ISSN 0957-4174, DOI https://doi.org/10.1016/j.eswa.2023.121616.

Full text not available from this repository.
Official URL: https://doi.org/10.1016/j.eswa.2023.121616

Abstract

Blast-induced overbreak in tunnels can cause severe damage and has therefore been a main concern in tunnel blasting. Researchers have developed many machine learning-based models to predict overbreak. Collecting overbreak data manually, however, can be challenging and might obtain insufficient or poorly structured data. Thus, this study aims to utilise a deep generative model, namely the Conditional Tabular Generative Adversarial Network (CTGAN), to establish an acceptable dataset for overbreak prediction. The CTGAN model was applied to overbreak data collected from paired tunnels: a left-line tunnel and a right-line tunnel. The overbreak dataset collected from the left-line tunnel-nominated as the true dataset-served to train the CTGAN model. Then the well-trained CTGAN model generated a synthetic overbreak dataset. Statistical-based approaches verified the similarity between the true and synthetic datasets; machine learning-based approaches verified the feasibility of using the synthetic dataset to train overbreak prediction model. Lastly, this study clarified how to resolve the problem of data shortage and data imbalance by leveraging the CTGAN model. The results evidence that the CTGAN model can effectively generate a high-quality synthetic overbreak dataset. The synthetic overbreak dataset not only greatly retains the properties of the true dataset but also effectively enhances its diversity. The way, integrating the true and synthetic overbreak datasets, can dramatically resolve the problem of data shortage and data imbalance in overbreak prediction. The findings in this study, therefore, highlight it as a promising perspective to resolve such a particular engineering problem.

Item Type: Article
Funders: UNSPECIFIED
Uncontrolled Keywords: Tunnel blasting; Overbreak; Environmental issue; Data augmentation; CTGAN; Machine learning
Subjects: Q Science > QA Mathematics > QA75 Electronic computers. Computer science
T Technology > TA Engineering (General). Civil engineering (General)
T Technology > TD Environmental technology. Sanitary engineering
Divisions: Faculty of Engineering > Department of Civil Engineering
Depositing User: Ms. Juhaida Abd Rahim
Date Deposited: 14 Jun 2024 02:49
Last Modified: 14 Jun 2024 02:49
URI: http://eprints.um.edu.my/id/eprint/44315

Actions (login required)

View Item View Item