Oversampling Method for Imbalanced Classification

Zhuoyuan Zheng; Yunpeng Cai; Ye Li

Authors

Zhuoyuan Zheng Guangxi Key Laboratory of Trusted Software, Guilin University of Electronic Technology, Guilin
Yunpeng Cai Shenzhen Institutes of Advanced Technology, Key Laboratory for Biomedical Informatics and Health Engineering, Chinese Academy of Sciences, Shenzhen
Ye Li Shenzhen Institutes of Advanced Technology, Key Laboratory for Biomedical Informatics and Health Engineering, Chinese Academy of Sciences, Shenzhen

Keywords:

Classification, imbalanced dataset, oversampling, SMOTE, SNOCC

Abstract

Classification problem for imbalanced datasets is pervasive in a lot of data mining domains. Imbalanced classification has been a hot topic in the academic community. From data level to algorithm level, a lot of solutions have been proposed to tackle the problems resulted from imbalanced datasets. SMOTE is the most popular data-level method and a lot of derivations based on it are developed to alleviate the problem of class imbalance. Our investigation indicates that there are severe flaws in SMOTE. We propose a new oversampling method SNOCC that can compensate the defects of SMOTE. In SNOCC, we increase the number of seed samples and that renders the new samples not confine in the line segment between two seed samples in SMOTE. We employ a novel algorithm to find the nearest neighbors of samples, which is different to the previous ones. These two improvements make the new samples created by SNOCC naturally reproduce the distribution of original seed samples. Our experiment results show that SNOCC outperform SMOTE and CBSO (a SMOTE-based method).

Downloads

Download data is not yet available.

Oversampling Method for Imbalanced Classification

Authors

Keywords:

Abstract

Downloads

Downloads

Published

How to Cite

Issue

Section

Information

Make a Submission

Keywords