Summary: | The feature selection method enhances machine learning performance by enhancing learning precision. Determining the optimal feature selection method for a given machine learning task involving big-dimension data is crucial. Therefore, the purpose of this study is to make a comparison of feature selection methods highlighting several filters (information gain, chi-square, ReliefF) and embedded (Lasso, Ridge) hybrid with logistic regression (LR). A sample size of n=100, 75 is chosen randomly, and the reduction features d=50, 22, and 10 are applied. The procedure for feature reduction makes use of the entire sample sizes. Each sample size's results are compared, including tests with no feature selection process. The results indicate that LR+ReliefF is the best method for mammary cancer data, whereas LR+IG is the best for prostatic cancer data, making the filter more suitable than embedded for big-dimension data. This study revealed that the sample's features and size influence the most effective method for selecting features from big-dimension data. Therefore, it provides insight into the most effective methods for particular features and sample sizes in high-dimensional data. © 2024, Institute of Advanced Engineering and Science. All rights reserved.
|