Comparison of Automatic Modeling Effects of Ymodel, Weka, Rapidminer

Objective:To compare the automatic modeling effects of Weka, Rapidminer, and Ymodel

Data to be used:5 pieces of data in total, 3 pieces of classification, and 2 pieces of regression

2 classic Kaggle cases and 3 real business data

Titanic DataClassificationKaggle
House Price PredictionRegressionKaggle
Credit Company User Overdue PredictionClassification
Claims prediction of insurance company policiesClassification
Second-hand car transaction price predictionRegression

Due to the limited data size of Rapidminer's free version of 10000 items, three real business data was sampled, with sample sizes controlled within a few thousand items. It is not possible to conduct large data volume testing.

Product introduction: Weka is open source, and the automatic modeling function is an extension module of Weka, which is free to use. Rapidminer is a commercial software. Although it has a free version, the auto model function will be charged.

Overall user experience:Ymodel has the fastest modeling speed. Rapidminer is relatively fast in model building, and when there are many variables, the modeling time increases significantly. Weka modeling requires setting the modeling time beforehand, and the modeling speed is also relatively slow. In Weka, sometimes it is necessary to manually handle some variable types in order to be recognized by automatic modeling. In terms of automatic modeling functionality, Weka's experience is relatively poor.

Testing method:All data is divided into a training set and a prediction set, and the prediction results are exported and scored uniformly.

Test results:

1. Titanic Survival Prediction - Classification

Training data: 802 items, 12 variables

The ratio of positive and negative samples is approximately 3:5

WekaRapidminerYmodel
Accuracy0.7220.7870.775
Precision0.8620.8090.857
Recall0.5560.7560.667
Specificity0.9090.8180.886
F10.6760.7820.75
AUC0.7930.847
Ranking321

It is unable to output probability values in Weka (or possibly not finding how to output), therefore unable to calculate AUC.

2. House Price Prediction - Regression

WekaRapidminerYmodel
Mse4.17E81.41E99.85E8
Rmse204303753931385
Mae141641945916378
Mape9.10811.3179.921
R20.8890.7550.829
Ranking132

3. Credit Company User Overdue Prediction - Classification

Training data: 8938 items, 56 variables

The ratio of positive and negative samples is approximately 1:8

WekaRapidminerYmodel
Accuracy0.8780.8800.804
Precision-0.4710.281
Recall00.0630.409
Specificity10.990.858
F1-0.1110.333
AUC0.7290.742
Ranking321

On this data, the Weka model failed and did not capture any positive sample.

4. Claims prediction of insurance company policies - classification

Training data: 3470 items, 29 variables

The ratio of positive and negative samples is approximately 1:7

WekaRapidminerYmodel
Accuracy0.9050.9490.882
Precision0.0510.0330.022
Recall0.2640.0690.139
Specificity0.9160.9650.895
F10.0860.0450.038
AUC0.6420.638
Ranking123

5. Second-hand car transaction price prediction

WekaRapidminerYmodel
Mse277992784667169429967
Rmse166729103070
Mae83515801537
Mape277554
R20.9410.8210.801
Ranking123

Overall evaluation:Among the 5 data samples used in this testing, the rankings vary depending on the data, but the difference in indexes is not significant, and the overall performance of Ymodel is quite good. In comparison, Weka performs well in regression model, Ymodel performs well in classification model, and Rapidminer is in the middle.

Leave a Reply

Discover more from esProc SPL Official Blog

Subscribe now to keep reading and get access to the full archive.

Continue reading