How to Use Binary Encoding to Handle Categorical Variables in Machine Learning

In machine learning, categorical variables are those that take on a limited number of discrete values, such as color, gender, or country of origin. To use categorical variables in a machine learning model, they need to be encoded into numerical values that the model can process. One common technique for encoding categorical variables is binary encoding, which creates binary columns for each category in a variable. In this tutorial, we will walk through how to encode categorical variables with binary encoding in Python.
Step 1: Importing Libraries
To use binary encoding in Python, we need to import the necessary libraries. We will use the category_encoders library, which contains the BinaryEncoder class that we will use to encode our categorical variables. We will also import the pandas library for data manipulation.
import pandas as pd
import category_encoders as ceStep 2: Loading Data
Next, we will load a sample dataset to work with. For this tutorial, we will use the titanic dataset from the seaborn library, which contains information about passengers on the Titanic ship. We will load the dataset using the load_dataset function from seaborn.
import seaborn as sns
titanic = sns.load_dataset('titanic')Here are the first few rows of the input dataset:

Step 3: Creating Categorical Variables
Before we can encode categorical variables with binary encoding, we need to create some categorical variables. For this tutorial, we will create a categorical variable for the class column in the titanic dataset. We can do this using the astype method in pandas to convert the column to the category data type.
titanic['class'] = titanic['class'].astype('category')Step 4: Encoding Categorical Variables with Binary Encoding
Now that we have a categorical variable to work with, we can encode it with binary encoding using the BinaryEncoder class from the category_encoders library. We will create an instance of the BinaryEncoder class and pass in the name of the column we want to encode as the cols parameter.
encoder = ce.BinaryEncoder(cols=['class'])
titanic_encoded = encoder.fit_transform(titanic)This will create new columns in the titanic_encoded dataset for each category in the class column, with a binary value of 1 or 0 indicating whether that category is present or not. The number of binary columns created for each categorical variable depends on the number of unique categories in that variable. For example, since the class column in the titanic dataset has 3 unique categories (First, Second, and Third), binary encoding will create 2 new columns to represent these categories.
Step 5: Examining Encoded Data
Finally, we can examine the encoded data to make sure that the binary encoding was done correctly. We can print out the first few rows of the titanic_encoded dataset to see the new binary columns that were created.
print(titanic_encoded.head())This will show us the following output:

As we can see, the class_0, class_1, and class_2 columns represent the three categories in the class column (First, Second, and Third). The binary encoding has converted these categories into binary values, with a 1 in the corresponding column indicating that the category is present.
Conclusion
Binary encoding is a useful technique for encoding categorical variables in machine learning. It creates binary columns for each category in a variable, allowing machine learning models to process categorical data as numerical data. In this tutorial, we walked through how to encode categorical variables with binary encoding in Python using the category_encoders library.




