2020-12-09

How to reduce space between bars in ggplot2

Title

Problem

We want to reduce the space between bars using geom_bar with ggplot2.

Vertical bar plot

Horizontal bar plot

library(ggthemes)
library(ggplot2)
df <- data.frame(x = c("firm", "vendor"), y = c(50, 20))

# Vertical
ggplot(df, aes(x = x, y = y)) + 
  geom_bar(stat = "identity", width = 0.4) + 
  theme_tufte() + 
  labs(x = "", y = "")

# Horizontal
ggplot(df, aes(x = x, y = y)) + 
  geom_bar(stat = "identity", width = 0.4) + 
  theme_tufte() +  coord_flip() +
  labs(x = "", y = "")

Solution

We just need to to introduce theme(aspect.ratio = xx), then we need to play with the ratio to find the desired width. To flip cartesian coordinates we use coord_flip()


# Vertical
ggplot(df, aes(x = x, y = y)) + 
  geom_bar(stat = "identity") + 
  theme_tufte() + theme(aspect.ratio = 2) +
  labs(x = "", y = "")

# Horizontal
ggplot(df, aes(x = x, y = y)) + 
  geom_bar(stat = "identity") + 
  theme_tufte() + theme(aspect.ratio = .2) +
  coord_flip() +
  labs(x = "", y = "")

Results

Vertical bar plot

Horizontal bar plot

Related posts

2020-12-08

Cómo combinar múltiples condiciones para filtrar un data frame usando "OR”

Title

Problema

Queremos filtrar un data frame basándonos en múltiples condiciones usando el operador "O". En nuestro ejemplo filtraremos aquellas filas del data frame donde la v1 sea menor que 0.5 o donde v2 sea igual a g.

Data frame original

           v1 v2
1  0.26550866  a
2  0.37212390  b
3  0.57285336  c
4  0.90820779  d
5  0.20168193  e
6  0.89838968  f
7  0.94467527  g
8  0.66079779  h
9  0.62911404  i
10 0.06178627  j

Resultado esperado

          v1 v2
1 0.26550866  a
2 0.37212390  b
3 0.20168193  e
4 0.94467527  g
5 0.06178627  j
set.seed(1)
df <- data.frame(v1 = runif(10), v2 = letters[1:10])

Solución

Hay múltiples opciones:

  • Funciones del paquete base
  • subset(df , v1 < 0.5 | v2 == "g")
    df[which(df$v1 < 0.5 | df$v2 == "g"), ]
    

  • Operatores [ y [[
  • df[df[1] < 0.5 | df[2] == "g", ] 
    df[df[[1]] < 0.5 | df[[2]] == "g", ] 
    df[df["v1"] < 0.5 | df["v2"] == "g", ]
    

    df$name is equivalent to df[["name", exact = FALSE]]

  • dplyr
  • library(dplyr)
    filter(df, v1 < 0.5 | v2 == "g")
    

  • sqldf
  • library(sqldf)
    sqldf('SELECT *
          FROM df 
          WHERE v1 < 0.5 OR v2 = "g")
    

Referencias

2020-12-07

How to combine multiple conditions to subset a data frame using “OR”?

Title

Problem

We want to subset a data frame based on multiple conditions using "OR". In our example we want to subset the data frame to include all rows where v1 is less than 0.5 or rows where v2 is equal to g.

Original data frame

           v1 v2
1  0.26550866  a
2  0.37212390  b
3  0.57285336  c
4  0.90820779  d
5  0.20168193  e
6  0.89838968  f
7  0.94467527  g
8  0.66079779  h
9  0.62911404  i
10 0.06178627  j

Expected output

          v1 v2
1 0.26550866  a
2 0.37212390  b
3 0.20168193  e
4 0.94467527  g
5 0.06178627  j
set.seed(1)
df <- data.frame(v1 = runif(10), v2 = letters[1:10])

Solution

There are multiple options:

  • Base functions
  • subset(df , v1 < 0.5 | v2 == "g")
    df[which(df$v1 < 0.5 | df$v2 == "g"), ]
    

  • Operators [ and [[
  • df[df[1] < 0.5 | df[2] == "g", ] 
    df[df[[1]] < 0.5 | df[[2]] == "g", ] 
    df[df["v1"] < 0.5 | df["v2"] == "g", ]
    

    df$name is equivalent to df[["name", exact = FALSE]]

  • dplyr
  • library(dplyr)
    filter(df, v1 < 0.5 | v2 == "g")
    

  • sqldf
  • library(sqldf)
    sqldf('SELECT *
          FROM df 
          WHERE v1 < 0.5 OR v2 = "g")
    

References

2020-12-04

How to remove whiskers in a boxplot in ggplot2

Title

Problem

We want to remove whiskers in a boxplot created with ggplot2.

library(ggplot2)
p <- ggplot(mtcars, aes(factor(cyl), mpg))
p + geom_boxplot()

Solution

We specify coef = 0, overwriting the default value (coef = 1.5).

p <- ggplot(mtcars, aes(factor(cyl), mpg))
p + geom_boxplot(outlier.size = 0, coef = 0)

Related posts

2020-12-03

How to remove whiskers in a boxplot with the R Base Package

Title

Problem

We want to remove whiskers in a boxplot created with the R base package.

boxplot(mpg ~ cyl, data = mtcars)

Solution

We pass the arguments: whisklty = 0, staplelty = 0.

whisklty - to control the whisker line type (default: "dashed").
staplelty - to control the staple (= end of whisker) line type.

boxplot(mpg ~ cyl, data = mtcars, whisklty = 0, staplelty = 0)

Related posts

2020-12-01

Mostrar las áreas de densidad deseadas con stat_density_2d en ggplot2

Title

Problema

Tenemos el siguiente diagrama de dispersión para dos variables categóricas.

Cuando creamos un gráfico de densidad en 2D, obtenemos varios contornos de densidad. Queremos controlar precisamente el número de contornos.

library(ggplot2)
set.seed(123)
plot_data <-
  data.frame(
    X = c(rnorm(300, 3, 2.5), rnorm(150, 7, 2)),
    Y = c(rnorm(300, 6, 2.5), rnorm(150, 2, 2)),
    Label = c(rep('A', 300), rep('B', 150))
  )

ggplot(plot_data, aes(X, Y, colour = Label)) + geom_point()
ggplot(plot_data, aes(X, Y)) +
  stat_density_2d(geom = "polygon", aes(alpha = ..level.., fill = Label))

Solución

  • Opción 1
  • Añadiendo stat_density_2d con el argumento bins (número de contornos) evitamos una sobrecagarda de información, controlamos y centramos la atención en un número concreto de contornos de densidad.

    ggplot(plot_data, aes(X, Y, group = Label)) +
      stat_density_2d(geom = "polygon",
                      aes(alpha = ..level.., fill = Label),
                      bins = 4) 
    
  • Opción 2
  • Asignando manualmente los colores, especificando NA para aquellos niveles que no queremos mostrar. La principal desventaja es que necesitamos saber por adelantado el número de valores que scale_fill_manual necesita. En nuestro ejemplo, necesitamos especificar 7 valores en la escala manual.

    ggplot(plot_data, aes(X, Y, group = Label)) +
      stat_density_2d(geom = "polygon", aes(fill = as.factor(..level..))) +
      scale_fill_manual(values = c(NA, NA, NA, "#BDD7E7", "#6BAED6", "#3182BD", "#08519C"))
    

Referencias

Show only high density areas with stat_density_2d with ggplot2

Title

Problem

We have the following scatterplot for two categorical variables.

When we create a 2D-density plot, we obtain overlapping densities. We want to control the number of contour bins.

library(ggplot2)
set.seed(123)
plot_data <-
  data.frame(
    X = c(rnorm(300, 3, 2.5), rnorm(150, 7, 2)),
    Y = c(rnorm(300, 6, 2.5), rnorm(150, 2, 2)),
    Label = c(rep('A', 300), rep('B', 150))
  )

ggplot(plot_data, aes(X, Y, colour = Label)) + geom_point()
ggplot(plot_data, aes(X, Y)) +
  stat_density_2d(geom = "polygon", aes(alpha = ..level.., fill = Label))

Solution

  • Option 1
  • By adding to stat_density_2d the argument bins (number of contour bins) we definitely avoid overplotting, control and draw the attention to a number of density areas in a very economical fashion.

    ggplot(plot_data, aes(X, Y, group = Label)) +
      stat_density_2d(geom = "polygon",
                      aes(alpha = ..level.., fill = Label),
                      bins = 4) 
    
  • Option 2
  • Assigning manually the colours, NA for those levels we do not want to plot. The main disadvantage is that we should know the number of values needed by scale_fill_manual in advance. In our example, we need to pass 7 values in manual scale.

    ggplot(plot_data, aes(X, Y, group = Label)) +
      stat_density_2d(geom = "polygon", aes(fill = as.factor(..level..))) +
      scale_fill_manual(values = c(NA, NA, NA, "#BDD7E7", "#6BAED6", "#3182BD", "#08519C"))
    

References

Nube de datos