We will load clustermole along with dplyr to help with summarizing the data.
You can use clustermole as a simple database and get a table of all cell type markers.
markers <- clustermole_markers(species = "hs")
markers
#> # A tibble: 521,262 × 8
#> celltype_full db species organ celltype gene_original gene n_genes
#> <chr> <chr> <chr> <chr> <chr> <chr> <chr> <int>
#> 1 (Pro-) Subiculum | … ScTy… "" Hipp… (Pro-) … ADAMTS2 ADAM… 10
#> 2 (Pro-) Subiculum | … ScTy… "" Hipp… (Pro-) … FN1 FN1 10
#> 3 (Pro-) Subiculum | … ScTy… "" Hipp… (Pro-) … KLHL1 KLHL1 10
#> 4 (Pro-) Subiculum | … ScTy… "" Hipp… (Pro-) … LIPM LIPM 10
#> 5 (Pro-) Subiculum | … ScTy… "" Hipp… (Pro-) … NPSR1 NPSR1 10
#> 6 (Pro-) Subiculum | … ScTy… "" Hipp… (Pro-) … NTS NTS 10
#> 7 (Pro-) Subiculum | … ScTy… "" Hipp… (Pro-) … RAB38 RAB38 10
#> 8 (Pro-) Subiculum | … ScTy… "" Hipp… (Pro-) … RXFP1 RXFP1 10
#> 9 (Pro-) Subiculum | … ScTy… "" Hipp… (Pro-) … STAC STAC 10
#> 10 (Pro-) Subiculum | … ScTy… "" Hipp… (Pro-) … TLE4 TLE4 10
#> # ℹ 521,252 more rowsEach row contains a gene and a cell type associated with it. The
gene column is the gene symbol (human or mouse), the
gene_original column is the gene symbol from the source
database, and the celltype_full column contains the
detailed cell type string including the species and the original
database.
Cell types by species
Check the number of cell types per species (not available for all cell types).
Cell types by organ
Check the number of available cell types per organ (not available for all cell types).
distinct(markers, celltype_full, organ) |> count(organ, sort = TRUE)
#> # A tibble: 383 × 2
#> organ n
#> <chr> <int>
#> 1 "" 3686
#> 2 "Brain" 798
#> 3 "Lung" 659
#> 4 "Liver" 534
#> 5 "Peripheral blood" 521
#> 6 "Skin" 497
#> 7 "Kidney" 428
#> 8 "Bone marrow" 427
#> 9 "Pancreas" 315
#> 10 "Breast" 292
#> # ℹ 373 more rowsPackage version
Check the package version since the database contents may change.
packageVersion("clustermole")
#> [1] '1.1.1.9000'