Access the Google Data Commons API V2. Data Commons provides programmatic access to statistical and demographic data from dozens of sources organized in a knowledge graph.
You can install datacommons from CRAN via:
install.packages("datacommons")You can install the development version of datacommons
from GitHub with:
# install.packages("pak")
pak::pak("tidy-intelligence/r-datacommons"):bulb: A detailed walkthrough of census data analysis using
datacommonsis available in the corresponding vignette.
Load the package:
library(datacommons)Get a free API key for Data Commons here. Set
the Data Commons API key as the DATACOMMONS_API_KEY
environment variable using the helper function and restart your R
session to load the key:
dc_set_api_key("YOUR_API_KEY")If you want to use a custom
Data Commons instance, then you can also set the
DATACOMMONS_BASE_URL environment varibale on the project or
global level:
dc_set_base_url("YOUR_BASE_URL")Get a data frame with US population data from World Development Indicators:
country_level <- dc_get_observations(
date = "all",
variable_dcids = "Count_Person",
entity_dcids = "country/USA",
return_type = "data.frame",
filter_facet_ids = "18369491376878146239"
)
head(country_level, 5)
#> entity_dcid entity_name variable_dcid variable_name date value
#> 1 country/USA United States Count_Person Total population 1960 180671000
#> 2 country/USA United States Count_Person Total population 1961 183691000
#> 3 country/USA United States Count_Person Total population 1962 186538000
#> 4 country/USA United States Count_Person Total population 1963 189242000
#> 5 country/USA United States Count_Person Total population 1964 191889000
#> facet_id facet_name
#> 1 18369491376878146239 WorldDevelopmentIndicators
#> 2 18369491376878146239 WorldDevelopmentIndicators
#> 3 18369491376878146239 WorldDevelopmentIndicators
#> 4 18369491376878146239 WorldDevelopmentIndicators
#> 5 18369491376878146239 WorldDevelopmentIndicatorsIf you want to get different population numbers from the US Census on the state level:
state_level <- dc_get_observations(
variable_dcids = "Count_Person",
date = 2021,
parent_entity = "country/USA",
entity_type = "State",
return_type = "data.frame",
filter_facet_ids = "8912910856362438925"
)
head(state_level, 5)
#> entity_dcid entity_name variable_dcid variable_name date value
#> 1 geoId/01 Alabama Count_Person Total population 2021 5039877
#> 2 geoId/02 Alaska Count_Person Total population 2021 732673
#> 3 geoId/04 Arizona Count_Person Total population 2021 7276316
#> 4 geoId/05 Arkansas Count_Person Total population 2021 3025891
#> 5 geoId/06 California Count_Person Total population 2021 39237836
#> facet_id facet_name
#> 1 8912910856362438925 USCensusPEP_Annual_Population
#> 2 8912910856362438925 USCensusPEP_Annual_Population
#> 3 8912910856362438925 USCensusPEP_Annual_Population
#> 4 8912910856362438925 USCensusPEP_Annual_Population
#> 5 8912910856362438925 USCensusPEP_Annual_PopulationContributions to datacommons are welcome! If you’d like
to contribute, please follow these steps: