Logo Questions Linux Laravel Mysql Ubuntu Git Menu
 

How to do named entity recognition (NER) using quanteda?

Tags:

r

quanteda

Having a dataframe with text

df = data.frame(id=c(1,2), text = c("My best friend John works and Google", "However he would like to work at Amazon as he likes to use python and stay at Canada")

Without any preprocessing

How is it possible to extract name entity recognition like this

Example results words

dfresults = data.frame(id=c(1,2), ner_words = c("John, Google", "Amazon, python, Canada")
like image 604
Nathalie Avatar asked Oct 17 '25 02:10

Nathalie


1 Answers

You can do this without quanteda, using the spacyr package -- a wrapper around the spaCy library mentioned in your linked article.

Here, I have slightly edited your input data.frame.

df <- data.frame(id = c(1, 2), 
                 text = c("My best friend John works at Google.", 
                          "However he would like to work at Amazon as he likes to use Python and stay in Canada."),
                 stringsAsFactors = FALSE)

Then:

library("spacyr")
library("dplyr")

# -- need to do these before the next function will work:
# spacy_install()
# spacy_download_langmodel(model = "en_core_web_lg")

spacy_initialize(model = "en_core_web_lg")
#> Found 'spacy_condaenv'. spacyr will use this environment
#> successfully initialized (spaCy Version: 2.0.10, language model: en_core_web_lg)
#> (python options: type = "condaenv", value = "spacy_condaenv")

txt <- df$text
names(txt) <- df$id

spacy_parse(txt, lemma = FALSE, entity = TRUE) %>%
    entity_extract() %>%
    group_by(doc_id) %>%
    summarize(ner_words = paste(entity, collapse = ", "))
#> # A tibble: 2 x 2
#>   doc_id ner_words             
#>   <chr>  <chr>                 
#> 1 1      John, Google          
#> 2 2      Amazon, Python, Canada
like image 104
Ken Benoit Avatar answered Oct 18 '25 20:10

Ken Benoit



Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!