A Technique to Detect and Categorize COVID-19 Disinformation from Amharic and English Twitter feeds among Ethiopian Twitter Users
Keywords:
COVID-19 ·Coronavirus, Disinformation Detection, Logistic regression, Amharic tweets, Tweet classificationAbstract
Since its announcement as pandemic by WHO, COVID-19 virus has been a devastating experience that the world has faced. Many have lost their precious lives to COVID-19 virus. As the virus is not well-known to medical communities during its eruption, many have put forward speculations as to how the virus comes into existence, how it is transmitted, how it could be prevented. Unsubstantiated rumors, fake remedies and falsehoods quickly surfaced social media platforms. Hence; identifying disinformation regarding COVID-19 for a regular person is challenging specially when much is not known and authorities does not have utilities to guide public opinions regarding the virus. If unverified rumors and speculations people hold regarding COVID-19 are shared on social media like twitter, people might subscribe to the ideas embedded in the rumors and might apply to the real world by considering these rumors a legitimate claim. This thesis work has proposed a technique that will detect disinformation twitter feeds/claims. The detection model works for both Amharic and English languages. Amharic dataset contains 205 non-disinformation claims and 196 disinformation claims; similarly, English dataset included 568 non-disinformation and 585 disinformation claims. The developed technique has been deployed on Flask web application to serve as COVID-19 disinformation detection platform. The developed detection technique has F-1 score of 90.48% for English dataset while the same model recorded F-1 score of 91.4% for Amharic dataset. This thesis work could be expanded further by incorporating other social media sources like, Facebook, WhatsApp, Telegram, etc. as the source COVID-19 conversations to further enrich the datasets.