語系:
繁體中文
English
說明(常見問題)
登入
回首頁
切換:
標籤
|
MARC模式
|
ISBD
Towards Machine Learning for Gulf Dialectical Arabic Malicious Content Detection in Social Media
紀錄類型:
書目-語言資料,印刷品 : Monograph/item
正題名/作者:
Towards Machine Learning for Gulf Dialectical Arabic Malicious Content Detection in Social Media/ Dema Alorini.
作者:
Alorini, Dema,
面頁冊數:
1 electronic resource (72 pages)
附註:
Source: Dissertations Abstracts International, Volume: 80-09, Section: B.
Contained By:
Dissertations Abstracts International80-09B.
標題:
Computer science. -
電子資源:
http://pqdd.sinica.edu.tw/twdaoapp/servlet/advanced?query=13425964
ISBN:
9780438969087
Towards Machine Learning for Gulf Dialectical Arabic Malicious Content Detection in Social Media
Alorini, Dema,
Towards Machine Learning for Gulf Dialectical Arabic Malicious Content Detection in Social Media
[electronic resource] /Dema Alorini. - 1 electronic resource (72 pages)
Source: Dissertations Abstracts International, Volume: 80-09, Section: B.
Dialectical Arabic (DA) is the native tongue of Arabic speakers and is used in communication whether in person or in social media, while Modern Standard Arabic (MSA) is only used in Arabic written forms, formal communication, media, but rarely in speeches. In Machine Translation (MT) most of the work and research did is targeting MSA without focusing on Gulf DA. The main problem of translating dialectical Arabic to English is the lack of linguistic resources. Therefore, a dataset is built consisting of 2000 Gulf DA sentences translated into English using Arab annotators. The resource of this dataset is from Twitter's tweets that covers all the dialects of Gulf including Saudi, Kuwaiti, Emirati, Bahraini, Omani, and Qatari using the location listed in the account. One of the largest domains for written communication is the online domain. Today, social media have become widely used among people of different ages and nationalities. The usage of social media is increasing rapidly in the Arab region. One of the popular social networking sites for sharing news and spreading propaganda is Twitter. Spammers use these sites to disseminate adult content and false political news through Arabic tweets. Within the Arab region, distributing adult materials is illegal and some governments attempted to block malicious URLs, but they fail most of the time. Tweets do not only contain information about opinions, news, and conversations, but also malicious content such as false information, malicious links, and other types of cyber threats. Therefore, those tweets need to be analyzed and identified in order to reduce risks associated with them. Tweets from the Gulf region are not written in the MSA which is used in most translation systems as an Arabic source. This dissertation presents both user and content attributes to differentiate between legitimate and illegitimate users. Then these attributes are used with machine learning algorithms to detect spam on Twitter. The algorithms use Naive Bayes (NB) and Support Vector Machine (SVM) classification methods. Results show that NB produces more accurate outcomes in detecting Arabic spam.
English
ISBN: 9780438969087Subjects--Topical Terms:
573171
Computer science.
Towards Machine Learning for Gulf Dialectical Arabic Malicious Content Detection in Social Media
LDR
:03585nam a22003733i 4500
001
1172816
005
20260622113211.5
006
m o d
007
cr|nu||||||||
008
260803s2018 miu||||||m |||||||eng d
020
$a
9780438969087
035
$a
(MiAaPQD)AAI13425964
035
$a
(MiAaPQD)howard:11571
035
$a
AAI13425964
040
$a
MiAaPQD
$b
eng
$c
MiAaPQD
$e
rda
100
1
$a
Alorini, Dema,
$e
author.
$3
1503413
245
1 0
$a
Towards Machine Learning for Gulf Dialectical Arabic Malicious Content Detection in Social Media
$c
Dema Alorini.
$h
[electronic resource] /
264
1
$a
Ann Arbor :
$b
ProQuest Dissertations & Theses,
$c
2018
300
$a
1 electronic resource (72 pages)
336
$a
text
$b
txt
$2
rdacontent
337
$a
computer
$b
c
$2
rdamedia
338
$a
online resource
$b
cr
$2
rdacarrier
500
$a
Source: Dissertations Abstracts International, Volume: 80-09, Section: B.
500
$a
Publisher info.: Dissertation/Thesis.
500
$a
Advisors: Rawat, Danda B. Committee members: Garuba, Moses; Peter, Keiller A.; Rawat, Danda B.; Rubaai, Ahmed; Shetty, Sachin.
502
$b
Ph.D.
$c
Howard University
$d
2018.
520
#
$a
Dialectical Arabic (DA) is the native tongue of Arabic speakers and is used in communication whether in person or in social media, while Modern Standard Arabic (MSA) is only used in Arabic written forms, formal communication, media, but rarely in speeches. In Machine Translation (MT) most of the work and research did is targeting MSA without focusing on Gulf DA. The main problem of translating dialectical Arabic to English is the lack of linguistic resources. Therefore, a dataset is built consisting of 2000 Gulf DA sentences translated into English using Arab annotators. The resource of this dataset is from Twitter's tweets that covers all the dialects of Gulf including Saudi, Kuwaiti, Emirati, Bahraini, Omani, and Qatari using the location listed in the account. One of the largest domains for written communication is the online domain. Today, social media have become widely used among people of different ages and nationalities. The usage of social media is increasing rapidly in the Arab region. One of the popular social networking sites for sharing news and spreading propaganda is Twitter. Spammers use these sites to disseminate adult content and false political news through Arabic tweets. Within the Arab region, distributing adult materials is illegal and some governments attempted to block malicious URLs, but they fail most of the time. Tweets do not only contain information about opinions, news, and conversations, but also malicious content such as false information, malicious links, and other types of cyber threats. Therefore, those tweets need to be analyzed and identified in order to reduce risks associated with them. Tweets from the Gulf region are not written in the MSA which is used in most translation systems as an Arabic source. This dissertation presents both user and content attributes to differentiate between legitimate and illegitimate users. Then these attributes are used with machine learning algorithms to detect spam on Twitter. The algorithms use Naive Bayes (NB) and Support Vector Machine (SVM) classification methods. Results show that NB produces more accurate outcomes in detecting Arabic spam.
546
$a
English
590
$a
School code: 0088
650
# 4
$a
Computer science.
$3
573171
690
$a
0984
710
2 #
$a
Howard University.
$b
Systems and Computer Science.
$e
degree granting institution.
$3
1503414
720
1
$a
Rawat, Danda B.
$e
degree supervisor.
773
0 #
$t
Dissertations Abstracts International
$g
80-09B.
790
$a
0088
791
$a
Ph.D.
792
$a
2018
856
4 0
$u
http://pqdd.sinica.edu.tw/twdaoapp/servlet/advanced?query=13425964
筆 0 讀者評論
多媒體
評論
新增評論
分享你的心得
Export
取書館別
處理中
...
變更密碼[密碼必須為2種組合(英文和數字)及長度為10碼以上]
登入