語系:
繁體中文
English
說明(常見問題)
登入
回首頁
切換:
標籤
|
MARC模式
|
ISBD
Automatic Lyrics Transcription of Polyphonic Music
紀錄類型:
書目-語言資料,印刷品 : Monograph/item
正題名/作者:
Automatic Lyrics Transcription of Polyphonic Music / Xiaoxue Gao.
作者:
Gao, Xiaoxue,
面頁冊數:
1 electronic resource (182 pages)
附註:
Source: Dissertations Abstracts International, Volume: 84-09, Section: B.
Contained By:
Dissertations Abstracts International84-09B.
標題:
Mathematics. -
電子資源:
http://pqdd.sinica.edu.tw/twdaoapp/servlet/advanced?query=30340048
ISBN:
9798374486216
Automatic Lyrics Transcription of Polyphonic Music
Gao, Xiaoxue,
Automatic Lyrics Transcription of Polyphonic Music
[electronic resource] /Xiaoxue Gao. - 1 electronic resource (182 pages)
Source: Dissertations Abstracts International, Volume: 84-09, Section: B.
Automatic Lyrics Transcription of polyphonic music (ALTP) aims to recognize the sung lyrics from singing vocals in the presence of instrumental music accompaniment, and it serves as a desired technology to develop for many applications such as automatic generation of karaoke lyrical content, music video subtitling, and query-by-singing. Despite considerable research effort, there exist several challenges in ALTP, such as complicated multi-step training, imperfect singing vocal extraction, complex structure of polyphonic music, music genre discrepancy problem and background music interference problem. This thesis aims to study advanced deep learning techniques and end-to-end (E2E) frameworks to overcome the challenges for ALTP.Starting from the traditional multi-step framework, the first contribution of this thesis focuses to remedy the imperfect singing vocal extraction and background music interference problems, where music-robust model is introduced to combine the music-removed features and the music-present features to form a music-robust feature for acoustic modeling. These two sets of features complement each other, as the music-removed features focus on the lyrical content without background music interference while the music-present features compensate for the distortion caused by imperfect vocal extraction. The combination of the two features performs better than they are used alone, thus improving the robustness of acoustic model to the background music.The second contribution of this thesis focuses on overcoming the multistep training problem by firstly building E2E frameworks for both automatic lyrics transcription of monophonic singing (ALTM) and ALTP using transformer. Instead of relying on the traditional multi-step Kaldi system, the proposed E2E frameworks jointly model acoustic, lexical and linguistic models in a single neural network. The proposed frameworks are known to be the first to use transformer to capture temporal contextual information for ALTM and ALTP. Along with the transformer models, we further address the complex polyphonic music problem by formulating multi-task multi-transcribers that interprets the music information by chords transcription as a secondary task. Knowing that lyrics and chords are essential to represent singing vocal and background music, multi-transcribers are proposed to disentangle the lyrics from the chords in polyphonic music for effective lyrics transcription in a single step. These methodologies show competitive performances for ALTM and ALTP towards E2E era.The third contribution resolves the music genre discrepancy problem by proposing a genre-specific adapter for ALTP. ALTP is challenging as the background music and the singing style vary across music genres, which affects lyrics intelligibility of the song in different ways. A genre-conditioned network is proposed to transcribe the lyrics of polyphonic music using genre-conditioned adapters. The proposed network adopts pre-trained model parameters and incorporates the genre adapters to capture different genre peculiarities for lyrics-genre pairs, thereby providing genre-related knowledge to help with music interference problem.The final contribution of this thesis attempts to handle the background music interference problem by developing E2E joint-training frameworks with an integrated architecture for ALTP. Typically, lyrics transcription can be performed by a two-step pipeline, i.e. singing vocal extraction frontend, followed by a lyrics transcriber backend, where the frontend and backend are trained separately.
English
ISBN: 9798374486216Subjects--Topical Terms:
527692
Mathematics.
Automatic Lyrics Transcription of Polyphonic Music
LDR
:05015nam a22004333i 4500
001
1172863
005
20260622113217.5
006
m o d
007
cr|nu||||||||
008
260803s2022 miu||||||m |||||||eng d
020
$a
9798374486216
035
$a
(MiAaPQD)AAI30340048
035
$a
(MiAaPQD)USingapore233974
035
$a
AAI30340048
040
$a
MiAaPQD
$b
eng
$c
MiAaPQD
$e
rda
100
1
$a
Gao, Xiaoxue,
$e
author.
$3
1503478
245
1 0
$a
Automatic Lyrics Transcription of Polyphonic Music
$c
Xiaoxue Gao.
$h
[electronic resource] /
264
1
$a
Ann Arbor :
$b
ProQuest Dissertations & Theses,
$c
2022
300
$a
1 electronic resource (182 pages)
336
$a
text
$b
txt
$2
rdacontent
337
$a
computer
$b
c
$2
rdamedia
338
$a
online resource
$b
cr
$2
rdacarrier
500
$a
Source: Dissertations Abstracts International, Volume: 84-09, Section: B.
500
$a
Advisors: Sam Ge, Shuzhi; Haizhou, LI.
502
$b
Ph.D.
$c
National University of Singapore (Singapore)
$d
2022.
520
#
$a
Automatic Lyrics Transcription of polyphonic music (ALTP) aims to recognize the sung lyrics from singing vocals in the presence of instrumental music accompaniment, and it serves as a desired technology to develop for many applications such as automatic generation of karaoke lyrical content, music video subtitling, and query-by-singing. Despite considerable research effort, there exist several challenges in ALTP, such as complicated multi-step training, imperfect singing vocal extraction, complex structure of polyphonic music, music genre discrepancy problem and background music interference problem. This thesis aims to study advanced deep learning techniques and end-to-end (E2E) frameworks to overcome the challenges for ALTP.Starting from the traditional multi-step framework, the first contribution of this thesis focuses to remedy the imperfect singing vocal extraction and background music interference problems, where music-robust model is introduced to combine the music-removed features and the music-present features to form a music-robust feature for acoustic modeling. These two sets of features complement each other, as the music-removed features focus on the lyrical content without background music interference while the music-present features compensate for the distortion caused by imperfect vocal extraction. The combination of the two features performs better than they are used alone, thus improving the robustness of acoustic model to the background music.The second contribution of this thesis focuses on overcoming the multistep training problem by firstly building E2E frameworks for both automatic lyrics transcription of monophonic singing (ALTM) and ALTP using transformer. Instead of relying on the traditional multi-step Kaldi system, the proposed E2E frameworks jointly model acoustic, lexical and linguistic models in a single neural network. The proposed frameworks are known to be the first to use transformer to capture temporal contextual information for ALTM and ALTP. Along with the transformer models, we further address the complex polyphonic music problem by formulating multi-task multi-transcribers that interprets the music information by chords transcription as a secondary task. Knowing that lyrics and chords are essential to represent singing vocal and background music, multi-transcribers are proposed to disentangle the lyrics from the chords in polyphonic music for effective lyrics transcription in a single step. These methodologies show competitive performances for ALTM and ALTP towards E2E era.The third contribution resolves the music genre discrepancy problem by proposing a genre-specific adapter for ALTP. ALTP is challenging as the background music and the singing style vary across music genres, which affects lyrics intelligibility of the song in different ways. A genre-conditioned network is proposed to transcribe the lyrics of polyphonic music using genre-conditioned adapters. The proposed network adopts pre-trained model parameters and incorporates the genre adapters to capture different genre peculiarities for lyrics-genre pairs, thereby providing genre-related knowledge to help with music interference problem.The final contribution of this thesis attempts to handle the background music interference problem by developing E2E joint-training frameworks with an integrated architecture for ALTP. Typically, lyrics transcription can be performed by a two-step pipeline, i.e. singing vocal extraction frontend, followed by a lyrics transcriber backend, where the frontend and backend are trained separately.
546
$a
English
590
$a
School code: 1883
650
# 4
$a
Mathematics.
$3
527692
650
# 4
$a
Fine arts.
$3
1112523
650
# 4
$a
Speech.
$3
566021
650
# 4
$a
Acoustics.
$3
670692
650
# 4
$a
Error analysis.
$3
1466327
650
# 4
$a
Chords (Music).
$3
1503480
650
# 4
$a
Neural networks.
$3
1011215
650
# 4
$a
Voice recognition.
$3
1413635
650
# 4
$a
Polyphony.
$3
1479337
650
# 4
$a
Educational objectives.
$3
1473651
650
# 4
$a
Genre.
$3
1108142
650
# 4
$a
Lyrics.
$3
1503479
650
# 4
$a
Music.
$3
649088
650
# 4
$a
Language.
$3
571568
690
$a
0679
690
$a
0413
690
$a
0986
690
$a
0800
690
$a
0357
690
$a
0405
710
2 #
$a
National University of Singapore (Singapore).
$3
1193097
720
1
$a
Sam Ge, Shuzhi
$e
degree supervisor.
720
1
$a
Haizhou, LI
$e
degree supervisor.
773
0 #
$t
Dissertations Abstracts International
$g
84-09B.
790
$a
1883
791
$a
Ph.D.
792
$a
2022
856
4 0
$u
http://pqdd.sinica.edu.tw/twdaoapp/servlet/advanced?query=30340048
筆 0 讀者評論
多媒體
評論
新增評論
分享你的心得
Export
取書館別
處理中
...
變更密碼[密碼必須為2種組合(英文和數字)及長度為10碼以上]
登入