Word Embeddings for Automatic Equalization in Audio Mixing

Venkatesh, S; Moffat, D; Miranda, Eduardo

dc.contributor.author	Venkatesh, S
dc.contributor.author	Moffat, D
dc.contributor.author	Miranda, Eduardo
dc.date.accessioned	2022-09-26T18:31:25Z
dc.date.issued	2022-09-12
dc.identifier.issn	1549-4950
dc.identifier.uri	http://hdl.handle.net/10026.1/19638
dc.description.abstract	In recent years, machine learning has been widely adopted to automate the audio mixing process. Automatic mixing systems have been applied to various audio effects such as gain-adjustment, equalization, and reverberation. These systems can be controlled through visual interfaces, providing audio examples, using knobs, and semantic descriptors. Using semantic descriptors or textual information to control these systems is an effective way for artists to communicate their creative goals. In this paper, we explore the novel idea of using word embeddings to represent semantic descriptors. Word embeddings are generally obtained by training neural networks on large corpora of written text. These embeddings serve as the input layer of the neural network to create a translation from words to EQ settings. Using this technique, the machine learning model can also generate EQ settings for semantic descriptors that it has not seen before. We compare the EQ settings of humans with the predictions of the neural network to evaluate the quality of predictions. The results showed that the embedding layer enables the neural network to understand semantic descriptors. We observed that the models with embedding layers perform better than those without embedding layers, but still not as good as human labels.
dc.format.extent	753-763
dc.language.iso	en
dc.publisher	Audio Engineering Society
dc.subject	Audio Mixing
dc.subject	Automatic Mixing
dc.subject	Equalization
dc.subject	Semantic Word Vectors
dc.title	Word Embeddings for Automatic Equalization in Audio Mixing
dc.type	journal-article
dc.type	Journal Article
plymouth.issue	9
plymouth.volume	70
plymouth.publisher-url	http://www.aes.org/e-lib/browse.cfm?elib=21887
plymouth.publication-status	Published online
plymouth.journal	Journal of the Audio Engineering Society
dc.identifier.doi	10.17743/jaes.2022.0047
plymouth.organisational-group	/Plymouth
plymouth.organisational-group	/Plymouth/Faculty of Arts, Humanities and Business
plymouth.organisational-group	/Plymouth/Users by role
plymouth.organisational-group	/Plymouth/Users by role/Academics
dcterms.dateAccepted	2022-07-25
dc.rights.embargodate	2022-10-1
dc.rights.embargoperiod	Not known
rioxxterms.funder	Engineering and Physical Sciences Research Council
rioxxterms.identifier.project	Radio Me: Real-time Radio Remixing for people with mild to moderate dementia who live alone, incorporating Agitation Reduction, and Reminders
rioxxterms.version	Version of Record
rioxxterms.versionofrecord	10.17743/jaes.2022.0047
rioxxterms.licenseref.uri	http://www.rioxx.net/licenses/all-rights-reserved
rioxxterms.licenseref.startdate	2022-09-12
rioxxterms.type	Journal Article/Review
plymouth.funder	Radio Me: Real-time Radio Remixing for people with mild to moderate dementia who live alone, incorporating Agitation Reduction, and Reminders::Engineering and Physical Sciences Research Council

Files in this item

Name:: 21887.pdf
Size:: 830.9Kb
Format:: PDF

View/Open

Name:: UoP_Deposit_Agreement v1.1 ...
Size:: 125.4Kb
Format:: PDF

View/Open

This item appears in the following Collection(s)

University of Plymouth Research Outputs

Show simple item record