Manas: mining software repositories to assist AutoML

Nguyen, Giang; Islam, Md Johirul; Pan, Rangeet; Rajan, Hridesh

Manas: mining software repositories to assist AutoML

dc.contributor.author	Nguyen, Giang
dc.contributor.author	Islam, Md Johirul
dc.contributor.author	Pan, Rangeet
dc.contributor.author	Rajan, Hridesh
dc.contributor.department	Department of Computer Science
dc.date.accessioned	2022-12-09T15:18:49Z
dc.date.available	2022-12-09T15:18:49Z
dc.date.issued	2022-05
dc.description.abstract	Today deep learning is widely used for building software. A software engineering problem with deep learning is that finding an appropriate convolutional neural network (CNN) model for the task can be a challenge for developers. Recent work on AutoML, more precisely neural architecture search (NAS), embodied by tools like Auto-Keras aims to solve this problem by essentially viewing it as a search problem where the starting point is a default CNN model, and mutation of this CNN model allows exploration of the space of CNN models to find a CNN model that will work best for the problem. These works have had significant success in producing high-accuracy CNN models. There are two problems, however. First, NAS can be very costly, often taking several hours to complete. Second, CNN models produced by NAS can be very complex that makes it harder to understand them and costlier to train them. We propose a novel approach for NAS, where instead of starting from a default CNN model, the initial model is selected from a repository of models extracted from GitHub. The intuition being that developers solving a similar problem may have developed a better starting point compared to the default model. We also analyze common layer patterns of CNN models in the wild to understand changes that the developers make to improve their models. Our approach uses commonly occurring changes as mutation operators in NAS. We have extended Auto-Keras to implement our approach. Our evaluation using 8 top voted problems from Kaggle for tasks including image classification and image regression shows that given the same search time, without loss of accuracy, Manas produces models with 42.9% to 99.6% fewer number of parameters than Auto-Keras' models. Benchmarked on GPU, Manas' models train 30.3% to 641.6% faster than Auto-Keras' models.
dc.description.comments	This proceeding is published as Giang Nguyen, Md Johirul Islam, Rangeet Pan, and Hridesh Rajan. 2022. Manas: mining software repositories to assist AutoML. In Proceedings of the 44th International Conference on Software Engineering (ICSE '22). Association for Computing Machinery, New York, NY, USA, 1368–1380. https://doi.org/10.1145/3510003.3510052.
dc.identifier.uri	https://dr.lib.iastate.edu/handle/20.500.12876/5w5p00oz
dc.language.iso	en
dc.publisher	Association for Computing Machinery
dc.rights	Copyright © 2022 Owner/Author. This work is licensed under a Creative Commons Attribution International 4.0 License.
dc.source.uri	https://doi.org/10.1145/3510003.3510052	*
dc.subject.disciplines	DegreeDisciplines::Physical Sciences and Mathematics::Computer Sciences::Software Engineering
dc.subject.keywords	Deep Learning
dc.subject.keywords	AutoML
dc.subject.keywords	Mining Software Repositories
dc.subject.keywords	MSR
dc.title	Manas: mining software repositories to assist AutoML
dc.type	Presentation
dspace.entity.type	Publication
relation.isAuthorOfPublication	4e3f4631-9a99-4a4d-ab81-491621e94031
relation.isOrgUnitOfPublication	f7be4eb9-d1d0-4081-859b-b15cee251456

File

Original bundle

Now showing 1 - 1 of 1

Name:: 2022-Rajan-ManasMining.pdf
Size:: 985.63 KB
Format:: Adobe Portable Document Format
Description:

Download

Collections

Conference Proceedings and Presentations