Uncertainty Estimation of Deep Neural Networks
- 44 views
Monday, October 15, 2018 - 02:30 pm
Meeting room 2267, Innovation Center
DISSERTATION DEFENSE
Department of Computer Science and Engineering
University of South Carolina
Author : Chao Chen
Advisor : Dr. Gabriel Terejanu
Date : Oct. 15th , 2018
Time : 2:30 pm
Place : Meeting room 2267, Innovation Center
Abstract
Normal neural networks trained with gradient descent and back-propagation have received great success in various applications. On one hand, point estimation of the network weights is prone to over-fitting problems and lacks important uncertainty information associated with the estimation. On the other hand, exact Bayesian neural network methods are intractable and non-applicable for real-world applications. To date, approximate methods have been actively under development for Bayesian neural networks, including but not limited to, stochastic variational methods, Monte Carlo dropouts, and expectation propagation. Though these methods are applicable for current large networks, there are limits of these approaches with either under estimation or over-estimation of uncertainty. Extended Kalman filters (EKFs) and unscented Kalman filters (UKFs), which are widely used in data assimilation community, adopt a different perspective of inferring the parameters. Nevertheless, EKFs are incapable of dealing with highly non-linearity, while UKFs are inapplicable for large network architectures.
Ensemble Kalman filters (EnKFs) serve as great methodology in atmosphere and oceanology disciplines targeting extremely high-dimensional, non-Gaussian, and nonlinear state-space models. So far, there is little work that applies EnKFs to estimate the parameters of deep neural networks. By considering neural network as a nonlinear function, we augment the network prediction with parameters as new states and adapt the state-space model to update the parameters. In the first work, we describe the ensemble Kalman filter, two proposed algorithms for training both fully-connected and Long Short-term Memory (LSTM) networks, and experiment it with a synthetic dataset, 10 UCI datasets, and a natural language dataset for different regression tasks. To further evaluate the effectiveness of the proposed training scheme, we trained a deep LSTM network with the proposed algorithm, and applied it on five real-world sub-event detection tasks. With a formalization of the sub-event detection task, we develop an outlier detection framework and take advantage of the Bayesian Long Short-term Memory (LSTM) network to capture the important and interesting moments within an event. In the last work, we develop a framework for student knowledge estimation using Bayesian network. By constructing student models with Bayesian network, we can infer the new state of knowledge on each concept given a student. With a novel parameter estimate algorithm, the model can also indicate misconception on each question. Furthermore, we develop a predictive validation metric with expected data likelihood of the student model to evaluate the design of questions.

Abstract: Cybersecurity is becoming one of the challenging problems in the connected world because of heterogeneity of networked systems and scale and complexity of cyberspace. Cyber- attacks are not only increasing in terms of numbers but also getting more sophisticated. Cyber- defense for prevention, detection and response to cyber-attacks is an on-going challenge that needs efforts to protect critical infrastructures and private information. Complexity and scale of cyberspace and heterogeneity of networked systems make cybersecurity even more challenging. Almost all organizations are vulnerable to (similar or same) cyber-attacks where information sharing could help prevent future cyber-attacks
This talk presents and evaluates an information sharing framework for cybersecurity with the goal of protecting confidential information and networked infrastructures from future cyber- attacks. The proposed framework leverages the blockchain concept where multiple organizations/agencies participate for information sharing (without violating their privacy) to secure and monitor their cyberspaces. This blockchain based framework is to constantly collect high resolution cyber-attack information across organizational boundaries of which the organizations have no specific knowledge or control over any other organizations' data or damage caused by cyber-attacks.
Bio:
Laurent L. Njilla received his B.S. in Computer Science from the University of Yaoundé 1 in Cameroon, the M.S. in Computer Engineering from the University of Central Florida (UCF) in 2005 and Ph.D. in Electrical Engineering from Florida International University (FIU) in 2015. He joined the Cyber Assurance Branch of the U.S. Air Force Research Laboratory (AFRL), Rome, New York, as a Research Electronics Engineer in 2015. Prior to joining the AFRL, he was a Senior Systems Analyst in the industry sector for more than 10 years. He is responsible for conducting basic research in the areas of hardware design, game theory applied to cyber security and cyber survivability, hardware Security, online social network, cyber threat information sharing, category theory, and blockchain technology. He is the Program Manager for the Cyber Security Center of Excellence (CoE) for the HBCU/MI and the Disruptive Information Technology Program at AFRL/RI. Dr. Njilla’s research has resulted in more than 50 peer-reviewed journal and conference papers and multiple awards including Air Force Notable Achievement Awards, the 2015 FIU World Ahead Graduate award and etc. He is a reviewer of multiple journals and serves on the technical program committees of several international conferences. He is a member of the National Society of Black Engineer (NSBE).
Please see
Prof. Amit P. Sheth
Abstract: While Bill Gates, Stephen Hawking, Elon Musk, Peter Thiel, and others engage in OpenAI discussions of whether or not AI, robots, and machines will replace humans, proponents of human-centric computing continue to extend work in which humans and machines partner in contextualized and personalized processing of multimodal data to derive actionable information.
This talk describes how maturing towards the emerging paradigms of semantic computing (SC), cognitive computing (CC), and perceptual computing (PC) provides a continuum through which to exploit the ever-increasing and growing diversity of data that could enhance people’s daily lives. SC and CC sift through raw data to personalize it according to context and individual users, creating abstractions that move the data closer to what humans can readily understand and apply in decision-making. PC, which interacts with the surrounding environment to collect data that is relevant and useful in understanding the outside world, is characterized by interpretative and exploratory activities that are supported by the use of prior/background knowledge. Using the examples of personalized digital health and a smart city, we will demonstrate how the trio of these computing paradigms form complementary capabilities that will enable the development of the next generation of intelligent systems. For background: