As the Holi semester comes to a close, I wanted to take a moment to express my appreciation for the students in my seminar course on Responsible AI.
The backdrop of this course is something very familiar to all of us. We are living at a time, where more than ever before, we are continuously grappling with questions on how to engage with AI systems and what kind of role we envision AI technologies playing in our collective futures. It is in no way an exaggeration to say that questions on what constitutes Responsible AI development and deployment are at the center of this conversation.
The primary purpose of this course was to provide an overview of the technical apparatus involved in conducting research on Responsible AI. But even with the technical focus, we pursued a multidisciplinary inquiry on this topic. Topics covered included influential works on algorithmic fairness, actionability, causation, robustness, interpretability, explainability, and alignment. These were complemented and contextualized by scholarship from economics, philosophy, political science, law, sociology, and cognitive science, with the aim of having holistic discussions on what constitutes Responsible AI and how we might get there.
This writeup intends to highlight the excellent work students produced for their course projects. The aim is both to showcase what they accomplished and to invite feedback, suggestions, and dialogue from the broader academic community. With this context, below are the major themes that emerged across these projects.
Several built tools, websites, and demos designed to educate people about the impacts of AI and help them reclaim agency over aspects of these systems that are often only quasi-consensual in nature
- Sanket Rout and Vinayak Vishvakarma’s project explored the meaning of consent when it comes to AI trained on internet data. Their web-based resource (available here) serves as a guide for artists and creators on the current state of AI regulation, and provides an overview of tools available to protect their work from AI-based replication.
- Arnav Raj built a simple-and-effective plugin (available here) to protect people’s private identifying information from being shared during their interactions with chatbots, while also ensuring that user experience is not degraded.
- Arindam Dongre and Pratik Chaudhri created a website on the environmental costs associated of chatbot interactions, drawing attention to the growing climate impact of modern AI tools.
Algorithmic Fairness and Social Choice
Others pursued a deeper engagement with the complexities of assessing and achieving algorithmic fairness.
- Girish Thakur and Davy Dhadut explored the consequences of shifting underlying data distributions on the gender bias of algorithmic classifiers, using a fascinating dataset of Indian judicial cases.
- Tharun Tej’s project informed on the impossibility results in algorithmic fairness literature, focusing on the fundamental infeasibility of satisfying multiple fairness criteria simultaneously.
- Sanchita Saha and Siddhesh Nagar undertook an investigation into what actually drives hate speech classification; whether automated classifiers genuinely account for the complex contexts of such speech, or whether they rely primarily on stereotypical triggers.
- Shashank Shekhar and Arge Saurabh’s project explored the social choice aspects associated with achevieing algorithmic fairness when developing ensemble models.
Beyond Predictive Models: Recourse and Causation
Some went beyond predictive models, and investigated actions motivated by algorithmic predictions
- Joel M and Raj Kumar’s project evaluated the efficacy of actions suggested by algorithmic recourse methods, finding hilarious and concerning cases of unrealistic recourse suggestions.
- Lakshya Batra and Advait Karnatak similarly studied the importance of assessing recourse variability over time in credit related datasets from Korea and India.
- Khushi Mishra and Karan Singh explored the underappreciated role of confounders in ML, highlighting the dangers of ignoring them when creating AI models for prediction and recourse.
Large Language Models: Alignment and Evaluation
Finally, many played with current LLM methods, assessing the methods we undertake to align LLM models to human values or to evaluate domain-specific LLM behavior.
- Hasan Mustafa and Somesh Agarwal simulated a society of AI agents, with active interactions between utility-maximizing agents and LLMs, investigating the behavioral and economic dynamics that emerge in this digital environment.
- Ajeet Kumar Singh and Atharv Tripathi study the reliability of critique-based improvement of LLMs, empirically exploring the possibility of confirmation biases.
- Abhay Sharma spoke about the alignment tax, and the difficulty of fine-tuning models to not produce behavior that is considered unacceptable without compromising overall performance.
- Dhruv Pawar and Anurag Tiwari studied the efficacy of self-correction methods in developing safe LLM models, moving towards the broader goal of figuring out when is self-correction an effective intervention for AI alignment.
- Pulkit Sheoran and Tanisha Bhaskar assessed hallucination rates of LLMs across different domains, demonstrating that chances of LLM hallucination are higher for interactions related to current affairs.
All of these projects deepen our understanding of AI and our interactions with it. But, from the perspective of the objectives of our course, I’d like to believe that they go further: these projects not only improve our understanding of the societal impacts of AI, but also resulted in impressive tools and research that can inform others.
Please do reach out to me or to the students directly if anything here catches your interest :)