RLHF (PPO) for less toxic summarization
Implemented a RLHF (PPO) pipeline for LoRA fine-tuning an LLM to generate less toxic dialogue summarization.
Implemented a RLHF (PPO) pipeline for LoRA fine-tuning an LLM to generate less toxic dialogue summarization.