Enabling Privacy-Preserving Inference in the Electricity Sector Using Large Language Models

Author Names:
Shenglong Liu, Yixin Li, Ge Zhang, Yuxiao Xia, and Zhenqi Guo
Author Affiliation:
Author Email:
yixin__li@163.com
Publication Date:
February 26, 2026

Page numbers:

DOI Number:

http://-

Abstract:

Large language models (LLMs) such as ChatGPT and GPT-4 have achieved remarkable success across tasks like image synthesis, text generation, speech interaction, and multimodal understanding. LLMs trained on massive real-world data pose serious privacy risks. Existing studies mainly address training-time protection (e.g., differential privacy), whereas inference-time vulnerabilities remain underexplored, especially in sensitive domains where user prompts may contain private information. In this work, we study privacy leakage during LLM inference in the electric power industry, where inputs often include sensitive fields such as ID numbers and electricity records. We show that LLMs can inadvertently memorize or regenerate such information, even without explicit prompts. Existing defenses—e.g., regex filters or alignment-based safety layers—are brittle under fine-tuning and insufficient for handling domain-specific patterns. To address this, we propose a dual-channel privacy protection framework that enforces semantic-level input-output sanitization using a shared detector and configurable policies (e.g., redaction, rejection, rewriting). Experiments on two realistic scenarios demonstrate that our approach effectively blocks leakage of sensitive fields while maintaining task performance, offering a practical solution for inference-time privacy in LLM deployments. https://mc.manuscriptcentral.com/jcmse Journal of Computational Methods in Science and Engineering For Peer Review Enabling Privacy-Preserving Inference in the Electricity Sector Using Large Language Models Shenglong Liu1, Yixin Li1*, Ge Zhang1, Yuxiao Xia1, and Zhenqi Guo2 1. Big Data Center of State Grid Corporation of China, Beijing, China. 2.School of Computer Science and Technology, Tongji University, Shanghai, China. * Correspondence to: yixin__li@163.com Abstract: Large language models (LLMs) such as ChatGPT and GPT-4 have achieved remarkable success across tasks like image synthesis, text generation, speech interaction, and multimodal understanding. LLMs trained on massive real-world data pose serious privacy risks. Existing studies mainly address training-time protection (e.g., differential privacy), whereas inference-time vulnerabilities remain underexplored, especially in sensitive domains where user prompts may contain private information. In this work, we study privacy leakage during LLM inference in the electric power industry, where inputs often include sensitive fields such as ID numbers and electricity records. We show that LLMs can inadvertently memorize or regenerate such information, even without explicit prompts. Existing defenses—e.g., regex filters or alignment-based safety layers—are brittle under finetuning and insufficient for handling domain-specific patterns. To address this, we propose a dual-channel privacy protection framework that enforces semantic-level input-output sanitization using a shared detector and configurable policies (e.g., redaction, rejection, rewriting). Experiments on two realistic scenarios demonstrate that our approach effectively blocks leakage of sensitive fields while maintaining task performance, offering a practical solution for inference-time privacy in LLM deployments.
Keywords:
generative artificial intelligence, sensitive data protection, electric power industry, large language models Abstract: Large language models (LLMs) such as ChatGPT and GPT-4 have achieved remarkable success across tasks like image synthesis, text generation, speech interaction, and multimodal understanding. LLMs trained on massive real-world data pose serious privacy risks. Existing studies mainly address training-time protection (e.g., differential privacy), whereas inference-time vulnerabilities remain underexplored, especially in sensitive domains where user prompts may contain private information. In this work, we study privacy leakage during LLM inference in the electric power industry, where inputs often include sensitive fields such as ID numbers and electricity records. We show that LLMs can inadvertently memorize or regenerate such information, even without explicit prompts. Existing defenses—e.g., regex filters or alignment-based safety layers—are brittle under fine-tuning and insufficient for handling domain-specific patterns. To address this, we propose a dual-channel privacy protection framework that enforces semantic-level input-output sanitization using a shared detector and configurable policies (e.g., redaction, rejection, rewriting). Experiments on two realistic scenarios demonstrate that our approach effectively blocks leakage of sensitive fields while maintaining task performance, offering a practical solution for inference-time privacy in LLM deployments. https://mc.manuscriptcentral.com/jcmse Journal of Computational Methods in Science and Engineering For Peer Review Enabling Privacy-Preserving Inference in the Electricity Sector Using Large Language Models Shenglong Liu1, Yixin Li1*, Ge Zhang1, Yuxiao Xia1, and Zhenqi Guo2 1. Big Data Center of State Grid Corporation of China, Beijing, China. 2.School of Computer Science and Technology, Tongji University, Shanghai, China. * Correspondence to: yixin__li@163.com Abstract: Large language models (LLMs) such as ChatGPT and GPT-4 have achieved remarkable success across tasks like image synthesis, text generation, speech interaction, and multimodal understanding. LLMs trained on massive real-world data pose serious privacy risks. Existing studies mainly address training-time protection (e.g., differential privacy), whereas inference-time vulnerabilities remain underexplored, especially in sensitive domains where user prompts may contain private information. In this work, we study privacy leakage during LLM inference in the electric power industry, where inputs often include sensitive fields such as ID numbers and electricity records. We show that LLMs can inadvertently memorize or regenerate such information, even without explicit prompts. Existing defenses—e.g., regex filters or alignment-based safety layers—are brittle under finetuning and insufficient for handling domain-specific patterns. To address this, we propose a dual-channel privacy protection framework that enforces semantic-level input-output sanitization using a shared detector and configurable policies (e.g., redaction, rejection, rewriting). Experiments on two realistic scenarios demonstrate that our approach effectively blocks leakage of sensitive fields while maintaining task performance, offering a practical solution for inference-time privacy in LLM deployments. Keywords: Generative artificial intelligence, sensitive data protection, electric power industry, large language models 1.
You need to register before accessing this content.
Scroll to Top