<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
    <title>anirban · basu - data-cleaning</title>
    <subtitle>Researcher in Computational Trust, Privacy, Security and Artificial Intelligence</subtitle>
    <link rel="self" type="application/atom+xml" href="https://anirbanbasu.netlify.app/tags/data-cleaning/atom.xml"/>
    <link rel="alternate" type="text/html" href="https://anirbanbasu.netlify.app"/>
    <generator uri="https://www.getzola.org/">Zola</generator>
    <updated>2026-04-24T00:00:00+00:00</updated>
    <id>https://anirbanbasu.netlify.app/tags/data-cleaning/atom.xml</id>
    <entry xml:lang="en">
        <title>Practical confidential data cleaning using trusted execution environments</title>
        <published>2025-08-04T00:00:00+00:00</published>
        <updated>2026-04-24T00:00:00+00:00</updated>
        
        <author>
          <name>Anirban Basu</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://anirbanbasu.netlify.app/publications/basu2025practical/"/>
        <id>https://anirbanbasu.netlify.app/publications/basu2025practical/</id>
        
        <content type="html" xml:base="https://anirbanbasu.netlify.app/publications/basu2025practical/">&lt;!-- citation: basu2025practical --&gt;
&lt;h3&gt;&lt;i&gt;&lt;u&gt;Anirban Basu&lt;/u&gt;, Masayuki Yoshino and Minako Toba&lt;/i&gt;&lt;/h3&gt;
&lt;div&gt;&lt;b&gt;Abstract:&lt;/b&gt; Data cleaning, also known as data cleansing or data scrubbing, is the process of identifying and correcting errors, inconsistencies, and inaccuracies in datasets. It is a crucial step in statistical analysis and machine learning as the quality of the input data directly affects the reliability and validity of the results obtained from any analysis. Data cleaning, when outsourced, poses privacy and confidentiality challenges. To address these, there has been recent research focus on privacy and confidentiality preserving data cleaning. In this paper, we propose a practical qualitative data cleaning system that preserves the privacy and confidentiality of the data utilising trusted execution environments. We have implemented our system in Python and deployed it using the Gramine library operating system on Intel Software Guard Extensions (SGX) hardware.&lt;/div&gt;
&lt;table class=&quot;table-publication-metadata&quot;&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;th scope=&quot;col&quot;&gt;address&lt;/th&gt;
&lt;td&gt;Chania, Greece&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;th scope=&quot;col&quot;&gt;booktitle&lt;/th&gt;
&lt;td&gt;Proceedings of the IEEE International Conference on Cyber Security and Resilience (CSR)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;th scope=&quot;col&quot;&gt;day&lt;/th&gt;
&lt;td&gt;04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;th scope=&quot;col&quot;&gt;doi&lt;/th&gt;
&lt;td&gt;&lt;a href=&quot;https://doi.org/10.1109/CSR64739.2025.11130151&quot; target=&quot;_blank&quot;&gt;10.1109/CSR64739.2025.11130151&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;th scope=&quot;col&quot;&gt;keywords&lt;/th&gt;
&lt;td&gt;data-privacy, statistical-analysis, operating-systems, machine-learning, cleaning, software, security, reliability, resilience, python, trusted-execution-environments, data-cleaning, security, privacy, confidentiality&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;th scope=&quot;col&quot;&gt;month&lt;/th&gt;
&lt;td&gt;08&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;th scope=&quot;col&quot;&gt;pages&lt;/th&gt;
&lt;td&gt;342-349&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;details&gt;
&lt;summary&gt;Cite this publication, using BibTeX&lt;/summary&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color: #F8F8F2; background-color: #272822;&quot; &gt;&lt;code data-lang=&quot;bibtex&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #F92672;&quot;&gt;@inproceedings&lt;/span&gt;&lt;span&gt;{&lt;/span&gt;&lt;span style=&quot;color: #A6E22E;text-decoration: underline;&quot;&gt;basu2025practical&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #66D9EF;&quot;&gt;  abstract&lt;/span&gt;&lt;span&gt; = {Data cleaning, also known as data cleansing or data scrubbing, is the process of identifying and correcting errors, inconsistencies, and inaccuracies in datasets. It is a crucial step in statistical analysis and machine learning as the quality of the input data directly affects the reliability and validity of the results obtained from any analysis. Data cleaning, when outsourced, poses privacy and confidentiality challenges. To address these, there has been recent research focus on privacy and confidentiality preserving data cleaning. In this paper, we propose a practical qualitative data cleaning system that preserves the privacy and confidentiality of the data utilising trusted execution environments. We have implemented our system in Python and deployed it using the Gramine library operating system on Intel Software Guard Extensions (SGX) hardware.},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #66D9EF;&quot;&gt;  address&lt;/span&gt;&lt;span&gt; = {Chania, Greece},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #66D9EF;&quot;&gt;  author&lt;/span&gt;&lt;span&gt; = {Basu, Anirban and Yoshino, Masayuki and Toba, Minako},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #66D9EF;&quot;&gt;  booktitle&lt;/span&gt;&lt;span&gt; = {Proceedings of the IEEE International Conference on Cyber Security and Resilience (CSR)},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #66D9EF;&quot;&gt;  day&lt;/span&gt;&lt;span&gt; = {4},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #66D9EF;&quot;&gt;  doi&lt;/span&gt;&lt;span&gt; = {10.1109/CSR64739.2025.11130151},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #66D9EF;&quot;&gt;  keywords&lt;/span&gt;&lt;span&gt; = {Data privacy;Statistical analysis;Operating systems;Machine learning;Cleaning;Software;Security;Reliability;Resilience;Python;trusted execution environments;data cleaning;security;privacy;confidentiality},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #66D9EF;&quot;&gt;  month&lt;/span&gt;&lt;span&gt; = {August},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #66D9EF;&quot;&gt;  pages&lt;/span&gt;&lt;span&gt; = {342-349},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #66D9EF;&quot;&gt;  title&lt;/span&gt;&lt;span&gt; = {Practical confidential data cleaning using trusted execution environments},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: #66D9EF;&quot;&gt;  year&lt;/span&gt;&lt;span&gt; = {2025},&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;}&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/details&gt;
</content>
        
    </entry>
</feed>
