<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Security on Simple Made Daily</title>
    <link>https://caioferreira.dev/tags/security/</link>
    <description>Recent content in Security on Simple Made Daily</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <lastBuildDate>Tue, 07 May 2024 00:00:00 +0000</lastBuildDate><atom:link href="https://caioferreira.dev/tags/security/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Como estudar na área de Tecnologia?</title>
      <link>https://caioferreira.dev/posts/aprendendo-em-tech/</link>
      <pubDate>Tue, 07 May 2024 00:00:00 +0000</pubDate>
      
      <guid>https://caioferreira.dev/posts/aprendendo-em-tech/</guid>
      <description>Introdução Nos últimos dois meses, várias pessoas me procuraram para saber como organizo meus estudos. Decidi, então, compilar um resumo das minhas técnicas e métodos. Espero que encontrem aqui ferramentas úteis!
Método de estudos Primeiramente, o método de estudo é algo muito pessoal. As práticas que mencionarei tiveram êxito para mim. A primeira e mais importante tarefa é começar a desenvolver o seu método:
Com que frequência você deseja estudar? Como você vai organizar o conhecimento?</description>
      <content:encoded><![CDATA[<h2 id="introdução">Introdução</h2>
<p>Nos últimos dois meses, várias pessoas me procuraram para saber como organizo meus estudos. Decidi, então, compilar um resumo das minhas técnicas e métodos. Espero que encontrem aqui ferramentas úteis!</p>
<h2 id="método-de-estudos">Método de estudos</h2>
<p>Primeiramente, o método de estudo é algo muito pessoal. As práticas que mencionarei tiveram êxito para mim. A primeira e mais importante tarefa é começar a desenvolver o seu método:</p>
<ul>
<li>Com que frequência você deseja estudar?</li>
<li>Como você vai organizar o conhecimento?</li>
<li>Que tipo de conteúdo mais te engaja (livros, vídeo aulas, cursos, etc)?</li>
<li>Quanta estrutura você quer dar aos seus estudos? Você funciona melhor com um planejamento detalhado e prazos, ou prefere algo mais flexível, seguindo os tópicos que mais capturam sua atenção no momento?</li>
</ul>
<p>Posso falar por mim: estudo todos os dias, leio artigos ou livros sempre que possível. Organizo meu conhecimento em notas no <strong>Obsidian</strong>, usando uma estrutura que combina o <strong>Zettelkasten</strong> com o método <strong>PARA</strong>. Prefiro, de longe, conteúdos escritos e não gosto de fazer grandes planejamentos; estudo o que tenho vontade, quando tenho vontade.</p>
<h2 id="fontes-de-informação">Fontes de informação</h2>
<p>Explico que existem duas formas de desenvolvimento: horizontal e vertical.</p>
<ul>
<li>O desenvolvimento horizontal foca em criar uma base de conhecimento em um universo, abrangendo vários tópicos. Isso não implica em um estudo superficial, mas em entender como os tópicos se relacionam em um contexto. Dependendo do universo, como em segurança da informação, esse tipo de estudo pode ser um dos mais desafiadores.</li>
<li>O desenvolvimento vertical, por sua vez, foca em aprofundar um tópico e aprender o máximo possível sobre ele. O objetivo é estudar como se quisesse se tornar um especialista naquele assunto.</li>
</ul>
<p>Na carreira, é essencial manter ambos os tipos de desenvolvimento ativos. Para isso, recomendo realizar uma curadoria constante de diferentes tipos de fontes de informação:</p>
<ul>
<li><strong>Newsletters</strong> (horizontal): para mim, são a melhor ferramenta para um profissional de tecnologia hoje. Faço uma seleção contínua das newsletters que assino. Ao receber uma nova publicação, seleciono links relevantes, adiciono-os à minha lista de leitura e os consumo gradualmente.</li>
<li><strong>Feeds de agregadores</strong> (horizontal): sites como Hackersnews e Lobste.rs têm feeds RSS organizados por tema. Funcionam como uma linha do tempo de uma rede social, com muito ruído, mas ocasionalmente você encontra artigos úteis. É crucial saber gerenciar fontes de alto e baixo ruído para que não interfiram na eficiência do seu consumo.</li>
<li><strong>Blogs/Publicações especializadas</strong> (vertical): prefiro separar essas fontes dos livros tradicionais porque geralmente são mais sucintas, têm mais exemplos e exercícios práticos e são mais atualizadas que os livros. Exemplos incluem: <a href="https://quii.gitbook.io/learn-go-with-tests">https://quii.gitbook.io/learn-go-with-tests</a> e <a href="https://www.fuzzingbook.org/">https://www.fuzzingbook.org/</a>.</li>
<li><strong>Papers</strong> (vertical): começar a ler artigos acadêmicos pode ser desafiador, mas é um passo importante para conhecer a vanguarda do desenvolvimento. Quanto antes começar, melhor.</li>
<li><strong>Livros e cursos</strong> (vertical): geralmente os utilizo para tópicos que evoluem mais lentamente, como algoritmos, matemática e criptografia. Nessas áreas, você encontrará publicações excelentes e ainda relevantes.</li>
</ul>
<p>Com o tempo, ao consumir conteúdos, você acumulará links, artigos, livros, ferramentas: todo tipo de referência. Cuide bem dessa coleção, escolha uma forma de organização, um aplicativo, o que preferir. Essa coleção se tornará seu jardim, de onde extrairá insumos para produzir conhecimento.</p>
<p>No entanto, é fundamental <strong>não cair</strong> na <a href="https://zettelkasten.de/posts/collectors-fallacy/">falácia do colecionador</a>, onde simplesmente acumulamos pilhas de materiais e nos damos por satisfeitos.</p>
<h2 id="organizando-conhecimento">Organizando conhecimento</h2>
<p>Um dos meus maiores arrependimentos foi não ter começado a manter notas organizadas mais cedo! Minha maior recomendação aqui é a simplicidade. É muito fácil cair na armadilha de otimizar a estrutura das suas notas e ser extremamente produtivo (eu já caí várias vezes), criando um sistema demasiadamente complicado.</p>
<p>Algumas coisas que faço para estruturar o conhecimento são:</p>
<ul>
<li>Notas de Literatura: fazer notas detalhadas e com suas próprias palavras custa muito tempo, então eu criei um tipo diferente de nota chamada Nota de Literatura. Nela, deposito informações relevantes da fonte de informação, por vezes copiando diretamente, outras vezes reescrevendo.
<ul>
<li>Assistente AI: embora o objetivo seja produzir notas mais simples e rápidas, a AI tem me ajudado a produzir notas de literatura com qualidade de notas de assunto. Eu uso um prompt que gera uma lista de perguntas sobre os tópicos abordados na fonte de informação. Isso tem tornado minhas notas de literatura muito mais eficientes. O prompt que uso está disponível publicamente: <a href="https://github.com/caiorcferreira/prompts/blob/main/patterns/thought_provoking.md">thought_provoking prompt</a>.</li>
</ul>
</li>
<li>Notas de Assunto (Subject/Evergreen notes): essas notas são o personagem principal. Elas têm uma vida longa, são atualizadas sempre que aprendo algo novo sobre o assunto e são escritas com minhas próprias palavras. A técnica mais importante para mim aqui é escrever como se estivesse explicando para outra pessoa.</li>
</ul>
<h2 id="aplicando-seu-método">Aplicando seu método</h2>
<p>Existem ainda desafios únicos durante o andamento dos seus estudos. Uma pergunta recorrente é: eu sei qual assunto quero estudar, sei qual tipo de estudo (horizontal ou vertical) e tenho uma curadoria de fontes - <strong>como escolher qual material consumir?</strong> Nesse caso, meu método é:</p>
<ol>
<li>Seleciono material relacionado ao tema que já possa ter armazenado;</li>
<li>Uso <a href="https://pt.linkedin.com/pulse/entendendo-o-google-hacking-na-pratica-e-otimizando-suas-janones">DORKs</a> para pesquisar material no histórico de newsletters e feeds de agregadores.</li>
<li>Dedico um tempo navegando entre os diferentes materiais selecionados, até encontrar um (ou alguns) com um bom equilíbrio entre didática e cobertura do assunto.</li>
</ol>
<p>Isso acaba levando a um começo mais lento, porém, na minha experiência, ele se paga, pois me dá a oportunidade de selecionar materiais de estudo de maior qualidade.</p>
<p>Outro desafio comum é: <strong>como tirar dúvidas sobre o assunto?</strong> Atualmente, LLMs têm sido de grande ajuda para mim (desde que você saiba como lidar com possíveis alucinações). Buscar por comunidades dedicadas ao assunto no Reddit também costuma ter bons resultados.
No entanto, a melhor ferramenta é sua própria coleção de materiais: busque por outros materiais que possam explicar o mesmo conceito a partir de diferentes pontos de vista. Eventualmente, você irá montar um quebra-cabeça, com uma compreensão muito mais abrangente do assunto.</p>
<h2 id="conclusão">Conclusão</h2>
<p>A jornada do aprendizado é profundamente pessoal e varia de indivíduo para indivíduo. As estratégias e métodos que compartilhei refletem minhas experiências e preferências pessoais. É importante encontrar o equilíbrio entre a absorção de informação e a criação de conhecimento, evitando a armadilha de apenas acumular referências.</p>
<p>Lembre-se de que a chave para um estudo eficaz não está apenas em selecionar os materiais certos, mas também em como você organiza e interage com essas informações ao longo do tempo. Espero que este guia sirva como um ponto de partida para você desenvolver um método de estudo que seja tão único e eficaz quanto suas próprias aspirações e necessidades de aprendizado.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>How to learn in the Tech field?</title>
      <link>https://caioferreira.dev/posts/learning-in-tech/</link>
      <pubDate>Tue, 07 May 2024 00:00:00 +0000</pubDate>
      
      <guid>https://caioferreira.dev/posts/learning-in-tech/</guid>
      <description>Introduction Over the past two months, several people have approached me to find out how I organize my studies. Therefore, I decided to compile a summary of my techniques and methods. I hope you find useful tools here!
Study Method Firstly, the study method is something very personal. The practices I will mention have been successful for me. The first and most important task is to start developing your method:</description>
      <content:encoded><![CDATA[<h2 id="introduction">Introduction</h2>
<p>Over the past two months, several people have approached me to find out how I organize my studies. Therefore, I decided to compile a summary of my techniques and methods. I hope you find useful tools here!</p>
<h2 id="study-method">Study Method</h2>
<p>Firstly, the study method is something very personal. The practices I will mention have been successful for me. The first and most important task is to start developing your method:</p>
<ul>
<li>How often do you want to study?</li>
<li>How will you organize knowledge?</li>
<li>What type of content engages you the most (books, video lessons, courses, etc)?</li>
<li>How much structure do you want to give your studies? Do you function better with detailed planning and deadlines, or do you prefer something more flexible, following the topics that capture your attention at the moment?</li>
</ul>
<p>I can speak for myself: I study every day, read articles or books whenever possible. I organize my knowledge in notes in <strong>Obsidian</strong>, using a structure that combines the <strong>Zettelkasten</strong> method with the <strong>PARA</strong> method. I vastly prefer written content and do not like making extensive plans; I study what I want, when I want.</p>
<h2 id="sources-of-information">Sources of Information</h2>
<p>I explain that there are two forms of development: horizontal and vertical.</p>
<ul>
<li>Horizontal development focuses on creating a knowledge base in a universe, covering various topics. This does not imply superficial study but understanding how the topics relate in a context. Depending on the universe, such as in information security, this type of study can be one of the most challenging.</li>
<li>Vertical development, on the other hand, focuses on deepening a topic and learning as much as possible about it. The goal is to study as if you wanted to become an expert on that subject.</li>
</ul>
<p>In your career, it is essential to keep both types of development active. For this, I recommend constantly curating different types of information sources:</p>
<ul>
<li><strong>Newsletters</strong> (horizontal): for me, they are the best tool for a technology professional today. I make a continuous selection of the newsletters I subscribe to. When I receive a new publication, I select relevant links, add them to my reading list, and consume them gradually.</li>
<li><strong>Feed aggregators</strong> (horizontal): sites like Hackersnews and Lobste.rs have RSS feeds organized by theme. They work like a social network timeline, with a lot of noise, but occasionally you find useful articles. It is crucial to know how to manage high and low noise sources so that they do not interfere with the efficiency of your consumption.</li>
<li><strong>Blogs/Specialized publications</strong> (vertical): I prefer to separate these sources from traditional books because they are usually more concise, have more examples and practical exercises, and are more up-to-date than books. Examples include: <a href="https://quii.gitbook.io/learn-go-with-tests">https://quii.gitbook.io/learn-go-with-tests</a> and <a href="https://www.fuzzingbook.org/">https://www.fuzzingbook.org/</a>.</li>
<li><strong>Papers</strong> (vertical): starting to read academic papers can be challenging, but it is an important step to know the cutting edge of development. The sooner you start, the better.</li>
<li><strong>Books and courses</strong> (vertical): I generally use them for topics that evolve more slowly, such as algorithms, mathematics, and cryptography. In these areas, you will find excellent publications that are still relevant.</li>
</ul>
<p>Over time, as you consume content, you will accumulate links, articles, books, tools: all kinds of references. Take good care of this collection, choose a way to organize it, an app, whatever you prefer. This collection will become your garden, from which you will draw inputs to produce knowledge.</p>
<p>However, it is crucial <strong>not to fall</strong> into the <a href="https://zettelkasten.de/posts/collectors-fallacy/">collector&rsquo;s fallacy</a>, where we simply accumulate piles of materials and are satisfied with that.</p>
<h2 id="organizing-knowledge">Organizing Knowledge</h2>
<p>One of my biggest regrets was not starting to keep organized notes earlier! My biggest recommendation here is simplicity. It&rsquo;s very easy to fall into the trap of optimizing your note structure and being extremely productive (I&rsquo;ve fallen several times), creating an overly complicated system.</p>
<p>Some things I do to structure knowledge are:</p>
<ul>
<li>Literature Notes: making detailed notes in your own words takes a lot of time, so I created a different type of note called Literature Note. In it, I deposit relevant information from the source of information, sometimes copying directly, other times rewriting.
<ul>
<li>AI Assistant: although the goal is to produce simpler and quicker notes, AI has helped me produce literature notes with the quality of subject notes. I use a prompt that generates a list of questions about the topics covered in the source of information. This has made my literature notes much more efficient. The prompt I use is publicly available: <a href="https://github.com/caiorcferreira/prompts/blob/main/patterns/thought_provoking.md">thought_provoking prompt</a>.</li>
</ul>
</li>
<li>Subject Notes (Subject/Evergreen notes): these notes are the main character. They have a long life, are updated whenever I learn something new about the subject, and are written in my own words. The most important technique for me here is writing as if I were explaining to someone else.</li>
</ul>
<h2 id="applying-your-method">Applying Your Method</h2>
<p>There are still unique challenges during the course of your studies. A recurring question is: I know which subject I want to study, I know what type of study (horizontal or vertical) and I have a curation of sources - <strong>how to choose which material to consume?</strong> In this case, my method is:</p>
<ol>
<li>I select material related to the theme that I may have already stored;</li>
<li>I use <a href="https://pt.linkedin.com/pulse/understanding-google-hacking-in-practice-and-optimizing-your-janones">DORKs</a> to search for material in the newsletters&rsquo;s archives and feed aggregators.</li>
<li>Then, I dedicate some time navigating among the selected materials, until I find one (or a few) with a good balance between didactic approach and subject coverage.</li>
</ol>
<p>This ends up leading to a slower start, but in my experience, it pays off, as it gives me the opportunity to select higher quality study materials.</p>
<p>Another common challenge is: <strong>how to ask questions about the subject?</strong> Currently, LLMs have been of great help to me (as long as you know how to deal with possible hallucinations). Searching for dedicated communities on Reddit often yields good results.
However, the best tool is your own collection of materials: look for other materials that may explain the same concept from different perspectives. Eventually, you will put together a puzzle, gaining a much more comprehensive understanding of the subject.</p>
<h2 id="conclusion">Conclusion</h2>
<p>The learning journey is deeply personal and varies from individual to individual. The strategies and methods I have shared reflect my experiences and personal preferences. It&rsquo;s important to find a balance between absorbing information and creating knowledge, avoiding the trap of just accumulating references.</p>
<p>Remember, the key to effective study is not only in selecting the right materials but also in how you organize and interact with these materials over time. I hope this guide serves as a starting point for you to develop a study method that is as unique and effective as your own learning aspirations and needs.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Notes on Dos and Don&#39;ts of Machine Learning in Computer Security</title>
      <link>https://caioferreira.dev/posts/notes-on-do-and-donts-of-ml-in-security/</link>
      <pubDate>Thu, 22 Jun 2023 00:00:00 +0000</pubDate>
      
      <guid>https://caioferreira.dev/posts/notes-on-do-and-donts-of-ml-in-security/</guid>
      <description>Following the subject from my last post, Reflections about Supervised Learning on Security, I put down some more thoughts about the implementation of learning-based systems in the Security domain.
This is my extension to the problems and recommendations presented on the paper Dos and Don&amp;rsquo;ts of Machine Learning in Computer Security (Quiring, et al, 2022). I encourage you to also read the paper, as it&amp;rsquo;s excellent and provide a lot of insights about how to better build machine learning models.</description>
      <content:encoded><![CDATA[<p>Following the subject from my last post, <a href="https://caioferreira.dev/posts/reflections-supervised-ml/reflections-supervised-learning-in-security/">Reflections about Supervised Learning on Security</a>, I put down some more thoughts about the implementation of learning-based systems in the Security domain.</p>
<p>This is my extension to the problems and recommendations presented on the paper <a href="https://mlsec.org/docs/2022-sec.pdf">Dos and Don&rsquo;ts of Machine Learning in Computer Security</a> (Quiring, et al, 2022). I encourage you to also read the paper, as it&rsquo;s excellent and provide a lot of insights about how to better build machine learning models.</p>
<h2 id="machine-learning-workflow">Machine learning workflow</h2>
<h3 id="data-collection">Data collection</h3>
<blockquote>
<p>Pitfalls: Sampling Bias, Label Inaccuracy</p>
</blockquote>
<p>As with any other domain, Security is also heavily affected by bad data quality used for training. However, it&rsquo;s often worse in Security because data acquisition of adversary activity is usually hard and the method&rsquo;s used to acquire it will bring a bias.</p>
<p>Let&rsquo;s take for example a honeypot implemented with a vulnerable Apache Server. Even though there are lots of bad actors in the wild, if you have a medium size environment, the volume of data they will produce attacking your honeypot will not come near to the volume of packets in your production network. Also, the TTPs you are going to see will come from threat actors that are used to leverage Apache Server in their kill chain, possibly leaving other adversaries, that may be focusing on other types of vectors, producing different kill chains, out of the dataset.</p>
<p>On top of that, we are usually dealing with lots of examples. So, if we have label inaccuracies, dealing with them is a lot harder. We can&rsquo;t just apply common techniques like using the mode to fill in, because in the Security domain, the adversary is actively trying to mimic the distribution of the benign cases. Therefore, we have very little space for noise, as it would blur even more the distinction between the outcomes.</p>
<p>Besides the recommendations presented in the article, I would add two more:</p>
<ul>
<li>User open and disseminate datasets whenever possible. Dataset sharing is still uncommon on the Security community. We already have some network and host datasets, but there are many more assets nowadays (logs for cloud, kubernetes, CDN, CI/CD, etc), and the technologies on networks and hosts are constantly evolving, posing the need for these datasets to be always updated.</li>
<li>Use a model design that depends less on adversary data, such as we discussed in the previous article. This will reduce the dependency on the adversary behavior and increase the sources of data available to use.</li>
</ul>
<h3 id="model-design-and-implementation">Model design and implementation</h3>
<blockquote>
<p>Pitfalls: Data Snooping, Spurious Correlations, Biased Parameter Selection</p>
</blockquote>
<p>Developing machine learning models is no easy task. Feature engineering, hyperparameters optimization, data preparation and appropriate splitting for validation (to avoid snooping). These are just some of the challenges when taking on this endeavor.</p>
<p>Using tools like AutoML may help reduce the burden, however it won&rsquo;t take care of everything. In the end, you still need to understand your data characteristics and how it&rsquo;s related to your problem.</p>
<p>But, some of these relations and characteristics tend to repeat. Time relations, for example, are extremely common and important in a lot of Security problems. Therefore, I suggest that every Security team doing machine learning on its own to think about how they can extract such aspects into reusable components.</p>
<p>This has the added benefit to scale the impact of the Security members that are more focused on implementing machine learning. Our domain has many different areas and enabling other teams and specialists to more easily and correctly implement models, even as proofs of concept, expands the possibilities of what the enterprise can achieve.</p>
<p>In the end, the best strategy to best tackle all these challenges and be more prepared to handle pitfalls it to avoid jumping to complex and sexy algorithms from the start, such as neural networks.</p>
<p>Using explanation techniques, as the article suggests, can help you catch spurious correlation, but even better is to use simpler and more understandable methods. I have been seen great results with simple probabilistic methods, such as with Histogram Based Outlier Score. Using a simple algorithm, we can immediately see what is driving its decision and catch faster these types of pitfalls.</p>
<h3 id="performance-evaluation">Performance evaluation</h3>
<blockquote>
<p>Pitfalls: Inappropriate Baseline, Inappropriate Performance Measures, Base Rate Fallacy</p>
</blockquote>
<p>As I mentioned before, most times when building a learning-based system in Security, we end up with imbalanced and noisy datasets. This demands a special care when choosing performance metrics to best reflect our model. Building on the previous suggestion, having common components that use by default metrics more fit to most problems in Security, such as precision, recall, and MCC, would reduce the change of human error.</p>
<p>However, even after computing correct performance metrics, we still need good baselines to compare them against. The article suggest goods options such as using simple methods or automated machine learning. I would add that comparing it with vendor products can also be a good experiments to fully understand the impact of replacing the third party with the model.</p>
<h3 id="operation">Operation</h3>
<blockquote>
<p>Pitfalls: Lab-Only Evaluation, Inappropriate Threat Model</p>
</blockquote>
<p>Although the article identifies Lab-Only Evaluation as a common pitfall in Security model, I think this is one of the topic where in Security we have the most options to address.</p>
<p>First, we could use adversary emulation frameworks and popular attack toolkits to launch real world attacks against a test environment and study the model response. Another option would be to set up a honeypot and evaluate how well the model performed by comparing to a human analyses of the evidences afterward.</p>
<p>This type of experiment would also feed back into our threat model. Given the hype of AI in latest months, more knowledge is beings shared and produced about security best practices and process for machine learning, however we can still consider the topic in its infancy.</p>
<p>One special vector that we are still starting to discuss is supply-chain attacks to models, specially with the popularization of transfer-learning for more complex algorithms such as LLMs.</p>
<p><a href="https://www.splunk.com/en_us/blog/security/paws-in-the-pickle-jar-risk-vulnerability-in-the-model-sharing-ecosystem.html">Splunk showed</a> that more than 80% of HuggingFace&rsquo;s models use pickle-serialized code, which is vulnerable to arbitrary code execution and code injection, although it&rsquo;s not possible to say if any of them is malicious.</p>
<p>But, besides low-level vulnerabilities like this, a transfer-learning based model can also inherit the biases (intentionally placed or not) from the original model. That is, if the original model is vulnerable to an adversarial example, there is a high risk that your new model is going to also be.</p>
<p>This opens the possible for two types of attacks:</p>
<ol>
<li>Malicious models: base models shared with adversarial examples trained, such as a large batch of images with a negative label having a red square, making the model learn that any image with a red square should be classified as negative, therefore implanting a bypass. This type of attack would be extremely difficult to detect.</li>
<li>Cross model generalization: as <a href="https://arxiv.org/pdf/1312.6199.pdf">Szegedy, et all points in their paper</a>, different models may be susceptible to the same adversarial examples, even when having different hyperparameters. In this case, the models would not even be so different, therefore an adversary could search for adversarial examples against the original model and just use them in the target system, with a high chance of success.</li>
</ol>
<h2 id="conclusion">Conclusion</h2>
<p>In conclusion, implementing learning-based systems in the security domain presents challenges that require careful consideration. Addressing issues related to data collection, model design and implementation, performance evaluation, and operational considerations is crucial. By implementing these recommendations and reflecting on these problems, organizations can enhance their capabilities in security.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Reflections about Supervised Learning on Security</title>
      <link>https://caioferreira.dev/posts/reflections-supervised-ml/reflections-supervised-learning-in-security/</link>
      <pubDate>Sun, 23 Apr 2023 00:00:00 +0000</pubDate>
      
      <guid>https://caioferreira.dev/posts/reflections-supervised-ml/reflections-supervised-learning-in-security/</guid>
      <description>Supervised learning is a technique that aims to learn a hypothesis function $h$ that fits a behavior observed in the real-world, which is governed by an unknown function $f$.
To learn this function, we use a set of example data points composed of inputs (also called features) and outcomes (sometimes called labels). These example data points were sampled from the real world behavior, i.e. from the function $f$, at some time in the past.</description>
      <content:encoded><![CDATA[<p>Supervised learning is a technique that aims to learn a hypothesis function $h$ that fits a behavior observed in the real-world, which is governed by an unknown function $f$.</p>
<p>To learn this function, we use a set of example data points composed of inputs (also called features) and outcomes (sometimes called labels). These example data points were sampled from the real world behavior, i.e. from the function $f$, at some time in the past. Our goal is that we can extract knowledge from the past to figure out this behavior on unseen data beforehand. This knowledge extracted from the set of examples is materialized on the $h$ function.</p>
<p>In Security, we can imagine some examples where this would be useful, like trying to learn if an HTTP request contains malicious payload or if some set of bytes is a malware or not. However, supervised learning is less often used in Security than in many other domains.</p>
<p>This happens because the principles of the supervised learning theory conflicts with the nature of Security, limiting its application. But, by understanding these principles, it is also possible to see how to best apply this technique and how it can maybe useful.</p>
<h2 id="stationary-assumption">Stationary assumption</h2>
<p>The most important principle in supervised learning is the stationary assumption. When the data that represents the real world behaviors follows the stationary assumption, it means that predicting the behavior using past example is approximately correctly.</p>
<p>The <strong>stationary assumption</strong> states that the behavior that is being learned don&rsquo;t change through time. This has some important consequences:</p>
<ol>
<li>We expect that each data point is independent of each other. This is important, because if there were causal effects between data points, then the features of a data point $x_1$, caused by $x_0$, would vary with a probability that is a combination of the probability distribution and the effect of $x_0$, hence the distribution would not remain the same over time, because $x_0$ and $x_1$ would vary in different ways. For example, in a box with 1 blue ball and 2 red balls, the probability of picking a blue ball when drawing the first one from the box is 33% while a red one would be 66%. However, if the first one is indeed blue, then the probability of drawing a red ball as the second one is 100%. The first data point (blue ball draw) changed the probability of the second data point.</li>
<li>We expect that each data point is identically distributed, i.e. each data point should be drawn from the same probability distribution. We could learn the shopping behavior using data ranging from Black Friday to New Years, however the users&rsquo; behavior in this time is completed different from the rest of the year, therefore the data used to learn has a different probability distribution than the unseen data on which we are going to make predictions.</li>
</ol>
<p>Any dataset that follows these two characteristics is said to hold the i.i.d assumption (independent and identically distributed). The importance for our training datasets on supervised learning to be i.i.d is because it connects the past to the future, without it, any inference made on the available data would be invalid.</p>
<p>Understand the i.i.d assumption is specially important for Security Machine Learning because it is one of the areas where causality and behavior shifts are most present. So, exactly how this assumption affects our ability to do Supervised Machine Learning?</p>
<h3 id="causality">Causality</h3>
<p>Many threat behaviors have causal nature, and therefore we should have a lot of care when preparing our datasets and choosing our validation methods.</p>
<p>A good example is malware classification, where you could have many samples from various families from different years. Each family generation influences each other, and sometimes they have similar characteristics.
A special bad situation that could happen with this is that during splitting of the dataset between training and testing, without taking into account the time relation of the families and samples, then you could end up training the model with future information and testing against past samples. This would produce falsely accurate results that would not generalize in the real-world.</p>
<p>There are ways to deal with this, but it depends on the type and strength of the causal relation between the data. For this case, ensuring that the newest malware samples are used for test should be enough.</p>
<h3 id="behavior-shift">Behavior Shift</h3>
<p>We could create a model to learn a threat behavior like the profile of a botnet, however once we started responding effectively to it, adversaries would adapt and our model would become useless because the new botnets would have a totally different behavior.
This would be the case of a change in the probability distribution from which the features are drawn, leading to the break of our assumption.</p>
<h2 id="looking-to-the-other-side">Looking to the other side</h2>
<p>Although causality may be addressed by good data preparation, preventing a model to be become outdated due to behavioral shift is almost impossible. However, this problem isn&rsquo;t a new one in Security, such that one best practices is to instead of trying to detect and block malicious action, defenders should define what a legitimate system behavior looks like and block everything else. This can be summarized as: allow lists are more secure than block lists.</p>
<p>We can apply the same philosophy for supervised learning, by modeling profiles of legitimate behavior, which usually are more stable. Then, new data points are classified against these multiple profiles models, finally a meta-classifier is used to choose which one is the best fit or if it&rsquo;s an outlier. This has the ability to catch any new threat behavior that deviates from the know legitimate profiles.</p>
<p>This combination of multiple models is called ensemble learning, which has shown to improve models performance, like when comparing a Decision Tree model to a Random Forest one.</p>
<p>As a side bonus, since this way of building models does not depend on knowing threat behaviors, it avoids the common problem in mining data for Security that usually produce highly unbalanced datasets. We can train each profile using the true positive data of others profiles as false examples, assuming each profile is mutually exclusive.</p>
<p>The main challenge in this approach is that adversaries may try to mimic legitimate profiles, however like the authors of Notos showed, this can be useful sometimes, as in their cases this would imply in adversaries using a more stable network infrastructure that would be easily defeated by static block lists. Therefore, forcing adversaries to mimic a legitimate profile could also reduce their capabilities.</p>
<h2 id="conclusion">Conclusion</h2>
<p>In conclusion, supervised learning is a powerful technique that allows us to extract knowledge from past data points to predict future behavior. However, in the Security domain, applying it comes with unique challenges due to the presence of causality and behavior shifts. Understanding the stationary assumption and the importance of having an i.i.d dataset is crucial, because it informs us how to best prepare our datasets, choose our validation methods and, most importantly, what are the best behaviors to be modeled using supervised learning.</p>
<p>Luckily, by modeling profiles of legitimate behavior, and catching new threat behavior that deviates from known legitimate profiles, we can build intelligent allow lists.</p>
<p>While there are challenges, with proper preparation and understanding, supervised learning can be a valuable tool in Security.</p>
<h2 id="references">References</h2>
<ul>
<li><a href="https://astrolavos.gatech.edu/articles/Antonakakis.pdf">Notos: Building a Dynamic Reputation System for DNS</a></li>
</ul>
]]></content:encoded>
    </item>
    
    <item>
      <title>Implementing a safe and sound API Key authorization middleware in Go</title>
      <link>https://caioferreira.dev/posts/golang-secure-api-key-middleware/</link>
      <pubDate>Sat, 05 Feb 2022 00:00:00 +0000</pubDate>
      
      <guid>https://caioferreira.dev/posts/golang-secure-api-key-middleware/</guid>
      <description>How to design a more secure API Key handling in Go</description>
      <content:encoded><![CDATA[<p><img loading="lazy" src="./cover.jpg" alt=""  />
</p>
<p>A common requirement that I face on multiple projects is to safeguard some API endpoints to administrative access, or to provide a secure way for other applications to consume our service in a controlled and traceable manner.</p>
<p>The usual solution for it is API Keys, a simple and effective authorization control mechanism that we can implement with a few lines of code. However, when doing, so we also need to be aware of threats and possible attacks that we may suffer, specially due to the usual privileges that these keys provides.</p>
<p>Therefore, we are going to analyze common points of concern and design a solution that improve our security posture while keeping it simple.</p>
<h2 id="api-keys-threats">API Keys threats</h2>
<p>There are two main concerns when implementing an API Key authorization scheme: <strong>key provisioning</strong> and <strong>timing attacks</strong>. Let&rsquo;s review each threat before designing solutions to address them.</p>
<h3 id="key-provisioning">Key Provisioning</h3>
<p>The key storage is directly related to how applications expect these secrets to be provided to them. Environment variables are the most common solution used on modern services since they are widely supported and don&rsquo;t incur a high reading cost (in contrast to files) allowing for dynamic changes to be easily detected.</p>
<p>However, the way developers usually define the environment variables are through scripts or configuration files, for example using a <a href="https://kubernetes.io/docs/concepts/configuration/secret/">Kubernetes Secret</a> manifest. This introduces a serious threat of API Keys being committed to git repositories, which in the event of data leakage from the internal VCS management system would expose these credentials.</p>
<p>Note: remember that once committed, even if the keys are deleted from the source files, the information is already on the repository history and is easily searchable with tools like <a href="https://github.com/trufflesecurity/truffleHog">TruffleHog</a>.</p>
<p>Therefore, <strong>please do not commit your API Keys to git</strong>!</p>
<h3 id="timing-attacks">Timing Attacks</h3>
<p>Once your application is configured with the available API Keys, you need to verify that the end-user provided key (let&rsquo;s call this the <em>user key</em>) is correct. Doing so with a naive algorithm, like using == operator, will make the verification end on the first incorrect character, hence reducing the time taken to respond.</p>
<p>A timing attack takes advantage of this scenario by trying to guess the correct characters of a secret based on how long the application took to respond. If the guess is right, the response will take slightly longer than if it&rsquo;s wrong.</p>
<p>Naturally, since equality checks are orders of magnitude faster than the network roundtrip, this type of attack is extremely difficult to perform because it depends on a statistical analysis of many response samples. By looking at the time distribution produced by two different characters, one can infer that if they are different, inferring that the greater one is the correct value. For an extensive discussion of statistical techniques to help perform this attack see <a href="https://www.blackhat.com/docs/us-15/materials/us-15-Morgan-Web-Timing-Attacks-Made-Practical-wp.pdf">Morgan, Morgan 2015</a>.</p>
<h2 id="middleware-design-and-implementation">Middleware design and implementation</h2>
<p>Having these threats in mind, we can design a suitable solution. Let&rsquo;s start with the most simple API Key middleware implementation possible and iterate from it.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-go" data-lang="go"><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="kd">func</span> <span class="nf">ApiKeyMiddleware</span><span class="p">(</span><span class="nx">cfg</span> <span class="nx">conf</span><span class="p">.</span><span class="nx">Config</span><span class="p">,</span> <span class="nx">logger</span> <span class="nx">logging</span><span class="p">.</span><span class="nx">Logger</span><span class="p">)</span> <span class="kd">func</span><span class="p">(</span><span class="nx">handler</span> <span class="nx">http</span><span class="p">.</span><span class="nx">Handler</span><span class="p">)</span> <span class="nx">http</span><span class="p">.</span><span class="nx">Handler</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">	<span class="nx">apiKeyHeader</span> <span class="o">:=</span> <span class="nx">cfg</span><span class="p">.</span><span class="nx">APIKeyHeader</span> <span class="c1">// string
</span></span></span><span class="line"><span class="cl"><span class="c1"></span>	<span class="nx">apiKeys</span> <span class="o">:=</span> <span class="nx">cfg</span><span class="p">.</span><span class="nx">APIKeys</span> <span class="c1">// map[string]string
</span></span></span><span class="line"><span class="cl"><span class="c1"></span>
</span></span><span class="line"><span class="cl">	<span class="nx">reverseKeyIndex</span> <span class="o">:=</span> <span class="nb">make</span><span class="p">(</span><span class="kd">map</span><span class="p">[</span><span class="kt">string</span><span class="p">]</span><span class="kt">string</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">	<span class="k">for</span> <span class="nx">name</span><span class="p">,</span> <span class="nx">key</span> <span class="o">:=</span> <span class="nx">apiKeys</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">		<span class="nx">reverseKeyIndex</span><span class="p">[</span><span class="nx">key</span><span class="p">]</span> <span class="p">=</span> <span class="nx">name</span>
</span></span><span class="line"><span class="cl">	<span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">	<span class="k">return</span> <span class="kd">func</span><span class="p">(</span><span class="nx">next</span> <span class="nx">http</span><span class="p">.</span><span class="nx">Handler</span><span class="p">)</span> <span class="nx">http</span><span class="p">.</span><span class="nx">Handler</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">		<span class="k">return</span> <span class="nx">http</span><span class="p">.</span><span class="nf">HandlerFunc</span><span class="p">(</span><span class="kd">func</span><span class="p">(</span><span class="nx">w</span> <span class="nx">http</span><span class="p">.</span><span class="nx">ResponseWriter</span><span class="p">,</span> <span class="nx">r</span> <span class="o">*</span><span class="nx">http</span><span class="p">.</span><span class="nx">Request</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">			<span class="nx">apiKey</span><span class="p">,</span> <span class="nx">err</span> <span class="o">:=</span> <span class="nf">bearerToken</span><span class="p">(</span><span class="nx">r</span><span class="p">,</span> <span class="nx">apiKeyHeader</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">			<span class="k">if</span> <span class="nx">err</span> <span class="o">!=</span> <span class="kc">nil</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">				<span class="nx">logger</span><span class="p">.</span><span class="nf">Errorw</span><span class="p">(</span><span class="s">&#34;request failed API key authentication&#34;</span><span class="p">,</span> <span class="s">&#34;error&#34;</span><span class="p">,</span> <span class="nx">err</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">				<span class="nf">RespondError</span><span class="p">(</span><span class="nx">w</span><span class="p">,</span> <span class="nx">http</span><span class="p">.</span><span class="nx">StatusUnauthorized</span><span class="p">,</span> <span class="s">&#34;invalid API key&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">				<span class="k">return</span>
</span></span><span class="line"><span class="cl">			<span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">			<span class="nx">_</span><span class="p">,</span> <span class="nx">found</span> <span class="o">:=</span> <span class="nx">reverseKeyIndex</span><span class="p">[</span><span class="nx">apiKey</span><span class="p">]</span>
</span></span><span class="line"><span class="cl">			<span class="k">if</span> <span class="p">!</span><span class="nx">found</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">				<span class="nx">hostIP</span><span class="p">,</span> <span class="nx">_</span><span class="p">,</span> <span class="nx">err</span> <span class="o">:=</span> <span class="nx">net</span><span class="p">.</span><span class="nf">SplitHostPort</span><span class="p">(</span><span class="nx">r</span><span class="p">.</span><span class="nx">RemoteAddr</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">				<span class="k">if</span> <span class="nx">err</span> <span class="o">!=</span> <span class="kc">nil</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">					<span class="nx">logger</span><span class="p">.</span><span class="nf">Errorw</span><span class="p">(</span><span class="s">&#34;failed to parse remote address&#34;</span><span class="p">,</span> <span class="s">&#34;error&#34;</span><span class="p">,</span> <span class="nx">err</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">					<span class="nx">hostIP</span> <span class="p">=</span> <span class="nx">r</span><span class="p">.</span><span class="nx">RemoteAddr</span>
</span></span><span class="line"><span class="cl">				<span class="p">}</span>
</span></span><span class="line"><span class="cl">				<span class="nx">logger</span><span class="p">.</span><span class="nf">Errorw</span><span class="p">(</span><span class="s">&#34;no matching API key found&#34;</span><span class="p">,</span> <span class="s">&#34;remoteIP&#34;</span><span class="p">,</span> <span class="nx">hostIP</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">				<span class="nf">RespondError</span><span class="p">(</span><span class="nx">w</span><span class="p">,</span> <span class="nx">http</span><span class="p">.</span><span class="nx">StatusUnauthorized</span><span class="p">,</span> <span class="s">&#34;invalid api key&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">				<span class="k">return</span>
</span></span><span class="line"><span class="cl">			<span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">			<span class="nx">next</span><span class="p">.</span><span class="nf">ServeHTTP</span><span class="p">(</span><span class="nx">w</span><span class="p">,</span> <span class="nx">r</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">		<span class="p">})</span>
</span></span><span class="line"><span class="cl">	<span class="p">}</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1">// bearerToken extracts the content from the header, striping the Bearer prefix
</span></span></span><span class="line"><span class="cl"><span class="c1"></span><span class="kd">func</span> <span class="nf">bearerToken</span><span class="p">(</span><span class="nx">r</span> <span class="o">*</span><span class="nx">http</span><span class="p">.</span><span class="nx">Request</span><span class="p">,</span> <span class="nx">header</span> <span class="kt">string</span><span class="p">)</span> <span class="p">(</span><span class="kt">string</span><span class="p">,</span> <span class="kt">error</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">	<span class="nx">rawToken</span> <span class="o">:=</span> <span class="nx">r</span><span class="p">.</span><span class="nx">Header</span><span class="p">.</span><span class="nf">Get</span><span class="p">(</span><span class="nx">header</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">	<span class="nx">pieces</span> <span class="o">:=</span> <span class="nx">strings</span><span class="p">.</span><span class="nf">SplitN</span><span class="p">(</span><span class="nx">rawToken</span><span class="p">,</span> <span class="s">&#34; &#34;</span><span class="p">,</span> <span class="mi">2</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">	<span class="k">if</span> <span class="nb">len</span><span class="p">(</span><span class="nx">pieces</span><span class="p">)</span> <span class="p">&lt;</span> <span class="mi">2</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">		<span class="k">return</span> <span class="s">&#34;&#34;</span><span class="p">,</span> <span class="nx">errors</span><span class="p">.</span><span class="nf">New</span><span class="p">(</span><span class="s">&#34;token with incorrect bearer format&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">	<span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">	<span class="nx">token</span> <span class="o">:=</span> <span class="nx">strings</span><span class="p">.</span><span class="nf">TrimSpace</span><span class="p">(</span><span class="nx">pieces</span><span class="p">[</span><span class="mi">1</span><span class="p">])</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">	<span class="k">return</span> <span class="nx">token</span><span class="p">,</span> <span class="kc">nil</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span></code></pre></div><p>A middleware is a function that takes an <code>http.Handler</code> and returns an <code>http.Handler</code>. In this code, the function <code>ApiKeyMiddleware</code> is a factory that creates an instance of the middleware with the provided configuration and logger. The <code>config.Config</code> is a struct populated from environment variables and <code>logging.Logger</code> is an interface that can be implemented using any logging library or the standard library. You could pass only the header and map of keys, but for clarity we choose to denote the dependency from this middleware to the configuration.</p>
<p>After extracting the fields that it relies on, the function creates a reverse index of the API Keys, which is originally a map from a key id/name to the key value. Using this reverse index it&rsquo;s trivial to verify if the user key is valid by doing a map lookup on line 18.</p>
<p>However, this approach expects the API Keys as plaintext values and is susceptible to timing attacks, because its validation algorithm is not constant time.</p>
<h3 id="using-key-hashes-for-validation">Using key hashes for validation</h3>
<p>To improve the key provisioning workflow, we can use a simple yet effective solution: expect the available keys to be hashes. Using this approach we can now commit our key hashes to our repository because even in the event of a data leak they could not be reversed to their original value.</p>
<p>Let&rsquo;s use the SHA256 hashing algorithm to encode our keys. For example, if one of them is <code>123456789</code> (please, do not use a key like this :D) then its hash will be:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-fallback" data-lang="fallback"><span class="line"><span class="cl">15e2b0d3c33891ebb0f1ef609ec419420c20e320ce94c65fbc8c3312448eb225
</span></span></code></pre></div><p>Now you can add this hash to your deployment script, Kubernetes Secret, etc., and commit it with peace of mind.</p>
<p>Next, we need to handle this new format on our middleware. This is what the code will look like now:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-go" data-lang="go"><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="kd">func</span> <span class="nf">ApiKeyMiddleware</span><span class="p">(</span><span class="nx">cfg</span> <span class="nx">conf</span><span class="p">.</span><span class="nx">Config</span><span class="p">,</span> <span class="nx">logger</span> <span class="nx">logging</span><span class="p">.</span><span class="nx">Logger</span><span class="p">)</span> <span class="kd">func</span><span class="p">(</span><span class="nx">handler</span> <span class="nx">http</span><span class="p">.</span><span class="nx">Handler</span><span class="p">)</span> <span class="nx">http</span><span class="p">.</span><span class="nx">Handler</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">	<span class="nx">apiKeyHeader</span> <span class="o">:=</span> <span class="nx">cfg</span><span class="p">.</span><span class="nx">APIKeyHeader</span> <span class="c1">// string
</span></span></span><span class="line"><span class="cl"><span class="c1"></span>	<span class="nx">apiKeys</span> <span class="o">:=</span> <span class="nx">cfg</span><span class="p">.</span><span class="nx">APIKeys</span> <span class="c1">// map[string]string
</span></span></span><span class="line"><span class="cl"><span class="c1"></span>
</span></span><span class="line"><span class="cl">	<span class="nx">reverseKeyIndex</span> <span class="o">:=</span> <span class="nb">make</span><span class="p">(</span><span class="kd">map</span><span class="p">[</span><span class="kt">string</span><span class="p">]</span><span class="kt">string</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">	<span class="k">for</span> <span class="nx">name</span><span class="p">,</span> <span class="nx">key</span> <span class="o">:=</span> <span class="nx">apiKeys</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">		<span class="nx">reverseKeyIndex</span><span class="p">[</span><span class="nx">key</span><span class="p">]</span> <span class="p">=</span> <span class="nx">name</span>
</span></span><span class="line"><span class="cl">	<span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">	<span class="k">return</span> <span class="kd">func</span><span class="p">(</span><span class="nx">next</span> <span class="nx">http</span><span class="p">.</span><span class="nx">Handler</span><span class="p">)</span> <span class="nx">http</span><span class="p">.</span><span class="nx">Handler</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">		<span class="k">return</span> <span class="nx">http</span><span class="p">.</span><span class="nf">HandlerFunc</span><span class="p">(</span><span class="kd">func</span><span class="p">(</span><span class="nx">w</span> <span class="nx">http</span><span class="p">.</span><span class="nx">ResponseWriter</span><span class="p">,</span> <span class="nx">r</span> <span class="o">*</span><span class="nx">http</span><span class="p">.</span><span class="nx">Request</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">			<span class="nx">apiKey</span><span class="p">,</span> <span class="nx">err</span> <span class="o">:=</span> <span class="nf">bearerToken</span><span class="p">(</span><span class="nx">r</span><span class="p">,</span> <span class="nx">apiKeyHeader</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">			<span class="k">if</span> <span class="nx">err</span> <span class="o">!=</span> <span class="kc">nil</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">				<span class="nx">logger</span><span class="p">.</span><span class="nf">Errorw</span><span class="p">(</span><span class="s">&#34;request failed API key authentication&#34;</span><span class="p">,</span> <span class="s">&#34;error&#34;</span><span class="p">,</span> <span class="nx">err</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">				<span class="nf">RespondError</span><span class="p">(</span><span class="nx">w</span><span class="p">,</span> <span class="nx">http</span><span class="p">.</span><span class="nx">StatusUnauthorized</span><span class="p">,</span> <span class="s">&#34;invalid API key&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">				<span class="k">return</span>
</span></span><span class="line"><span class="cl">			<span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">			<span class="nx">_</span><span class="p">,</span> <span class="nx">ok</span> <span class="o">:=</span> <span class="nf">apiKeyIsValid</span><span class="p">(</span><span class="nx">apiKey</span><span class="p">,</span> <span class="nx">reverseKeyIndex</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">			<span class="k">if</span> <span class="p">!</span><span class="nx">ok</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">				<span class="nx">hostIP</span><span class="p">,</span> <span class="nx">_</span><span class="p">,</span> <span class="nx">err</span> <span class="o">:=</span> <span class="nx">net</span><span class="p">.</span><span class="nf">SplitHostPort</span><span class="p">(</span><span class="nx">r</span><span class="p">.</span><span class="nx">RemoteAddr</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">				<span class="k">if</span> <span class="nx">err</span> <span class="o">!=</span> <span class="kc">nil</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">					<span class="nx">logger</span><span class="p">.</span><span class="nf">Errorw</span><span class="p">(</span><span class="s">&#34;failed to parse remote address&#34;</span><span class="p">,</span> <span class="s">&#34;error&#34;</span><span class="p">,</span> <span class="nx">err</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">					<span class="nx">hostIP</span> <span class="p">=</span> <span class="nx">r</span><span class="p">.</span><span class="nx">RemoteAddr</span>
</span></span><span class="line"><span class="cl">				<span class="p">}</span>
</span></span><span class="line"><span class="cl">				<span class="nx">logger</span><span class="p">.</span><span class="nf">Errorw</span><span class="p">(</span><span class="s">&#34;no matching API key found&#34;</span><span class="p">,</span> <span class="s">&#34;remoteIP&#34;</span><span class="p">,</span> <span class="nx">hostIP</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">				<span class="nf">RespondError</span><span class="p">(</span><span class="nx">w</span><span class="p">,</span> <span class="nx">http</span><span class="p">.</span><span class="nx">StatusUnauthorized</span><span class="p">,</span> <span class="s">&#34;invalid api key&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">				<span class="k">return</span>
</span></span><span class="line"><span class="cl">			<span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">			<span class="nx">next</span><span class="p">.</span><span class="nf">ServeHTTP</span><span class="p">(</span><span class="nx">w</span><span class="p">,</span> <span class="nx">r</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">		<span class="p">})</span>
</span></span><span class="line"><span class="cl">	<span class="p">}</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1">// apiKeyIsValid checks if the given API key is valid and returns the principal if it is.
</span></span></span><span class="line"><span class="cl"><span class="c1"></span><span class="kd">func</span> <span class="nf">apiKeyIsValid</span><span class="p">(</span><span class="nx">rawKey</span> <span class="kt">string</span><span class="p">,</span> <span class="nx">availableKeys</span> <span class="kd">map</span><span class="p">[</span><span class="kt">string</span><span class="p">][]</span><span class="kt">byte</span><span class="p">)</span> <span class="p">(</span><span class="kt">string</span><span class="p">,</span> <span class="kt">bool</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">	<span class="nx">hash</span> <span class="o">:=</span> <span class="nx">sha256</span><span class="p">.</span><span class="nf">Sum256</span><span class="p">([]</span><span class="nb">byte</span><span class="p">(</span><span class="nx">rawKey</span><span class="p">))</span>
</span></span><span class="line"><span class="cl">	<span class="nx">key</span> <span class="o">:=</span> <span class="nb">string</span><span class="p">(</span><span class="nx">hash</span><span class="p">[:])</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">	<span class="nx">name</span><span class="p">,</span> <span class="nx">found</span> <span class="o">:=</span> <span class="nx">reverseKeyIndex</span><span class="p">[</span><span class="nx">apiKey</span><span class="p">]</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">	<span class="k">return</span> <span class="nx">name</span><span class="p">,</span> <span class="nx">found</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1">// bearerToken function omitted..
</span></span></span></code></pre></div><p>Here we extracted the logic to validate the key into a function that, before checking the equality of the user key against the available ones, encodes the user key using the same SHA256 algorithm.</p>
<p>This simple step improved a lot our security posture without adding much complexity. Now we can have the benefits of version control, like change history and easy detection when someone changes a key hash.</p>
<p>This approach works well when there are few keys to be managed, and you want to follow a GitOps approach. However, if you need to scale the key management, allow for self-service key requests and automatic rotation, you may want to look for a solution like <a href="https://www.vaultproject.io">Hashicorp Vault</a>. Even using an external secret store I still believe this strategy, to rely on key hashes to be valid, because your external secret store can persist both the original key and the hash, and the access policy for the application can have fewer privileges in such a way that it can only read the hashes.</p>
<h3 id="constant-time-key-verification">Constant time key verification</h3>
<p>Once we have a better strategy to provision our keys, we need to defend ourselves against them being exfiltrated by timing attacks. The solution for this kind of vulnerability is to use an algorithm that takes the same time to produce a result whether the keys are equal or not. This is called a constant time comparison, and the Go Standard Library offers us an implementation in the <code>crypto/subtle</code> package that is perfect to solve most of our problems. Hence, we can update our code to use this package:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-go" data-lang="go"><span class="line"><span class="cl"><span class="kd">func</span> <span class="nf">ApiKeyMiddleware</span><span class="p">(</span><span class="nx">cfg</span> <span class="nx">conf</span><span class="p">.</span><span class="nx">Config</span><span class="p">,</span> <span class="nx">logger</span> <span class="nx">logging</span><span class="p">.</span><span class="nx">Logger</span><span class="p">)</span> <span class="p">(</span><span class="kd">func</span><span class="p">(</span><span class="nx">handler</span> <span class="nx">http</span><span class="p">.</span><span class="nx">Handler</span><span class="p">)</span> <span class="nx">http</span><span class="p">.</span><span class="nx">Handler</span><span class="p">,</span> <span class="kt">error</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">	<span class="nx">apiKeyHeader</span> <span class="o">:=</span> <span class="nx">cfg</span><span class="p">.</span><span class="nx">APIKeyHeader</span>
</span></span><span class="line"><span class="cl">	<span class="nx">apiKeys</span> <span class="o">:=</span> <span class="nx">cfg</span><span class="p">.</span><span class="nx">APIKeys</span>
</span></span><span class="line"><span class="cl">	<span class="nx">apiKeyMaxLen</span> <span class="o">:=</span> <span class="nx">cfg</span><span class="p">.</span><span class="nx">APIKeyMaxLen</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">	<span class="nx">decodedAPIKeys</span> <span class="o">:=</span> <span class="nb">make</span><span class="p">(</span><span class="kd">map</span><span class="p">[</span><span class="kt">string</span><span class="p">][]</span><span class="kt">byte</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">	<span class="k">for</span> <span class="nx">name</span><span class="p">,</span> <span class="nx">value</span> <span class="o">:=</span> <span class="k">range</span> <span class="nx">apiKeys</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">		<span class="nx">decodedKey</span><span class="p">,</span> <span class="nx">err</span> <span class="o">:=</span> <span class="nx">hex</span><span class="p">.</span><span class="nf">DecodeString</span><span class="p">(</span><span class="nx">value</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">		<span class="k">if</span> <span class="nx">err</span> <span class="o">!=</span> <span class="kc">nil</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">			<span class="k">return</span> <span class="kc">nil</span><span class="p">,</span> <span class="nx">err</span>
</span></span><span class="line"><span class="cl">		<span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">		<span class="nx">decodedAPIKeys</span><span class="p">[</span><span class="nx">name</span><span class="p">]</span> <span class="p">=</span> <span class="nx">decodedKey</span>
</span></span><span class="line"><span class="cl">	<span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">	<span class="k">return</span> <span class="kd">func</span><span class="p">(</span><span class="nx">next</span> <span class="nx">http</span><span class="p">.</span><span class="nx">Handler</span><span class="p">)</span> <span class="nx">http</span><span class="p">.</span><span class="nx">Handler</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">		<span class="k">return</span> <span class="nx">http</span><span class="p">.</span><span class="nf">HandlerFunc</span><span class="p">(</span><span class="kd">func</span><span class="p">(</span><span class="nx">w</span> <span class="nx">http</span><span class="p">.</span><span class="nx">ResponseWriter</span><span class="p">,</span> <span class="nx">r</span> <span class="o">*</span><span class="nx">http</span><span class="p">.</span><span class="nx">Request</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">			<span class="nx">ctx</span> <span class="o">:=</span> <span class="nx">r</span><span class="p">.</span><span class="nf">Context</span><span class="p">()</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">			<span class="nx">apiKey</span><span class="p">,</span> <span class="nx">err</span> <span class="o">:=</span> <span class="nf">bearerToken</span><span class="p">(</span><span class="nx">r</span><span class="p">,</span> <span class="nx">apiKeyHeader</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">			<span class="k">if</span> <span class="nx">err</span> <span class="o">!=</span> <span class="kc">nil</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">				<span class="nx">logger</span><span class="p">.</span><span class="nf">Errorw</span><span class="p">(</span><span class="s">&#34;request failed API key authentication&#34;</span><span class="p">,</span> <span class="s">&#34;error&#34;</span><span class="p">,</span> <span class="nx">err</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">				<span class="nf">RespondError</span><span class="p">(</span><span class="nx">w</span><span class="p">,</span> <span class="nx">http</span><span class="p">.</span><span class="nx">StatusUnauthorized</span><span class="p">,</span> <span class="s">&#34;invalid API key&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">				<span class="k">return</span>
</span></span><span class="line"><span class="cl">			<span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">			<span class="k">if</span> <span class="nx">_</span><span class="p">,</span> <span class="nx">ok</span> <span class="o">:=</span> <span class="nf">apiKeyIsValid</span><span class="p">(</span><span class="nx">apiKey</span><span class="p">,</span> <span class="nx">decodedAPIKeys</span><span class="p">);</span> <span class="p">!</span><span class="nx">ok</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">				 <span class="nx">hostIP</span><span class="p">,</span> <span class="nx">_</span><span class="p">,</span> <span class="nx">err</span> <span class="o">:=</span> <span class="nx">net</span><span class="p">.</span><span class="nf">SplitHostPort</span><span class="p">(</span><span class="nx">r</span><span class="p">.</span><span class="nx">RemoteAddr</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">					<span class="k">if</span> <span class="nx">err</span> <span class="o">!=</span> <span class="kc">nil</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">						<span class="nx">logger</span><span class="p">.</span><span class="nf">Errorw</span><span class="p">(</span><span class="s">&#34;failed to parse remote address&#34;</span><span class="p">,</span> <span class="s">&#34;error&#34;</span><span class="p">,</span> <span class="nx">err</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">						<span class="nx">hostIP</span> <span class="p">=</span> <span class="nx">r</span><span class="p">.</span><span class="nx">RemoteAddr</span>
</span></span><span class="line"><span class="cl">					<span class="p">}</span>
</span></span><span class="line"><span class="cl">					<span class="nx">logger</span><span class="p">.</span><span class="nf">Errorw</span><span class="p">(</span><span class="s">&#34;no matching API key found&#34;</span><span class="p">,</span> <span class="s">&#34;remoteIP&#34;</span><span class="p">,</span> <span class="nx">hostIP</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">					<span class="nf">RespondError</span><span class="p">(</span><span class="nx">w</span><span class="p">,</span> <span class="nx">http</span><span class="p">.</span><span class="nx">StatusUnauthorized</span><span class="p">,</span> <span class="s">&#34;invalid api key&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">					<span class="k">return</span>
</span></span><span class="line"><span class="cl">			<span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">			<span class="nx">next</span><span class="p">.</span><span class="nf">ServeHTTP</span><span class="p">(</span><span class="nx">w</span><span class="p">,</span> <span class="nx">r</span><span class="p">.</span><span class="nf">WithContext</span><span class="p">(</span><span class="nx">ctx</span><span class="p">))</span>
</span></span><span class="line"><span class="cl">		<span class="p">})</span>
</span></span><span class="line"><span class="cl">	<span class="p">},</span> <span class="kc">nil</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1">// apiKeyIsValid checks if the given API key is valid and returns the principal if it is.
</span></span></span><span class="line"><span class="cl"><span class="c1"></span><span class="kd">func</span> <span class="nf">apiKeyIsValid</span><span class="p">(</span><span class="nx">rawKey</span> <span class="kt">string</span><span class="p">,</span> <span class="nx">availableKeys</span> <span class="kd">map</span><span class="p">[</span><span class="kt">string</span><span class="p">][]</span><span class="kt">byte</span><span class="p">)</span> <span class="p">(</span><span class="kt">string</span><span class="p">,</span> <span class="kt">bool</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">	<span class="nx">hash</span> <span class="o">:=</span> <span class="nx">sha256</span><span class="p">.</span><span class="nf">Sum256</span><span class="p">([]</span><span class="nb">byte</span><span class="p">(</span><span class="nx">rawKey</span><span class="p">))</span>
</span></span><span class="line"><span class="cl">	<span class="nx">key</span> <span class="o">:=</span> <span class="nx">hash</span><span class="p">[:]</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">	<span class="k">for</span> <span class="nx">name</span><span class="p">,</span> <span class="nx">value</span> <span class="o">:=</span> <span class="k">range</span> <span class="nx">availableKeys</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">		<span class="nx">contentEqual</span> <span class="o">:=</span> <span class="nx">subtle</span><span class="p">.</span><span class="nf">ConstantTimeCompare</span><span class="p">(</span><span class="nx">value</span><span class="p">,</span> <span class="nx">key</span><span class="p">)</span> <span class="o">==</span> <span class="mi">1</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">		<span class="k">if</span> <span class="nx">contentEqual</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">			<span class="k">return</span> <span class="nx">name</span><span class="p">,</span> <span class="kc">true</span>
</span></span><span class="line"><span class="cl">		<span class="p">}</span>
</span></span><span class="line"><span class="cl">	<span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">	<span class="k">return</span> <span class="s">&#34;&#34;</span><span class="p">,</span> <span class="kc">false</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1">// bearerToken function omitted...
</span></span></span></code></pre></div><p>Now, the function <code>apiKeyIsValid</code> uses <code>subtle.ConstantTimeCompare</code> to verify the user key against each available key. Since <code>subtle.ConstantTimeCompare</code> operates upon byte slices we don&rsquo;t cast our hash to string anymore and also our reversed index has gone in place of a decoded map.</p>
<p>The decoding is necessary because the string representation of our key hashes are actually a hexadecimal encoding of the binary value. Hence, we cannot just cast the string to byte slice because Go assumes all strings to be UTF-8 encoded.</p>
<blockquote>
<p>Note: for an example on how using a cast instead of the correct decoding function, the result of <code>[]byte(&quot;09&quot;)</code> is <code>110000111001</code> while <code>hex.DecodeString(&quot;09&quot;)</code> produces <code>1001</code>. Check out the live example <a href="https://go.dev/play/p/CPy16o7hvDO">here</a>.</p>
</blockquote>
<p>The major disadvantage of this solution is that now we need to iterate over all available keys before finding out if the key is incorrect. This doesn&rsquo;t scale well if there are too many keys, however one simple workaround would be to require the client to send an extra header with the key ID/name, e.g. <code>X-App-Key-ID</code>, with which you can find the key in <code>O(1)</code> and then apply the constant time comparison.</p>
<p>However, there is one subtle (<em>pun intended</em>) behavior from <code>subtle.ConstantTimeCompare</code> that we must be aware before deploying our solution to production. When the byte slices have different lengths, the functions returns earlier without performing the bitwise operations. This is natural because it does an XOR between each pair of bits from each slice, and with slices of different sizes, there would be bits from one slice without a matching pair to be combined with. <strong>Because of it, an adversary could measure that keys with the wrong length have a smaller response time than keys with the correct length, hence leaking the key length</strong>. It would only be a vulnerability if you use a short key that is easily brute-forced, but with a simple 30 character key using the UTF-8 printable characters you would have <code>30^95 = 2.12089515 × 10^140</code> possible keys.</p>
<p>Finally, we&rsquo;ve built a simple, secure and efficient API Key solution that should handle a lot of uses cases without additional infrastructure or complexity. Using a basic understanding of threats and the Golang standard library, we could do a security-oriented design instead of leaving security as an after-though in an iterative way.</p>
<hr>
<p>Photo by <a href="https://unsplash.com/@silas_crioco?utm_source=unsplash&utm_medium=referral&utm_content=creditCopyText">Silas Köhler</a> on <a href="https://unsplash.com/s/photos/key?utm_source=unsplash&utm_medium=referral&utm_content=creditCopyText">Unsplash</a>.</p>
]]></content:encoded>
    </item>
    
  </channel>
</rss>
