<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en_US"><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://stevenpisani.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://stevenpisani.com/" rel="alternate" type="text/html" hreflang="en_US" /><updated>2026-09-28T18:15:08-04:00</updated><id>https://stevenpisani.com/feed.xml</id><title type="html">Steve Pisani</title><subtitle>Steve Pisani&apos;s corner of the internet. Data architect in Philadelphia — plus writing, experiments, books, and whatever else I&apos;m tinkering with.</subtitle><author><name>Steve Pisani</name></author><entry><title type="html">Building Modern Data Pipelines: Lessons from the Trenches</title><link href="https://stevenpisani.com/2025/01/27/Building-Modern-Data-Pipelines.html" rel="alternate" type="text/html" title="Building Modern Data Pipelines: Lessons from the Trenches" /><published>2025-01-27T05:00:00-05:00</published><updated>2025-01-27T05:00:00-05:00</updated><id>https://stevenpisani.com/2025/01/27/Building-Modern-Data-Pipelines</id><content type="html" xml:base="https://stevenpisani.com/2025/01/27/Building-Modern-Data-Pipelines.html"><![CDATA[<p>After years of building data pipelines across different organizations—from Fortune 100 companies to scrappy startups—I’ve learned that the technical implementation is often the easy part. The real challenges lie in designing systems that are maintainable, scalable, and actually solve business problems.</p>

<h2 id="the-foundation-understanding-your-data">The Foundation: Understanding Your Data</h2>

<p>Before writing a single line of code, spend time understanding your data landscape. I’ve seen too many projects fail because teams jumped straight into implementation without properly mapping their data sources, understanding data quality issues, or defining clear success metrics.</p>

<p><strong>Key questions to ask:</strong></p>
<ul>
  <li>What are the upstream data sources and their reliability patterns?</li>
  <li>What’s the acceptable latency for different use cases?</li>
  <li>How will you handle schema evolution?</li>
  <li>What are the downstream dependencies?</li>
</ul>

<h2 id="design-principles-that-actually-matter">Design Principles That Actually Matter</h2>

<h3 id="1-idempotency-is-non-negotiable">1. Idempotency is Non-Negotiable</h3>

<p>Every pipeline component should be idempotent. If you run the same process twice with the same inputs, you should get the same outputs. This isn’t just good practice—it’s essential for debugging, recovery, and maintaining sanity during 3 AM incidents.</p>

<h3 id="2-fail-fast-and-fail-clearly">2. Fail Fast and Fail Clearly</h3>

<p>Design your pipelines to fail quickly when something goes wrong, and make sure the error messages are actionable. Nothing is worse than a pipeline that silently produces incorrect data or fails with cryptic error messages.</p>

<h3 id="3-observability-from-day-one">3. Observability from Day One</h3>

<p>Monitoring isn’t something you add later—it’s part of the architecture. Every pipeline should emit metrics about data volume, processing time, and data quality. Your future self (and your teammates) will thank you.</p>

<h2 id="the-tools-dont-matter-as-much-as-you-think">The Tools Don’t Matter (As Much As You Think)</h2>

<p>I’ve built successful pipelines with everything from cron jobs and Python scripts to sophisticated orchestration platforms like Airflow and Prefect. The tool choice matters less than having clear requirements and good engineering practices.</p>

<p>That said, here are some patterns that have served me well:</p>

<ul>
  <li><strong>Start simple</strong>: Begin with the simplest solution that could work, then add complexity as needed</li>
  <li><strong>Embrace SQL</strong>: Modern SQL engines are incredibly powerful—don’t reinvent the wheel</li>
  <li><strong>Version everything</strong>: Code, schemas, configurations, and even your data when possible</li>
</ul>

<h2 id="common-pitfalls-to-avoid">Common Pitfalls to Avoid</h2>

<h3 id="the-big-bang-migration">The “Big Bang” Migration</h3>

<p>I’ve never seen a successful “big bang” data migration. Always plan for incremental migration with parallel systems running during the transition period.</p>

<h3 id="over-engineering-for-scale">Over-Engineering for Scale</h3>

<p>Build for your current scale plus one order of magnitude, not for Google-scale unless you’re actually Google. Premature optimization in data pipelines often leads to unnecessary complexity.</p>

<h3 id="ignoring-data-quality">Ignoring Data Quality</h3>

<p>Data quality issues compound over time. Implement data quality checks early and make them part of your pipeline, not an afterthought.</p>

<h2 id="looking-forward">Looking Forward</h2>

<p>The data engineering landscape continues to evolve rapidly. New tools and paradigms emerge regularly, but the fundamental principles remain constant: understand your requirements, design for maintainability, and always keep the end user in mind.</p>

<p>What challenges have you faced building data pipelines? I’d love to hear about your experiences and lessons learned.</p>

<hr />

<p><em>Have questions about data pipeline architecture or want to discuss a specific challenge? Feel free to reach out—I’m always happy to chat about data engineering problems.</em></p>]]></content><author><name>Steve</name></author><summary type="html"><![CDATA[Key insights and best practices for building scalable, maintainable data pipelines in modern organizations.]]></summary></entry><entry><title type="html">The Art of SQL Optimization: Beyond the Basics</title><link href="https://stevenpisani.com/2025/01/20/The-Art-of-SQL-Optimization.html" rel="alternate" type="text/html" title="The Art of SQL Optimization: Beyond the Basics" /><published>2025-01-20T09:30:00-05:00</published><updated>2025-01-20T09:30:00-05:00</updated><id>https://stevenpisani.com/2025/01/20/The-Art-of-SQL-Optimization</id><content type="html" xml:base="https://stevenpisani.com/2025/01/20/The-Art-of-SQL-Optimization.html"><![CDATA[<p>SQL optimization is often treated as a dark art, but it doesn’t have to be. After optimizing queries across different database systems and scales, I’ve found that most performance issues stem from a few common patterns. Let me share some advanced techniques that go beyond the usual “add an index” advice.</p>

<h2 id="understanding-query-execution-plans">Understanding Query Execution Plans</h2>

<p>Before optimizing anything, you need to understand how your database is executing queries. Every major database system provides tools to examine execution plans:</p>

<ul>
  <li><strong>PostgreSQL</strong>: <code class="language-plaintext highlighter-rouge">EXPLAIN ANALYZE</code></li>
  <li><strong>SQL Server</strong>: <code class="language-plaintext highlighter-rouge">SET STATISTICS IO ON</code></li>
  <li><strong>MySQL</strong>: <code class="language-plaintext highlighter-rouge">EXPLAIN FORMAT=JSON</code></li>
</ul>

<p>The key is learning to read these plans and identify bottlenecks. Look for:</p>
<ul>
  <li>Sequential scans on large tables</li>
  <li>Nested loop joins with high row counts</li>
  <li>Sort operations on large datasets</li>
  <li>Hash joins that spill to disk</li>
</ul>

<h2 id="advanced-indexing-strategies">Advanced Indexing Strategies</h2>

<h3 id="partial-indexes">Partial Indexes</h3>

<p>Instead of indexing entire columns, create indexes on subsets of data:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1">-- Instead of indexing all orders</span>
<span class="k">CREATE</span> <span class="k">INDEX</span> <span class="n">idx_orders_status</span> <span class="k">ON</span> <span class="n">orders</span><span class="p">(</span><span class="n">status</span><span class="p">);</span>

<span class="c1">-- Index only active orders</span>
<span class="k">CREATE</span> <span class="k">INDEX</span> <span class="n">idx_active_orders</span> <span class="k">ON</span> <span class="n">orders</span><span class="p">(</span><span class="n">customer_id</span><span class="p">,</span> <span class="n">order_date</span><span class="p">)</span> 
<span class="k">WHERE</span> <span class="n">status</span> <span class="o">=</span> <span class="s1">'active'</span><span class="p">;</span>
</code></pre></div></div>

<h3 id="covering-indexes">Covering Indexes</h3>

<p>Include frequently accessed columns in your index to avoid table lookups:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">CREATE</span> <span class="k">INDEX</span> <span class="n">idx_orders_covering</span> <span class="k">ON</span> <span class="n">orders</span><span class="p">(</span><span class="n">customer_id</span><span class="p">,</span> <span class="n">order_date</span><span class="p">)</span> 
<span class="n">INCLUDE</span> <span class="p">(</span><span class="n">total_amount</span><span class="p">,</span> <span class="n">status</span><span class="p">);</span>
</code></pre></div></div>

<h3 id="expression-indexes">Expression Indexes</h3>

<p>Index computed values that you frequently filter on:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">CREATE</span> <span class="k">INDEX</span> <span class="n">idx_orders_month</span> <span class="k">ON</span> <span class="n">orders</span><span class="p">(</span><span class="k">EXTRACT</span><span class="p">(</span><span class="k">MONTH</span> <span class="k">FROM</span> <span class="n">order_date</span><span class="p">));</span>
</code></pre></div></div>

<h2 id="query-rewriting-techniques">Query Rewriting Techniques</h2>

<h3 id="window-functions-vs-self-joins">Window Functions vs. Self-Joins</h3>

<p>Replace expensive self-joins with window functions:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1">-- Instead of this expensive self-join</span>
<span class="k">SELECT</span> <span class="n">o1</span><span class="p">.</span><span class="n">customer_id</span><span class="p">,</span> <span class="n">o1</span><span class="p">.</span><span class="n">order_date</span><span class="p">,</span> <span class="n">o1</span><span class="p">.</span><span class="n">total_amount</span>
<span class="k">FROM</span> <span class="n">orders</span> <span class="n">o1</span>
<span class="k">JOIN</span> <span class="p">(</span>
    <span class="k">SELECT</span> <span class="n">customer_id</span><span class="p">,</span> <span class="k">MAX</span><span class="p">(</span><span class="n">order_date</span><span class="p">)</span> <span class="k">as</span> <span class="n">max_date</span>
    <span class="k">FROM</span> <span class="n">orders</span>
    <span class="k">GROUP</span> <span class="k">BY</span> <span class="n">customer_id</span>
<span class="p">)</span> <span class="n">o2</span> <span class="k">ON</span> <span class="n">o1</span><span class="p">.</span><span class="n">customer_id</span> <span class="o">=</span> <span class="n">o2</span><span class="p">.</span><span class="n">customer_id</span> 
    <span class="k">AND</span> <span class="n">o1</span><span class="p">.</span><span class="n">order_date</span> <span class="o">=</span> <span class="n">o2</span><span class="p">.</span><span class="n">max_date</span><span class="p">;</span>

<span class="c1">-- Use window functions</span>
<span class="k">SELECT</span> <span class="n">customer_id</span><span class="p">,</span> <span class="n">order_date</span><span class="p">,</span> <span class="n">total_amount</span>
<span class="k">FROM</span> <span class="p">(</span>
    <span class="k">SELECT</span> <span class="n">customer_id</span><span class="p">,</span> <span class="n">order_date</span><span class="p">,</span> <span class="n">total_amount</span><span class="p">,</span>
           <span class="n">ROW_NUMBER</span><span class="p">()</span> <span class="n">OVER</span> <span class="p">(</span><span class="k">PARTITION</span> <span class="k">BY</span> <span class="n">customer_id</span> <span class="k">ORDER</span> <span class="k">BY</span> <span class="n">order_date</span> <span class="k">DESC</span><span class="p">)</span> <span class="k">as</span> <span class="n">rn</span>
    <span class="k">FROM</span> <span class="n">orders</span>
<span class="p">)</span> <span class="n">ranked</span>
<span class="k">WHERE</span> <span class="n">rn</span> <span class="o">=</span> <span class="mi">1</span><span class="p">;</span>
</code></pre></div></div>

<h3 id="exists-vs-in">EXISTS vs. IN</h3>

<p>For large datasets, <code class="language-plaintext highlighter-rouge">EXISTS</code> often performs better than <code class="language-plaintext highlighter-rouge">IN</code>:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1">-- Instead of IN</span>
<span class="k">SELECT</span> <span class="o">*</span> <span class="k">FROM</span> <span class="n">customers</span> 
<span class="k">WHERE</span> <span class="n">customer_id</span> <span class="k">IN</span> <span class="p">(</span><span class="k">SELECT</span> <span class="n">customer_id</span> <span class="k">FROM</span> <span class="n">orders</span> <span class="k">WHERE</span> <span class="n">order_date</span> <span class="o">&gt;</span> <span class="s1">'2024-01-01'</span><span class="p">);</span>

<span class="c1">-- Use EXISTS</span>
<span class="k">SELECT</span> <span class="o">*</span> <span class="k">FROM</span> <span class="n">customers</span> <span class="k">c</span>
<span class="k">WHERE</span> <span class="k">EXISTS</span> <span class="p">(</span><span class="k">SELECT</span> <span class="mi">1</span> <span class="k">FROM</span> <span class="n">orders</span> <span class="n">o</span> 
              <span class="k">WHERE</span> <span class="n">o</span><span class="p">.</span><span class="n">customer_id</span> <span class="o">=</span> <span class="k">c</span><span class="p">.</span><span class="n">customer_id</span> 
              <span class="k">AND</span> <span class="n">o</span><span class="p">.</span><span class="n">order_date</span> <span class="o">&gt;</span> <span class="s1">'2024-01-01'</span><span class="p">);</span>
</code></pre></div></div>

<h2 id="data-type-optimization">Data Type Optimization</h2>

<p>Choose the right data types—it matters more than you think:</p>

<ul>
  <li>Use <code class="language-plaintext highlighter-rouge">INT</code> instead of <code class="language-plaintext highlighter-rouge">BIGINT</code> when possible</li>
  <li>Choose appropriate <code class="language-plaintext highlighter-rouge">VARCHAR</code> lengths</li>
  <li>Consider <code class="language-plaintext highlighter-rouge">ENUM</code> types for limited value sets</li>
  <li>Use proper date/time types instead of strings</li>
</ul>

<h2 id="partitioning-strategies">Partitioning Strategies</h2>

<p>For very large tables, partitioning can dramatically improve performance:</p>

<h3 id="range-partitioning">Range Partitioning</h3>
<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1">-- Partition by date range</span>
<span class="k">CREATE</span> <span class="k">TABLE</span> <span class="n">orders_2024</span> <span class="k">PARTITION</span> <span class="k">OF</span> <span class="n">orders</span>
<span class="k">FOR</span> <span class="k">VALUES</span> <span class="k">FROM</span> <span class="p">(</span><span class="s1">'2024-01-01'</span><span class="p">)</span> <span class="k">TO</span> <span class="p">(</span><span class="s1">'2025-01-01'</span><span class="p">);</span>
</code></pre></div></div>

<h3 id="hash-partitioning">Hash Partitioning</h3>
<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1">-- Distribute data evenly across partitions</span>
<span class="k">CREATE</span> <span class="k">TABLE</span> <span class="n">orders_hash_1</span> <span class="k">PARTITION</span> <span class="k">OF</span> <span class="n">orders</span>
<span class="k">FOR</span> <span class="k">VALUES</span> <span class="k">WITH</span> <span class="p">(</span><span class="n">MODULUS</span> <span class="mi">4</span><span class="p">,</span> <span class="n">REMAINDER</span> <span class="mi">0</span><span class="p">);</span>
</code></pre></div></div>

<h2 id="monitoring-and-maintenance">Monitoring and Maintenance</h2>

<h3 id="query-performance-monitoring">Query Performance Monitoring</h3>

<p>Set up monitoring for:</p>
<ul>
  <li>Slow query logs</li>
  <li>Query execution statistics</li>
  <li>Index usage statistics</li>
  <li>Lock contention</li>
</ul>

<h3 id="regular-maintenance-tasks">Regular Maintenance Tasks</h3>

<ul>
  <li>Update table statistics regularly</li>
  <li>Rebuild fragmented indexes</li>
  <li>Monitor and clean up unused indexes</li>
  <li>Analyze query patterns and adjust accordingly</li>
</ul>

<h2 id="database-specific-optimizations">Database-Specific Optimizations</h2>

<h3 id="postgresql">PostgreSQL</h3>
<ul>
  <li>Use <code class="language-plaintext highlighter-rouge">pg_stat_statements</code> for query analysis</li>
  <li>Leverage materialized views for complex aggregations</li>
  <li>Consider <code class="language-plaintext highlighter-rouge">BRIN</code> indexes for time-series data</li>
</ul>

<h3 id="sql-server">SQL Server</h3>
<ul>
  <li>Use columnstore indexes for analytical workloads</li>
  <li>Implement query store for performance tracking</li>
  <li>Consider in-memory OLTP for high-throughput scenarios</li>
</ul>

<h2 id="the-human-factor">The Human Factor</h2>

<p>Remember that the most optimized query is useless if it doesn’t solve the right business problem. Always:</p>

<ol>
  <li>Understand the business requirements</li>
  <li>Measure before optimizing</li>
  <li>Test with realistic data volumes</li>
  <li>Document your optimization decisions</li>
</ol>

<h2 id="conclusion">Conclusion</h2>

<p>SQL optimization is both an art and a science. While these techniques can dramatically improve performance, remember that premature optimization is the root of all evil. Focus on the queries that matter most to your users and business outcomes.</p>

<p>The best optimization is often the simplest one—sometimes the answer isn’t a more complex query, but rather a different approach to the problem entirely.</p>

<hr />

<p><em>What SQL optimization challenges are you facing? I’d love to hear about your experiences and discuss specific scenarios.</em></p>]]></content><author><name>Steve</name></author><summary type="html"><![CDATA[Advanced SQL optimization techniques that go beyond basic indexing and query structure.]]></summary></entry><entry><title type="html">SQL Guide</title><link href="https://stevenpisani.com/2021/02/15/SQL-Style-Guide.html" rel="alternate" type="text/html" title="SQL Guide" /><published>2021-02-15T07:00:00-05:00</published><updated>2021-02-15T07:00:00-05:00</updated><id>https://stevenpisani.com/2021/02/15/SQL%20Style%20Guide</id><content type="html" xml:base="https://stevenpisani.com/2021/02/15/SQL-Style-Guide.html"><![CDATA[<p>When writing SQL, it is important to stay consistant to help anyone reading your queries (including <em>future</em> you) quickly understand the logic in your code and pick out errors or opportunities for improvement. By sticking to a style guide, you can remove many of the headaches associated with reading, writing, and reviewing SQL in a team or by yourself.</p>

<p>These suggestions are not written in stone and, if you have a suggestion or disagreement, I would love to chat about it.</p>

<blockquote>
  <p><strong>Update:</strong> you don’t have to memorize any of this. <a href="/assets/files/.sqlfluff" download=".sqlfluff">Download my <code class="language-plaintext highlighter-rouge">.sqlfluff</code> config</a> and let <a href="https://sqlfluff.com">SQLFluff</a> enforce it, or paste a query into the <a href="/lab/sql-formatter">SQL formatter in the lab</a> to see it in action.</p>
</blockquote>

<h3 id="example">Example</h3>
<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight"><code>  <span class="k">with</span> <span class="n">date_spine</span> <span class="k">as</span> <span class="p">(</span>
    <span class="k">select</span>
      <span class="k">day</span>
    <span class="k">from</span> <span class="n">date_tables</span>
  <span class="p">)</span>

  <span class="p">,</span> <span class="n">account_revenue</span> <span class="k">as</span> <span class="p">(</span>
    <span class="k">select</span>
      <span class="n">account_id</span>
      <span class="p">,</span> <span class="n">created_date</span> <span class="k">as</span> <span class="n">joined_date</span>
      <span class="p">,</span> <span class="n">revenue</span>
    <span class="k">from</span> <span class="n">customers</span><span class="p">.</span><span class="n">revenue</span>
    <span class="k">where</span> <span class="n">revenue</span> <span class="o">&gt;</span> <span class="mi">0</span>
      <span class="k">and</span> <span class="nb">date</span><span class="p">(</span><span class="n">created_date</span><span class="p">)</span> <span class="o">&gt;=</span> <span class="p">(</span><span class="k">select</span> <span class="k">min</span><span class="p">(</span><span class="k">g</span><span class="p">.</span><span class="n">created_at</span><span class="p">)</span> <span class="k">as</span> <span class="n">first_at</span> <span class="k">from</span> <span class="n">customers</span><span class="p">.</span><span class="n">groups</span> <span class="k">as</span> <span class="k">g</span><span class="p">)</span>
  <span class="p">)</span>

  <span class="k">select</span>
    <span class="k">c</span><span class="p">.</span><span class="n">audience_id</span>
    <span class="p">,</span> <span class="n">d</span><span class="p">.</span><span class="k">day</span> <span class="k">as</span> <span class="n">date_of_interest</span>
    <span class="p">,</span> <span class="k">sum</span><span class="p">(</span><span class="n">ar</span><span class="p">.</span><span class="n">revenue</span><span class="p">)</span> <span class="k">as</span> <span class="n">daily_audience_revenue</span>
  <span class="k">from</span> <span class="n">customers</span><span class="p">.</span><span class="n">groups</span> <span class="k">as</span> <span class="k">c</span>
  <span class="k">inner</span> <span class="k">join</span> <span class="n">date_spine</span> <span class="k">as</span> <span class="n">d</span>
    <span class="k">on</span> <span class="k">c</span><span class="p">.</span><span class="n">created_at</span> <span class="o">&lt;=</span> <span class="n">d</span><span class="p">.</span><span class="k">day</span>
  <span class="k">left</span> <span class="k">outer</span> <span class="k">join</span> <span class="n">account_revenue</span> <span class="k">as</span> <span class="n">ar</span>
    <span class="k">on</span> <span class="k">c</span><span class="p">.</span><span class="n">account_id</span> <span class="o">=</span> <span class="n">ar</span><span class="p">.</span><span class="n">account_id</span>
  <span class="k">where</span> <span class="k">c</span><span class="p">.</span><span class="n">created_at</span> <span class="o">&lt;=</span> <span class="n">ar</span><span class="p">.</span><span class="n">joined_date</span>
  <span class="k">group</span> <span class="k">by</span>
    <span class="k">c</span><span class="p">.</span><span class="n">audience_id</span>
    <span class="p">,</span> <span class="n">d</span><span class="p">.</span><span class="k">day</span>
  <span class="k">order</span> <span class="k">by</span>
    <span class="k">c</span><span class="p">.</span><span class="n">audience_id</span>
    <span class="p">,</span> <span class="n">d</span><span class="p">.</span><span class="k">day</span>
</code></pre></div></div>

<hr />

<h1 id="rules">Rules</h1>

<p>CONTEXT DECOUPLED UNDERSTANDABILITY!!</p>

<p>The goal of constant SQL formatting is to improve development, review, and understanding time.</p>

<h2 id="capitalization">Capitalization</h2>

<ul>
  <li>
    <p><em>no capitalizing of keywords</em></p>

    <p>keywords are already highlighted in the editor. the additional keystrokes are redundant.</p>
  </li>
</ul>

<h2 id="aliasing--naming">Aliasing &amp; Naming</h2>

<ul>
  <li><em>Always use</em> <em><code class="language-plaintext highlighter-rouge">as</code></em> <em>when aliasing columns</em></li>
  <li>Use snake_case</li>
  <li><em>Use</em> <em><code class="language-plaintext highlighter-rouge">is_</code></em> <em>prefix when naming boolean fields</em></li>
  <li><em>Alway rename aggregates and function fields</em></li>
  <li><em>If joins, alias all tables and include when addressing fields</em></li>
  <li><em>Use distinct aliases</em></li>
  <li><em>Use meaningful CTE names</em></li>
</ul>

<h2 id="alignment">Alignment</h2>

<ul>
  <li><em>Left align new lines at their respective hierarchal level</em></li>
  <li>For single conditions (<code class="language-plaintext highlighter-rouge">where</code>, <code class="language-plaintext highlighter-rouge">on</code>, <code class="language-plaintext highlighter-rouge">when</code>, etc.), leave on same line. For multiple, use new line</li>
</ul>

<h2 id="spacing">Spacing</h2>

<ul>
  <li>
    <p><em>2 spaces for indents</em></p>

    <p>easier to manually do. no super long lines due to front spacing</p>
  </li>
  <li>
    <p>S<em>tart columns below the</em> <em><code class="language-plaintext highlighter-rouge">select</code></em> <em>indented</em></p>

    <p>Readability.</p>
  </li>
  <li>Break long lists of <code class="language-plaintext highlighter-rouge">in</code> values into multiple lines</li>
  <li>
    <p>A<em>lways end on new line</em></p>

    <p><a href="https://unix.stackexchange.com/questions/18743/whats-the-point-in-adding-a-new-line-to-the-end-of-a-file">unix expects all text files to end with \n</a></p>
  </li>
</ul>

<h2 id="separators---and">Separators (<code class="language-plaintext highlighter-rouge">,</code> &amp; <code class="language-plaintext highlighter-rouge">and</code>)</h2>

<ul>
  <li><em>put em in front</em></li>
</ul>

<h2 id="grouping">Grouping</h2>

<ul>
  <li><em>Grouping columns should go first in the</em> <em><code class="language-plaintext highlighter-rouge">select</code></em></li>
  <li><em>Group using column names or numbers but not both; prefer names to numbers</em></li>
</ul>

<h2 id="ctes--subqueries">CTEs &amp; Subqueries</h2>

<ul>
  <li>Use CTEs over subqueries</li>
</ul>

<hr />

<h1 id="enforce-it-automatically">Enforce it automatically</h1>

<p>Style guides only work if nobody has to think about them. Everything above (except the naming judgment calls) is encoded in a <a href="https://sqlfluff.com">SQLFluff</a> config:</p>

<p><a class="btn" href="/assets/files/.sqlfluff" download=".sqlfluff">Download .sqlfluff</a></p>

<p>Drop it in the root of your project (next to <code class="language-plaintext highlighter-rouge">dbt_project.yml</code> if you’re using dbt), then:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pip <span class="nb">install </span>sqlfluff
sqlfluff lint models/   <span class="c"># what's wrong</span>
sqlfluff fix models/    <span class="c"># fix what can be fixed</span>
</code></pre></div></div>

<p>It defaults to the Snowflake dialect and the Jinja templater. Change <code class="language-plaintext highlighter-rouge">dialect</code> for your warehouse, and install <code class="language-plaintext highlighter-rouge">sqlfluff-templater-dbt</code> if you want it to compile dbt refs.</p>]]></content><author><name>Steve</name></author><summary type="html"><![CDATA[When writing SQL, it is important to stay consistant to help anyone reading your queries (including future you) quickly understand the logic in your code and pick out errors or opportunities for improvement. By sticking to a style guide, you can remove many of the headaches associated with reading, writing, and reviewing SQL in a team or by yourself.]]></summary></entry><entry><title type="html">Rockets and Horses</title><link href="https://stevenpisani.com/2020/11/05/Rockets-and-Horses.html" rel="alternate" type="text/html" title="Rockets and Horses" /><published>2020-11-05T10:52:00-05:00</published><updated>2020-11-05T10:52:00-05:00</updated><id>https://stevenpisani.com/2020/11/05/Rockets%20and%20Horses</id><content type="html" xml:base="https://stevenpisani.com/2020/11/05/Rockets-and-Horses.html"><![CDATA[<p>One of my favorite bits of Internet lore is the connection between the design of the rocket boosters on the space shuttle and the size of Roman war chariots. Here is the rough retelling:</p>

<blockquote>
  <h3><strong><em>“</em></strong></h3>
  <p>The US standard railroad gauge (distance between the rails) is 4ft, 8.5in. This gauge is used because the English built railroads to that gauge and US railroads were built by English expatriates.</p>

  <p><br /><em>Why did the English build railroads to that gauge?</em></p>

  <p>Because the first rail lines were built by the same people who built the pre-railroad tramways, and that’s the gauge they used.</p>

  <p><br /><em>Why did those wheelwrights use that gauge then?</em></p>

  <p>Because the people who built the horse-drawn trams used the same jigs and tools that they used for building wagons, which used that wheel spacing.</p>

  <p><br /><em>Why did the wagons use that odd wheel spacing?</em></p>

  <p>For the practical reason that any other spacing would break an axle on some of the old, long distance roads, because this is the measure of the old wheel ruts.</p>

  <p><br /><em>So who built these old rutted roads?</em></p>

  <p>The first long distance roads in Europe were built by Imperial Rome for their legions and used ever since. The initial ruts were first made by Roman war chariots, which were of uniform military issue. The Imperial Roman chariots were made to be just wide enough to accommodate the back-ends of two war horses.</p>

  <p><br />This story does not end there, however.</p>

  <p>Look at a NASA Space Shuttle and the two big booster rockets attached to the sides of the main fuel tank. These are solid rocket boosters or SRBs. The SRBs are made by Thiokol at their factory at Utah. The engineers who designed the SRBs might have preferred to make them a bit fatter, but the SRBs had to be shipped by train from the factory to the launch site in Florida. The railroad line from the factory runs through a tunnel in the mountains and the SRBs have to fit through that tunnel. The tunnel is slightly wider than the railroad track.</p>

  <p>So, the major design feature of what is arguably the world’s most advanced transportation system was determined by the width of a horse’s ass.</p>
</blockquote>

<p>Now, this story may not be <em>100%</em> <a href="https://www.snopes.com/fact-check/railroad-gauge-chariots/">factual</a> but, I like to believe it is pretty damn close and provides plenty of food for thought.</p>

<p>The headline in my mind is that <code class="language-plaintext highlighter-rouge">design decisions propagate</code>.</p>

<p>The initial constrain of “accommodate the back-ends of two war horses” got passed along over and over simply because the next step relied on the previous. So by the time NASA was designing their rockets it had to be taken into account.</p>

<p>At a macro level, we (humanity) have been building on top of what came before us for all of history. At each stage in the design tree of horse’s ass to space ship, no one said “welp gotta keep this to the original ass spec”. They stayed pretty close to it because it was already built. No need to re-do the work. What gets my attention though is that the rocket’s size could have been larger if the constraint wasn’t there. Would that have made them more effective? I don’t know but it doesn’t sit well with me that the design was limited based on the needs of the inital design.</p>

<p>Back down on a micro level, the decisions <em>we</em> make can propagate farther than we anticipated. Writing new backend code for your company’s web app? Well whatever comes next is probably going to build on top of it. Designing a database structure? Whatever comes next is probably going to build on top of it.</p>

<p>All of this is my long-winded way of saying: 
<em>when you’re making something, take the time to understand constraints that might be piggy-backing in your work and if they are holding you back.</em></p>

<h6 id="note--i-dont-remember-where-i-first-heard-the-srb-to-horse-story-but-i-came-across-it-most-recently-in-joe-celkos-sql-for-smarties-which-i-highly-recommend">Note: <br /> I don’t remember where I first heard the SRB to horse story but I came across it most recently in <a href="https://www.amazon.com/Joe-Celkos-SQL-Smarties-Programming/dp/0128007613?sa-no-redirect=1&amp;pldnSite=1">Joe Celkos’ “SQL for Smarties”</a> which I highly recommend.</h6>

<h6 id="note2-tangentially-related-here-is-richard-feynmans-explaination-of-why-train-acels-dont-need-differentials-httpsyoutubewawdvbifkos">Note2: Tangentially related, here is Richard Feynman’s explaination of why train acels don’t need differentials: https://youtu.be/WAwDvbIfkos</h6>]]></content><author><name>Steve</name></author><summary type="html"><![CDATA[One of my favorite bits of Internet lore is the connection between the design of the rocket boosters on the space shuttle and the size of Roman war chariots. Here is the rough retelling:]]></summary></entry></feed>