| | This new guide from Hex breaks down what context actually means for AI analytics, from semantic models and metadata to existing analyses and business rules and shows how to build it over time without turning it into a massive documentation project. *Message from this week's sponsor, Hex Technologies |
| | |
| | 🔗 Note: if you don’t see links in this email, it’s also available on the web: dataelixir.com/archive/issue-588.html |
| | |
| | There’s a lot more here than the title suggests. This analysis of nearly 89,000 job posts from Hacker News tracks 14 years of changes in remote work, programming languages, data and ML roles, salaries, and seniority. Includes links to the notebook and data. MLJAR Blog | Piotr Płoński |
| | |
| | A lot of modern data infrastructure was designed around hardware constraints that don't exist anymore. In this post, Andy Warfield uses DuckDB to make the case that increasingly powerful single machines are changing where analytics happens and when distributed systems are actually needed. All Things Distributed | Andy Warfield |
| | |
| | AI can generate a good-looking dashboard in minutes, but that doesn’t mean it knows how to design one. In this post, Adam Kucharski walks through a vibe-coded dashboard to show what’s actually needed to make data easier to understand. Adam Kucharski |
| | |
| | Linux executables are packed with structured metadata, but you typically need specialized tools to inspect it. In this post, Farid Zakaria asks what happens if you model all of that as relational data instead. Then he builds an executable format backed by SQLite, which makes its internal structure queryable with SQL. Clever. Farid Zakaria |
| | |
| | Generic coding agents understand code, but not necessarily the data project around it. dbt Wizard CLI brings AI into the terminal with context about lineage, tests, metrics, and dependencies, so it can make and validate changes with a better understanding of how everything fits together. Learn more -> // sponsored |
| | |
| | | | Polars 2.0 is coming soon, and it’s less about new features than better defaults. Lazy queries move to the streaming engine automatically, while stricter type handling and error behavior are designed to catch problems earlier. Polars Blog | Ritchie Vink |
| | |
| | Tangle is an open-source project that brings a visual editor to ML pipelines without hiding the underlying code. It lets you connect reusable containerized components, run experiments, inspect results, and cache intermediate steps so you can iterate without rerunning the entire workflow. TangleML |
| | |
| | | Eleven years ago, the authors of the Dataflow model helped popularize streaming data concepts like windows, triggers, and watermarks. Looking back now, they argue that many of those ideas exposed too much machinery to users. This updated view looks a lot more like databases and incremental views. Tyler Akidau, et al. |
| | |
| | Flash Papers is a growing collection of 100+ research papers that have been turned into interactive explainers. Topics include things like attention, RAG, quantization, scaling laws, diffusion, RLHF, LoRA, and more. Flash Papers |
| | |
| | | Woxi is an attempt to reimplement the Wolfram Language in Rust, including symbolic math, graphics, notebooks, and a lot of the surrounding ecosystem. It's open-source and it runs in a browser too, so you can poke around without installing anything. Woxi |
| | |
| | | This week's issue sponsored by: | | |
|