| | Become an expert in data storytelling with the online Master of Translational Data Analytics from Ohio State’s Translational Data Analytics Institute. This interdisciplinary program blends machine learning, user experience, data visualization, and design thinking to prepare you to uncover and present insights from data. *Message from this week's sponsor, The Ohio State University |
| | |
| | 🔗 Note: if you don’t see links in this email, it’s also available on the web: dataelixir.com/archive/issue-589.html |
| | |
| | Robin Linacre has been using AI agents for everything from code migration to research experiments and has learned that the hard part of doing data science is increasingly not writing code, but designing the tests, benchmarks, and feedback loops that let the agent determine if its work is actually good. Robin Linacre |
| | |
| | This is less an argument that Pandas is bad and more an argument that the options have changed. With tools like Polars and DuckDB, workloads that previously pushed analysts toward distributed systems can often run comfortably on one machine. Eddie Atkinson |
| | |
| | NumPy makes array operations look simple, but there’s a lot happening under the hood. This post follows np.add() through NumPy’s internals and shows how things like strides, memory layout, and SIMD can have a big effect on performance. Veit Heller |
| | |
| | Data Colada is back with another dataset that doesn’t behave the way real data should. This time, they start by trying to reconstruct students’ final grades from component scores, find a strange mismatch, and follow it through a series of increasingly suspicious patterns. Data Colada | Uri, Joe, & Leif |
| | |
| | This is a useful case study in data-quality tradeoffs. Pew Research tested several ways to filter bogus respondents from online polls and found that better detection doesn’t necessarily produce better results. False positives matter too, and some screening rules distorted the sample more than the bad responses did. Pew Research Center | Andrew Mercer |
| | |
| | | | This is a fun collection of ideas about what makes data visualization work. In this post, Duncan Bradley starts with quotes from well-known visualization practitioners, then digs into the thinking behind them, from “clarify, don’t simplify” to why annotation and context matter as much as the chart itself. Graph Paper | Duncan Bradley |
| | |
| | | Knap is an open-source templating language for generating Markdown from structured data. It supports JSON and CSV input, variables, filters, loops, conditionals, batch generation, and both CLI and JavaScript APIs. Steph Ango |
| | |
| | Mnemiq is an open-source framework for evaluating and improving text-to-SQL on your own database. This launch post compares 28 configurations across models, retrieval strategies, semantic enrichment, verification, and decoding, with some big differences depending on the dataset. Agentic Fabriq |
| | |
| | | A team from Oruk took 499 neurons from a fruit fly's wiring diagram, copied their connections into software, and fed the network recordings of human voices. A small output layer learned to predict the emotion labels people had given those recordings, which worked, but what's really interesting is what happened when they scrambled the wiring. Nathan Roll |
| | |
| | | This week's issue sponsored by: | | |
|