Metadata

OpenRefine: API Basics and Repeat Operations

Even as an experienced OpenRefine user, I didn't know much about the API. I began investigating it when I needed to repeat operations across a large set of spreadsheets. While I found an alternative that's simpler for now, I wanted to share what I learned in case it's of interest to others.

Open Refine: Repeating Groups of Operations (the Easy Way)

Even though I've been a regular OpenRefine user for ages, I had never paid attention to a very important pair of buttons in the list of operations performed. As it turns out, it's extremely easy to perform a whole set of operations once and repeat it on other projects with the same data structure.

Open Refine: Blanking Down Only Within Records

Blanking down is a great way to remove repeating rows within an OpenRefine record. But if the field's contents are the same between records, a single blank down action can wipe out rows of data you wanted to retain once in each record. You can solve this using row.record.index to add the record number to rows beforehand. This writeup contains the two transformations you'll need and a brief walkthrough on how to use them.

Deeper Dive into Estimating BTAA Sociology Serials Holdings with WMS APIs, Z39.50, and Spreadsheets

A deep dive into the methods I used to estimate institutional holdings for over 5000 journals in the Big Ten. I walk through steps from ensuring I had ISSNs for both print and electronic versions of a resource to the multiple methods I used for querying.

Apple Health's Weird Workout Assumptions

I connected my Strong app for powerlifting to Apple Health and learned all kinds of weird things about Apple Health's design and assumptions.

Regex Resources for Catalogers and Others in the Library

After last week's talk on making use of the period between now and fundamental shifts in cataloging, regular expressions (regex) came up during the Q&A as one of the things I think it's helpful for catalogers to learn. Afterward, one of the attendees emailed me asking for more specific recommendations on how to learn them. I'm sharing the resources I pulled together, since I think they could be of use to more folks.

MARC Misconceptions: Which Language of Sound Track? Use of the 041$a

The MARC 041$a may initially seem like an important field to index for catalog faceting or collection assessment. However, its use for handling optional dubs in DVDs, Blu-rays, and streaming media may confuse patrons if we index it at face value. It may also contribute to incorrect assessments regarding the languages of material held in our collections. This post provides further context and recommendations for assessing language of the media in our catalog.

MARC Misconceptions: Incorrect ISSN or the 022$y

The MARC field for Incorrect ISSN has ended up taking on two very different functions. This post explores both functions and their history a little and explains why they've made automated processing of the field completely impossible.

Weighing Fields in Library Catalog Search, or, The Hillbilly Elegy Problem

An overview of the challenges in assigning field weights in a library catalog. This uses the example of a book about Hillbilly Elegy to explain demonstrate where challenges arise.

Repository Ouroboros

If you are somewhere on the repository ouroboros, you are not alone, you are not broken, and you are not hopelessly behind.