Showing posts with label Indexing. Show all posts
Showing posts with label Indexing. Show all posts

10 April 2015

Indices: Explore all Options

It's one o'clock in the morning. Do you know who else has indexed your data set?

Record collections are coming online with astonishing frequency. Many times we sigh with resignation as company B follows company A, which published the same collection a year or so previously. Later, company C provides the same  online collection. Aside from competition for customers, what's the point? The point is, researchers can benefit from independently indexed collections.

from Wikimediacommons
There are several reasons for searching the same or similar collections on more than one website. Different companies may have:
  • their own proprietary image enhancement technology that may significantly improve viewing of otherwise identical images (compare, for example, enhanced 1940 U.S. census images on a variety of websites);
  • advanced search options and tools allowing one to focus one's search energies;
  • a variety of methods for moving within datasets to browse for images of interest (for example, I much prefer to browse manifests on sites that allow me to jump around among the images, not force me to advance only one image at a time; I also like to be able to rotate census images so I can read through the street names quickly when I am searching through an enumeration district); 
But none of this is important if you cannot find the image or have no idea where to look within the collection. 

Indexers often get a bad rap. Yes, indexing is a prime area where errors may be introduced. Ancestry has been criticized for using foreign indexers. FamilySearch has been criticized for not allowing input for corrections to their indices. A common complaint from researchers whose ancestors came from Eastern Europe is that Ellis Island manifests have been indexed by researchers who have no familiarity with those surnames and places. Because of that, some of the indexing errors seem bizarre to those of us who have some familiarity.

The truth is, despite these issues, indexing is the heart and soul of genealogy research. Yes the computer has revolutionized research. But, if not for indexing, much of that would not have been possible.

We all have our favorite websites and search tools, but when all else fails in your search, bail and go to another site that has the same collection, indexed via a different set of indexers.[1]

I will admit, that I have been an ardent follower of the freely accessible Italian Genealogy Group and German Genealogy Group indices of New York City vital records. These entities partnered in a volunteer effort to index records available through the New York City Municipal Archives. I would sometimes use the Steve Morse One-Step search forms to access them. But, even after Ancestry put indices of these same records online, I still saw most benefit in staying with the ItalianGen/GermanGen indices.

Unfortunately, one thing we do not see often enough with complex record sets is independently derived indices that include different/additional information. I have lamented previously that I did not see much value added when FamilySearch decided to initiate their own indexing project for New York manifests. I wanted to see additional information indexed, such place of birth, closest relative in the old country and address and/or name of the person the immigrant planned to meet at their destination.

New York City vital records? Enter FamilySearch. 

On 20 March 2015 they added [2]:
As I usually do, I immediately tested the new indices with one of my unusual family surnames: Liebross. I figured after all these years I'd pretty much exhausted the Liebross vital records collections in New York City. But, FamilySearch added an element to their indexing of death records that had not been included by previous indexers: parents' names (parents' names are also included in birth and marriage record searches).

ItalianGen/GermanGen and Ancestry indices have coded first name, age at death, date, certificate number and county. Results in the marriage index also provide easy access to spouse names.


Using the Steve Morse search form one can also get results that include the FHL microfilm number. One would not see parents' names until one had acquired the original record. Where there were several people indexed with the same or similar names, this created a bit of a crap shoot. Many of us have ordered records we thought might be correct only to discover that the parents names were not. 

The new FamilySearch index not only adds a bit more certainty to the process of record acquisition, but also to the hope of finding new records.

Results of my recent Liebross search surprised me. Early in my family history research I'd found that my great aunt Rose who, I thought, had never married, had indeed married (in 1926) and later divorced (in 1931) a dentist named Nathan J. Bernstein. I'd located the marriage certificate indexed on ItalianGen.[3]

In FamilySearch's new NYC Municipal Deaths index database, I waded through the expected indexed records and then, towards the end, noted records where Liebross was the deceased's mother's surname.

Oh, my! Rose Liebross Bernstein had a baby who had died: Ira Howard Bernstein.
He was born and died between census enumerations. It is unlikely I ever would have found him. While I have since found him in the same cemetery (Mt. Lebanon) as the rest of the family, he and his parents were not buried in the same plots.  

[It is interesting to note that when I searched on Liebross in the FamilySearch marriage index I received no hits. Several Liebross family members are identified in the ItalianGen marriage index search results. This is another example of why one should use more than one index in one's searches.]

Of course indexers always make choices. While FamilySearch has included much more information that previous indexers, cause of death is not indexed; nor identification of the informant, the doctor, funeral home, etc. I have no particular criticism of that.  

The results provide more information than may be searched from FamilySearch's search box. It would be nice if the search box allowed for searching on particular dates (or even months) of death and birth. Right now one may only search on a range of years. Why not allow searches on particular addresses (the smallest geographical unit one may now search on is city)?

Regardless, this new index is huge. I have already found my great aunt's previously unknown child. By searching on family surnames, one may be able to find death records for women whose married names were previously unknown. 

If FamilySearch allowed more specific time or area searches, one one might be able to conduct research into community deaths in one small area of the city. Think of the context one might develop for understanding one's family and their lives at particular times and places.

I now await delivery of a copy of the original death record for Ira Howard Bernstein from the New York City Municipal Archives.

Let's hope new indices keep coming from a variety of sources. As researchers we must try new indices to expand our opportunities for success. 

Notes:
1. It would be nice if companies were up-front about how their collections were indexed? In addition to collection descriptions on websites, they should include how indices were derived. That way one might be able to tell if indices on different websites were independently developed or copied from another (also accessible) source.
2. The data sets do not actually contain death records from 1949 or marriage records from 1938. The records end the year before those designations. It would be good if FamilySearch corrected that so researchers do not think they might find death records from 1949 and marriage certificates from 1938.
3. Queens County, New York, marriage certificate no. 3319 (1926), Nathan Judas Bernstein and Rose Liebross, 14 November 1926; Municipal Archives, New York.

09 December 2013

We are what we index

I'm amazed. Why decide to do another index of a database and provide the same information that everyone else provides? I know I'm late to the show - particularly since FamilySearch has been indexing passenger manifests for quite a while now - but why didn't FamilySearch have their volunteer indexers include some of the most valuable pieces of information on the manifest? Specifically, the names and addresses of those left behind, those to whom the immigrant is going, and place of birth. I think they've missed the mark. [1]

Don't get me wrong: I am pleased that FamilySearch chose to work on a new index for manifests. Dueling indices from several online genealogy record providers are becoming the norm and I welcome them. I have found it useful in my research to occasionally jump from one index provider to another. Different indexers may record variations on the hand-written names; some search engines provide more robust searching options; search engines may offer or use differing sources or parameters of soundex (American, Daitch-Mokotoff, Beider-Morse); some databases are easier to navigate than others; and companies vary in how well they enhance their online images (even if the originals are from the same source). In fact, when teaching beginning genealogy classes, I encourage budding researchers to go beyond one favorite genealogy site to look at the same record sets on another - especially if having trouble finding a record in one website's index.

Short of OCR or recent promises of computerized reading of handwritten records, search engines are limited by the underlying indexes. [2] With existing manifest database indices, finding information on relatives in the new country and the old is dependent upon finding records via passenger search. One may search via name or place of origin (residence) for a particular individual. Where residence may be different than birth place or location of relatives/friends in the old country, you can't get there from here (!). One must ford through individual records to find the differences.

210 Grand Street - now a Chinese restaurant
Research would be greatly eased with  additional searchable information. For example, I am interested in community emigration. I would love to find out how many people (and who) from the shtetl of Labun wound up in the early twentieth century at 210 Grand Street, Manhattan. I have several relatives who did. I have also found a few other, perhaps unrelated, landsman who did, as well. Many seem to have become glaziers in New York City. Were my Malzmann family members the center of this gang of glaziers? Or, was this more a town-based enterprise? I can now search on the town of residence and see from the manifests where people are heading. But were there people heading to that address who were not from that town? Right now, no way to check.

If I could search by address in the USA, knowing the addresses of relatives who'd already made the journey, I might be able to find some relatives for whom I have not been able to locate manifests. I might be able to locate hitherto unknown relatives. I might also be able to find more relations if I could search on those addresses in the old country. [3] This kind of information is critical for working with and via the FAN (Friends, Associates and Neighbors) principle in our genealogical research - particularly in large cities where relatives may or may not wind up living near each other.

Recognizing the utility of this additional information, the Ukraine Special Interest Group (SIG) of JewishGen started to remedy this situation by indexing records on a town-by-town basis. [4] An indexer must first search the Ellis Island database (and others) for people who resided before emigration in the indexer's community of interest. Then they must sort through to make sure they indeed have the correct town; identify the Jewish people; and index the records including family's/friend's and addresses on both sides of the ocean.

There are some advantages to doing indexing in this manner. Those indexing the records have some familiarity with overseas town names and immigrant names. So, there should be fewer transcription mistakes. But, it's slow going.

The communities affiliated with individual passengers in this new Ukraine SIG indexing project are usually locations of last residence rather than birth. This is due to the fact that the underlying existing index on the Ellis Island website does not usually include town of birth. So we are likely missing a set of people. We won't be able to capture those who were born in our village of interest, but most recently resided elsewhere.

Another problem I see with the Ukraine SIG project is that if one wanted to later build on this database, filling in with records not previously indexed so that others besides town-oriented Jewish genealogists would benefit, it would be time-consuming to locate and index the previously non-indexed records. In fact, it probably would not be worth the effort to account for those records that had been picked through and indexed. Just start over. But, for right now, considering my research interests, it's the best indexing project going.

So, here's my thought: when planning a new index, find out what's really useful and go for it. Adding to the available indexed information may be the most powerful argument for participating in an otherwise parallel indexing project. I wish FamilySearch had done that.

Notes:
1. I am not indexing for FamilySearch right now. I was an enthusiastic (and productive) indexer for FamilySearch during the 1940 Census indexing juggernaut. I enjoyed the experience. They have an awesome interface and I'd do it again. I am on hiatus from indexing - working on other projects. I suppose I could be wrong, but I didn't see any evidence that FamilyFamilySearch is indexing additional information.
2. Optical Character Recognition - computerized finding aid for mining words in typed or typeset documents. Mocavo has recently announced they are getting ever closer to developing a computer program to decipher hand-written records.
3. As an example, see my series (still a work in progress) on Fannie Greenfield. My first post in that series is here.
4. I actually started doing this on my own several years ago for my Yurovshchina/Labun community website (http://www.kehilalinks.jewishgen.org/yurovshchina/index.html). However, I would be remiss to not mention that I am currently on the board of the Ukraine SIG.

25 February 2012

Saturday Night Genealogy Fun - FamilySearch Indexing

Randy Seaver, prolific blogger of Genea-Musings, suggests amusing activities for "Saturday Night Genealogy Fun" every weekend.  Today he suggested signing up and preparing for indexing the 1940 Census (set to be available after 72 years on 2 April 2012). I've never participated in his Saturday activities before but, I'd been thinking it was time to prepare for learning how to index on FamilySearch.  There's no time like the present!

FamilySearch Indexing: Piece of Cake!
I signed on, downloaded the software to my Mac, watched the video and set to it. After indexing three records from the "US, Arkansas - Second Registration Draft Cards" record set, I tackled the test for 40 records from the "US - 1940 Federal Census." Piece of cake [hmmm time to take a few minutes off to bake some brownies]!

[After the brownies were in the oven] I completed 10 more records from the "US, Texas - Deaths, 1890-1976." Then, finished off with a new challenge to me (since I never have to look at UK census records in my own research) and selected "UK, England and Wales - 1871 Census," 25 records [brownies are done]

I've done a bunch of indexing for the Italian Genealogy Group, which specializes in New York City records. For that, my indexing skills only required basic facility with Microsoft Excel spreadsheets (as well as eye-crossing copies of photocopied records). I found the FamilySearch indexing program easy to use and master. The Quality Checker (which kicks in after one has finished one's batch) and pull down menus are slick features and definitely helped when I was working on UK records and wasn't entirely familiar with some of the town names.

In all (including the US Census test) I completed 78 names and garnered 98 points (whatever) [and one brownie and a glass of milk.]

Thank you, Randy, for getting me moving on indexing [bon appetit!].

-------------------
My next challenge is to get ready for the 2 April unveiling of the unindexed 1940 Federal Census. Estimates are that it will take about six months before we volunteers have finished indexing the Census records.  During that time and during the hours I reserve for my own research, I plan to locate as many relatives as possible in the Census. To do that without an index will require making a good calculation, based on my prior research, of where each family was residing in about 1940 and applying the tools developed and posted on SteveMorse.org to determine the 1940 Federal Census Enumeration Districts (ED). I'll be posting about the 1940 address/ED database I am building for my relatives.

The URL for this post is: http://extrayad.blogspot.com/2012/02/saturday-night-genealogy-fun.html