Friday, July 13, 2012
2012 Project for name extraction
I have just finished assisting my society in getting out an issue of our magazine following the retirement of our editor of some 15 years. I mainly worked on the page layout of the articles and so had to polish my skills in using MS Word to properly space and align photos to avoid large blocks of unnecessary "white space".
I really don't know how the previous editor managed the indexing, but I was committed to using the FDEX compiler. The problem was to automate the extraction list from some 60 pages of content. Although the "non-contiguous selection" method works well (described on the web site) I decided to write a name extractor that would give me a rough list of all terms in which consecutive words all begin with upper case letters. In other words, in the sentence "Joe Doe took Sue Smith to lunch, the list would display:
Joe Doe
Sue Smith
At this time the extractor does not distinguish between upper case words used as a sentence first-word. This can be resolved by logic that looks at the sentence first-word, compares it to a list of common first-words (like 'The', 'Then', 'A', etc.) and ignores extracting them.
The project title for this extractor is FDEX2, but it is in the early stages of development and not yet posted to the web site. However, I will soon integrate it with a link from the "Tools" page and hope you will leave feedback for suggestions and reports of your experience in using it.
Please register to receive future notices on news of this developing project
Friday, April 23, 2010
Questions? Comments?
I welcome questions and suggestions on Format-A-Dex in general and the FDEX compiler and other tools specifically. Leave your Q & S as a comment to this post so others may read, benefit and respond.
Note: Comments are moderated to avoid spam-bots.
Note: Comments are moderated to avoid spam-bots.
Is it free?
Some early testers have asked if Format-A-Dex is a free service. At this time it is. Obviously the more who use it, the more feedback I get and the more stable and robust it becomes. My first consideration is that it be used by non-profit societies, particularly those who support publishing for the genealogical and historical interests communities. If support for it gets overwhelming, or it appears to be used by too many for commercial purposes, I will re-visit this policy.
Announcing Format-A-Dex
Format-A-Dex is a set of tools and guides to simplify and streamline the process of indexing a journal, newsletter, society magazine or any other document that is fact-filled with names, places and events. Because it was designed by a writer/editor involved in genealogy, Format-A-Dex is focused on the needs of editors in genealogical and historical societies.
At the site there is an on-line index formatting tool, a name reverser tool, and a tutorial on how to extract items from an MS-Word document for "Type Less" indexing. There are examples as well as links to other relevant web sites. You will also find a link to this blog for updated information on this project.
At the site there is an on-line index formatting tool, a name reverser tool, and a tutorial on how to extract items from an MS-Word document for "Type Less" indexing. There are examples as well as links to other relevant web sites. You will also find a link to this blog for updated information on this project.
***
Subscribe to:
Posts (Atom)