Showing posts with label Domain-Specific Language. Show all posts
Showing posts with label Domain-Specific Language. Show all posts

Monday, November 30, 2009

DSL 101 – Domain Specific Languages

Today, I spent quite a while helping a friend to architect at high-level a DSL (Domain Specific Language).

Even though at the time he did not realize what he wanted to do was create a DSL. TO him all he needed was a way to do excel like formula expressions inside his code, but in a more English like syntax well as English as a chemical scientist can be I guess.

To this end, I thought it might be worthwhile explaining what a DSR is, since I will build a few of them for the project’s I mentioned that I am playing with for fun on this blog as I learn F#.

So just what is a DSL:

A domain-specific language is a programming/specification language customized to a particular problem domain.

This is quite an old concept as special purpose languages for modeling have always existed but recently have had a surge of emergence once again due to most newly created systems needing a more flexible way to extend logic used without the need of a programmer.

Creating a DSL and the tools required to support it can be worthwhile if the language allows a particular type of problem to be expressed more clearly than existing languages would allow. An example of this is the issue of applying simple logic used to route a document for approval in a document management system.

Another example of the new generation of DSL’s are products/tools like cucumber (from the Ruby world to do BDD), and in the F#/C# world we have tools that are DSL' based such as NaturalSpec (also a BDD tool this time for F#).

Sunday, November 22, 2009

My First Big Data Integration project in the early 1990’s

I am not sure how many people read this blog? However I have been reminiscing on some of my favorite data integration Projects/products from my past.  (I will outline this project and the type of work involved later., it taught me a lot about data integration from multiple related sources into a single normalized master version of the data, and then using that normalized view to output data in any shape required!)

I worked on this product very early in my professional career; The product basically allowed you to take data from any mainstream project management system, store it in a common normalized format, (think of a high level logical model of project management entities and relationships plus logic for common project management tasks like rolling-up or down project metrics and WBS) and then pushing out data after some processing, integration and normalization to any project system format that we supported.

Just for the record some of the products we supported at the time Microsoft Project, Primavera and Microsoft Excel to name but a few.

A key selling point at the time was the ability to take changes to a project plan in any of the major project management tools on the market, normalize them into a common format, do some project related processing to the plans, and then push out a normalized view of the project data (a single master copy if will) to end user system formats. Then once changes were made once more they were updated in the normalized logical model and then publish back out to supported systems once again.

It was pretty impressive for a good 12+ years ago!

Wednesday, November 18, 2009

Merge-Match processing – Matching Rule breakdown…

As I stated in a previous blog entry, I would document all of the matching rule and logic structures that I know of and have encountered in the past 20 years or so of my career. Most of these terms and techniques are common to data warehousing. In-fact I have made this list a lot shorter than I was planning and have just kept to the basics for now. Yes, I expect to implement all of the basics below in the API I am creating.

the Structure is simple, all logic used to define a match is based on an archetypical logic structure, either simple logic or conditional. Then you have  the scoring and comparison scoring mechanisms that are used to match disparate data into logic matched items. I am not documenting here how these are done or implemented in code just yet.

Simple Matching Rules

  • Match All             Matches all rows within a match group
  • Match None         Turns off matching.

Conditional Matching Rules

  • Conditional match rules specify the conditions under which records match.
  • A conditional matching rule allows you to combine multiple attribute comparisons. When more than one attribute is involved in a rule, two records are considered to be a match only if all comparisons are potential match’s.

Comparison Scoring that can be used to indicate a match

Similarity Scoring 

  • A minimum similarity score required for two pieces of data which is used as the basis of a potential match. For example: A value of 10 indicates an exact match, and a value of 0 indicates a non-match. There are many ways of calculating this and as always there are great examples and ideas on the best measures all over the web.

Blank Scoring

  • A way of dealing with empty/null values in matching, this may involve automatically treating them as valid matches or as invalid ones. -  also a number value maybe given if part of an elaborate scoring system.

Comparison Scoring

  • Each attribute in a conditional match rule is assigned a comparison algorithm, which specifies how the attribute values are compared. Multiple attributes may be compared in one rule with a separate comparison algorithm selected for each. The real power of complex matching is in the Conditional rule aspect of matching.
  • Example types of Comparison could be:
    • Exact Match – The most common and easiest. Attributes match if their values are exactly the same. For example, "Dog" and "dog!" would not match, because the second string is not capitalized and contains an extra character. Some systems hash this value others just mark up a score in a related table. I aim to support both methods. This type of matching is valid for all data types. 
    • Soundex Comparison - Converts the data to a Soundex representation and then compares the text strings. If the Soundex representations match, then the two attribute values are a potential match. Not a very good way to do things to be honest, and pretty useless in any language other than English.
    • Abbreviation/Acronym Comparison – Quite simply a lookup, these are very domain and data source specific so care should be used. For example, "International Business Machines" would match "IBM", “Management” would match “Mgmt.”, and so on...