Showing posts with label Rules Engine. Show all posts
Showing posts with label Rules Engine. Show all posts

Friday, November 27, 2009

An issue of Context…

When creating a rule/logic/workflow engine there are various things that need to be considered in it’s architecture.

One of the most overlooked aspects of such systems is that of the CONTEXT in which a rule is run, and how it affects what a rule can do.

To elaborate a little more, rules normally know nothing of the environment that they run in, rules just execute logic based upon data that should be available when the rule is evaluated for its result. (a result is normally a logical true/false, but it can be other values and data types depending on its usage and context).

So what do I mean by Context, quite simple the context for an executing rule is the environment that it runs in.

A context normally defines what data is available to a rule for example:

    • Global Values
    • Results of nested logic that is used in the rule
    • System data that is applicable to the rule being evaluated
    • Functions or other rules that can be called by the rule

A rules engine when it is instantiated/run should go through the following processing steps:

  • Setup context for the rule (normally defined by a flag or driven by a workflow system)
  • Execute a rule or batch of rules in context
  • Tear down context (sometimes a context is persisted depending on the type of rule logic and system it is in)

Later on, I will be creating a simple logic/rule engine that conforms to the above and deals with the issues in creating and synchronizing a context and its data.

I hope this has been at least a little helpful, soon the real fun and more technical stuff will be started.

Thursday, November 26, 2009

Example Logic for a simple Rule Creation UI

Below is an example of a simple flow charts for  a couple of standard pieces of logic.(yes this rule structure will be in the API I am thinking of creating, and it is used in both a workflow API and also a simple rule engine/expert system.)

The format displayed below is common and is found in quite a few logic building systems, for example windows workflow foundation.

Copy from a single source field to a target field

image

Copy from two source fields and combine to a target field

image 

I have not shown the exact logic/process of how this rule could be created/stored/implemented.

That can wait till I cover a simple architecture that can be used to create a basic rule engine and then produce a simple code based API that will execute simple rules like the one demonstrated.

One thing I will say, is the logic structure and all you see is meta-type and meta-data driven. the meta-type/data will then be converted into a format that can be later interpreted by the API to process the logic in question.

Sunday, November 22, 2009

My First Big Data Integration project in the early 1990’s

I am not sure how many people read this blog? However I have been reminiscing on some of my favorite data integration Projects/products from my past.  (I will outline this project and the type of work involved later., it taught me a lot about data integration from multiple related sources into a single normalized master version of the data, and then using that normalized view to output data in any shape required!)

I worked on this product very early in my professional career; The product basically allowed you to take data from any mainstream project management system, store it in a common normalized format, (think of a high level logical model of project management entities and relationships plus logic for common project management tasks like rolling-up or down project metrics and WBS) and then pushing out data after some processing, integration and normalization to any project system format that we supported.

Just for the record some of the products we supported at the time Microsoft Project, Primavera and Microsoft Excel to name but a few.

A key selling point at the time was the ability to take changes to a project plan in any of the major project management tools on the market, normalize them into a common format, do some project related processing to the plans, and then push out a normalized view of the project data (a single master copy if will) to end user system formats. Then once changes were made once more they were updated in the normalized logical model and then publish back out to supported systems once again.

It was pretty impressive for a good 12+ years ago!

Friday, November 20, 2009

A history of Knowledge…

I was asked yesterday, just how much Data Warehousing and indeed BI knowledge I have. And how long have I worked on issues relating to workflow systems, logic systems, expert systems, compilers, interpreters and data consolidation technologies and products. So here goes a brief overview.

I started out in IT a good 20 years ago, and during that time I must admit I have seen a great many technologies.

I was fortunate in that I was introduced to large and very large data (and sets of data) when I worked as a Principle consultant for a market leading Workflow and Document Management solutions provider in the 90’s during this time I worked on what was at the time regarded as one of the leading workflow systems used in some banks for document management and also business process.

Remember that we are talking 1990’s here, this is before the advent of tools like WWF (Windows Workflow Foundation), what was also remarkable about that system which I also helped to architect solutions with and work with its code base; Is the fact that it worked natively in Java and also on the Microsoft platforms also. It was a ‘C’ based core API, with extended API’s in Java and at the time Visual Basic of all things.

I will say this however, it opened my eyes to logic/rules engines in a very big way and started my learning process on how to write:

  • complex text parsing/scanning systems
  • Interpreters (script engines and logic engines)
  • workflow systems (automated, and exception driven)
  • data integration from multiple sources into a single master source/copy of data.

It was also while working at this company, I helped a large group integrate many data sources of information we are talking between 50-100 separate and disparate sets of data into a large normalized data structure and sub-structures that had a single master copy of data, audited data related to how the master copy was created and references to all source’s used in that master. ( I had emails, data feeds, satellite feeds, and direct entry to match/merge for a specific business driven purpose). I also worked on implementing a full workflow system for a national postal service which used our workflow/rule engine API’s/UI’s to manage its postal sorting/routing and delivery –> an immense amount of data and processing that most people take for granted, it was a very humbling process)

At the time it caused me no end of pain to figure out with the teams I was working with, but in the end I learnt the true power of how integration of multiple-sources into one could be done, and indeed how a real workflow system and rules engine can be applied.

Ironically, a good 12 years later I am still working in the same area of technology, I have implemented interpreters/compilers and rule engines in many places, as-well as parsing engines, integration engines and workflow systems. (And far to many metric based systems related to data imported/transformed or used over time…)

I guess that’s why I am now documenting some of the more basic and publically common architectures for doing these kind of systems now.

I do admit that personally, I really enjoy the challenges in these systems, especially combining interpreter architecture with workflow, rules engines and parsing of data. and this blog is based on a new journey that I am making as I look further into these area’s.

It is also nice to see that the industry as a whole in the past 3 years has moved to adoptions of DLR’s, Workflows and business rule engines for most aspects of software delivery.

One of the more notable data warehousing systems I have a lot of understanding and respect for is Oracle’s 10g,11g OWB product (data warehouse builder), it has in my opinion one of the most flexible matching and single master copy of data processing/rule engine systems I have ever encountered.

In-fact the match/merge API I am thinking of creating here, is driven a lot by what I have seen in that product and many other referenced products in the EDM/MDM/ETL/E-TL and BI market place.

All of the techniques and process I am showing here is actually pretty old techniques (quite a lot originated in the golden computing era of the 70’s), I am just using more up to date technology.

If you Google on any of these technologies and techniques, you will find many solutions to these problems, almost all are now public domain and food for thought for those creating their own systems

Wednesday, November 18, 2009

Scanning and Parsing…

Over the past 4 years I have been looking at interpreter and compiler creation specifically in the area of Scanning and Parsing.

One of the designs architectures that I designed some years ago was based on parallel processing of a stream of data (a file, and communications port/channel, etc.).

It is just this design that I aim to bring to life in the API’s I am creating over the next god knows how many months.

The concept is simple, but like many concepts there is a lot more to it that the text here implies; However, I do believe I have solved most of the architectural and performance based issues. This is due to both some rather large processing advances and superior development tools now on the market.

the concept itself is based on a pointer based parsing system, that uses various conditional logic to step forwards, backwards and indeed around various parts of a stream as it is being processed.

Traditionally this kind of Scanning/Parsing system is notoriously slow. but by looking at certain interpreter/compiler creation and performance optimization tricks, I think it may work at an acceptable performance level for real world systems… Guess I will find out in time.

Funnel Metaphor Matching and Merging...

As I am designing/architecting a matching and merging API albeit a very basic one, I am leaning towards a Funnel like metaphor of how data is matched and merged into logical sets of normalized data.

The metaphor which for various reasons I will not put a diagram online for, is very similar to that of a sales pipeline as an idea of reference. (I have not been able to create a diagram that I am happy with, and also since the API is not created and only in early design. the Diagram is likely going to be wrong. I am not even sure this way will work.. guess I will know soon enough…)

As a sideline I am working on maybe naming this technology API/concept, something along the lines of the metaphor concept of funnel; But alas I am not getting any good ideas…

Maybe someone in the Blogosphere has a good idea?

Rules are thus:

  • It is an API used to Merge-Match data.
  • It is designed for developers.
  • It has no UI.

Please post your idea’s and keep them polite and clean ^^

Importance of Data Cleansing prior to matching…

Yet another key aspect in relation to matching is that of cleansing data. By cleansing I mean normalizing/correcting data as much as possible. Where possible you want to have reliable data in the data sources being matched.

Should this not be the case, all matches are potentially incorrect.

yes it seems obvious, but many seem not take this into account even when doing simple matching processing for duplicate data in a single source, let alone across multiple data sources.

As part of the merge-match blog entries I am making I will also elaborate on various cleansing and data quality checks that could and can be made. there are a great many techniques here, most are from hardcore ETL systems and Data warehousing systems.

And as always a lot of reference data and indeed systems can be found online.

Definitions of Logic Types that can be used in Matching…

Since I am delving quite deeply into Merge-Matching techniques it makes sense that I also define the 3 main types of Logical processes that can be used to decide logically what could or could not be a potential match of data source elements.

Plus its good for me to re-confirm my own understanding, I am open to other logic types that could be used in advanced matching as long as they are not too esoteric in nature.

Here list in order of my preference:

  • Deterministic Logic - Deterministic Logic gives an equal weighting to the different types of information available and uses the overall weighting values as a way to make a logical decision. For example a decision is determined based on a real known value/score.
  • Probabilistic Logic - Probabilistic Logic normally uses some statistical information that has been collected to make a decision. For Example it is more likely that this is the right decision based upon experience.
  • Fuzzy Logic - Fuzzy logic is meant to resemble basic human reasoning and its use of partial or approximate information weighted against uncertainty to come to a logical decision. For example, its sounds like that item so maybe it is that item.

Merge-Match processing – Matching Rule breakdown…

As I stated in a previous blog entry, I would document all of the matching rule and logic structures that I know of and have encountered in the past 20 years or so of my career. Most of these terms and techniques are common to data warehousing. In-fact I have made this list a lot shorter than I was planning and have just kept to the basics for now. Yes, I expect to implement all of the basics below in the API I am creating.

the Structure is simple, all logic used to define a match is based on an archetypical logic structure, either simple logic or conditional. Then you have  the scoring and comparison scoring mechanisms that are used to match disparate data into logic matched items. I am not documenting here how these are done or implemented in code just yet.

Simple Matching Rules

  • Match All             Matches all rows within a match group
  • Match None         Turns off matching.

Conditional Matching Rules

  • Conditional match rules specify the conditions under which records match.
  • A conditional matching rule allows you to combine multiple attribute comparisons. When more than one attribute is involved in a rule, two records are considered to be a match only if all comparisons are potential match’s.

Comparison Scoring that can be used to indicate a match

Similarity Scoring 

  • A minimum similarity score required for two pieces of data which is used as the basis of a potential match. For example: A value of 10 indicates an exact match, and a value of 0 indicates a non-match. There are many ways of calculating this and as always there are great examples and ideas on the best measures all over the web.

Blank Scoring

  • A way of dealing with empty/null values in matching, this may involve automatically treating them as valid matches or as invalid ones. -  also a number value maybe given if part of an elaborate scoring system.

Comparison Scoring

  • Each attribute in a conditional match rule is assigned a comparison algorithm, which specifies how the attribute values are compared. Multiple attributes may be compared in one rule with a separate comparison algorithm selected for each. The real power of complex matching is in the Conditional rule aspect of matching.
  • Example types of Comparison could be:
    • Exact Match – The most common and easiest. Attributes match if their values are exactly the same. For example, "Dog" and "dog!" would not match, because the second string is not capitalized and contains an extra character. Some systems hash this value others just mark up a score in a related table. I aim to support both methods. This type of matching is valid for all data types. 
    • Soundex Comparison - Converts the data to a Soundex representation and then compares the text strings. If the Soundex representations match, then the two attribute values are a potential match. Not a very good way to do things to be honest, and pretty useless in any language other than English.
    • Abbreviation/Acronym Comparison – Quite simply a lookup, these are very domain and data source specific so care should be used. For example, "International Business Machines" would match "IBM", “Management” would match “Mgmt.”, and so on...