Friday, January 09, 2009

Prototype, Proof-of-Concept, and Pilot. Oh my!

These terms are often used in IT contexts often without much consideration to nuances. Whatever you call it, the first important thing is that you understand and state your objective. The second is that you meet it.

My suggested definitions.

Pilot:
This is an implementation of a system that is often functional complete. It is typically deployed in production, but usually constrained to a small number of users. Although we hope everything is perfect, there is an expectation that there will be faults that require rework - otherwise we could have just gone full production. The fault maybe in deployment, code, design, or in user experience. A pilot is typically time-boxed. A pilot is usually fully productionalized from an operational perspective.

Beta
: Similar in many aspects to Pilot. There is a lesser expectation that it is functionally complete, but it typically is. With a Beta, there is a much more explicit understanding that it is not final. It will change in the final release. A Beta is often supported, if at all, by a different organization than a production instance, typically developers. Traditionally a beta was not to be used for production, however some companies are making it part of their normal process - the never ending beta. Something that Google has done many times.

Release Candidate:
A software build and might be view as living in the middle between Beta and Pilot. A release candidate may be promoted to pilot or production.

Proof of Concept:
I view this as a very narrow and well defined activity. There is a well defined concept. The objective of the activity is to prove that the concept is viable in some aspect. Functionally it is only complete enough to meet the objective. The resulting code is not intended to be used for anything else - although most programmers will harvest some aspects for other things. Agile methods talk of 'early pain'. Significant projects has technology aspects that well be challenging and be a source of risk. It is desirable to execute on those aspects first, if there is a problem you want to know about it early so that you can change your plans or maybe cut your losses. This is what a proof of concept is about; if you are going to fail, fail early. Many times I have seen people propose a 'proof of concept', with no idea what concept they wish to prove; often what they mean is that they want to start coding. PoCs are not implemented in production.

Prototype
: Perhaps this one has the most varied definitions: Experimental Prototype, Engineering Prototype, etc. In software development, a prototype is a rudimentary working model of a product or information system, usually built for demonstration purposes or as part of the development process. As part of an SDLC approach, a simple version of the system is built, tested, and then reworked iteratively until ready for use. Prototypes are not usually implemented in production. Go read the wikipedia article http://en.wikipedia.org/wiki/Prototype

Do you agree? Are there any characteristics of any of these terms that you think would help define them.

Monday, January 05, 2009

Unstructured, Semi-Structured, and Structured Data

I originally wrote this over two years ago and have intended to post it ever since. Not too late.

Often in the IT world we hear or even use these terms. But what do they really mean? Here is my view.

All bits and bytes that we deal with in the IT world we considered to be data (at least). It all has some form of syntax and structure, so what do we really mean, and why is it useful to distinguish between them?

These three classifications represent a continuum which spans from unstructured to structured data that represents the degree to which the data's semantic model (meaning) matches our processing requirements. In general what we are trying to describe is the readiness of the data to be processed in a particular business context .

For example, if the data in question is the raw audio recordings from the call centre, and the business context is we need to review all verbal instructions spoken by customer "Joe Smith" last year over the phone, we may consider that recordings to be unstructured. We have no easy way to process the request.

If we have augmented those recordings with additional data from other systems and have added customer number and call timestamp to the recordings (or an index) then we would consider that to be semi-structured data. Although we could quickly sift through the millions of minutes of recordings to get Joe's subset, somebody would still have to listen to the recordings to find the things that Joe said.

The structured data, in this case, could be represented by the actual transaction records that the call centre agent created in response to Joe's instructions.

Likewise, A TIFF image might be considered structured data within the Context of a GIS application (geographic information system), but might be considered unstructured within a mortgage appraisal application. (Perhaps even GIS would consider it unstructured since they might ideally wish to run queries over an image set to find all lakes larger than a certain size. That would be hard on untagged TIFF images.)

All the data we typically deal with has a known syntax, even if that syntax is only really understood by MS Word. And although a Word document may have semantic meaning to a human, that semantic meaning is not easily extract by computer. We consider a Word document to be unstructured (in most cases).

An Excel spreadsheet may have a well defined layout of rows and columns. Although Excel may find it easy to 'understand' its content, other programs may or may not. If they layout is regular and complete, programs other than Excel maybe able to extract that data from the spreadsheet and do useful things with it. We would consider that to be semi-structured. I suggest that the 'semi' aspect of the term introduces the concept of a degree of uncertainty. Perhaps this is because it source is not well controlled and the form (layout) may change and it suddenly becomes unstructured in our context.

Structured data has an aspect of surety about it . We know that there are 'fields', we know where they are, we know what values to expect. We know how to understand it. We expect there to be some kind of formal model which defines this structure, and we expect that there will be controls in place that enforce our expectations. We may often visualise such data as being a relational form stored in a RDBMS. But that is not a requirement.

All that said, here are my definitions:

Unstructured Data: Data which does not have the appropriate semantic structure which allows for computer processing within a particular business context .

Semi-structured Data: Data which has some form of semantic structure which would allow for a degree of computer process within a particular business context, but may need some human assistance. It may apply some heuristics but the process may fail due to volatility of the structure or incorrect assumptions about the structure.

Structured Data: Data which is well positioned to be reliably processed by computer within a particular business context. It has a well-defined and rigorously controlled syntactic and semantic structure. The elements of the data have a well defined datatype and rules about valid values and ranges. The meaning of these data elements is well understood in isolation as well and their relationships to other elements. Elements are also traceable to their originating sources and that path is verifiable.

Monday, December 22, 2008

Synovus to merge Americus and Albany banks - Atlanta Business Chronicle:

Synovus to merge Americus and Albany banks - Atlanta Business Chronicle:: Something doesn't seem right here. Consider these two quotes.

"Once completed, the combined banks will have $649 million in total assets in seven locations, Columbus-based Synovus said."

"On Dec. 19, Synovus received $967.8 million in funding from the federal government’s Treasury Troubled Asset Relief Program (TARP)."

They have received more in TARP funding than they have in assets!

Wednesday, October 08, 2008

Quantum cryptography unbreakable?

Laser cracks 'unbreakable' quantum communications

I have maintained for years that there is no security scheme that is uncrackable it is just a matter of are you smart enough and perhaps have enough money.


If a vendor tries to pitch a product to me and claims it is completely secure indirectly also tells me they don't really understand security and that their people are not smart enough to even consider how to attack it and succeed.


Quantum crypto was hailed as being uncrackable. I didn't believe it even though I knew relatively little about quantum anything.


Finally remember the Maginot Line.

An opportunity to pick up either space or customers, or both!

Bank Mergers Spell Change for Branches in Manhattan

Nervous customers may wish to jump to a competitor - perhaps seeking to diversify their deposits to fit within the FDIC coverage, and empty brnach locations provide plenty of opportunities - if we can talk ourselves into the spaces.

Monday, October 06, 2008

'Discomgoogolation' - this is why I have a blackberry!

Feel stressed if you can't get online? You could have 'discomgoogolation'

I do hate to admit it, but while in Alaska and Yukon in July there were days on end when I had no internet connection. It was somewhat stressful. I guess I should do that more often.

Sunday, March 23, 2008

Thinking into the future


Be Like the Internet - 8 steps to success in a post 2.0 world


From: Thor, 10 months ago





This is v1.0 of a presentation that Lane Becker and Thor Muller are workshopping. It was delivered for the first time at WebVisions in Portland, ... less Oregon.


SlideShare Link

Saturday, March 01, 2008

Software Architect Guy ...

A whiteboard in the shower. Brilliant.

Video: ODC2008 Architect Guy

If you like that, you may also enjoy "Greg the Architect".

Tuesday, February 26, 2008

I now have a New Jersey Phone Number

Maybe not a big deal, but i find it interesting. I signed up for Grand Central - a google provided free phone service. I could tell you what the phone number is, but I don't want to publish it here, if you want to give me a call, you still can. click the button.

You will be asked for a number to call and my phone will ring.

Friday, July 06, 2007

I am neither architect nor engineer but an ecologist and gardener

This discussion has been going on elsewhere, and it resonates well with me. It is useful to remember that all four terms are analogies. The thought space about my job as being a software architect or engineer are quite old, the ecologist and gardener ideas are refreshing.

Monday, June 25, 2007

Applicability of "Google’s three rules"

Robin Harris wrote about "Google's three rules" for the data centre, in short they are:
  1. Be cheap
  2. Embrace failure
  3. Architect for scale.
Great rules and they appear to be working well for Google. Does that mean that your organization should adopt the same rules? Not necessarily. As you read through Harris' discussion a couple of key success factors jump out at me that I suggest are pre-requisites of the Google approach.

  1. "Free or home-made software": If you are paying license or support fees for the software you are running that has a huge impact on how you approach provisioning. For many Enterprise Data Centres (EDCs) I would expect that they have already discovered that the cost of the software stack can easily exceed the cost of cheap servers several times. That is very much approaching Gates' "hardware will be free" prediction. Schwarz said a similar thing as well. What remains that drives cost is the software stack, and if you want to put large multipliers in front of that cost, then it has better be a small number.
  2. "Google hired some of the best minds in the business": rather than investing in software licenses or expensive hardware, they are investing in brain power. It seems to be a reasonable approach to me? Have they figured out how to scale brain power? Certainly there are lots of untapped talent around the world that should satisfy demand for a while - how long depends on how many companies take the same approach. The more significant limiter is not supply but rather their ability to manage the bureaucracy that will evolve. But this is not about Google, but your organization. Can you acquire brain power and scale it?
I am not saying the Google's approach is wrong - but if you don't satisfy those pre-reqs and forge forward with their approach, you will have problems. Lots little and yet expensive problems called servers.

Tuesday, May 08, 2007

Event Processing

I've come upon lots of good reading lately on complex event processing. Interesting stuff indeed. Although there are some really cool things that can be done with this kind of technology, I think operational costsare a significant stumbling block for organizations like mine.

The article "Event Processing in 2007 and beyond" provides a good overview and some examples.

This site is dedicated to the topic and contains much good material.

Design patterns is a favourite topic of mine, and CEP Design Patterns provides a design pattern spin.

Wednesday, May 02, 2007

Multi-processor and Programming Languages

"now that multi-core is in rapid adoption, the issue of multi-processor programming is beginning to loom large." I wasn't the first to say it, but it is reasuring to see others say it to.

With respect to StreamIt, it is a completely different programming model than the previously reported Fortress Language.

Thursday, April 19, 2007

A legitimate alternative to passwords?

| Tech Sanity Check: has a commemtary about yet another Multi-factor Authentication scheme.

vidoop has this interesting scheme for multi-factor. The interesting twist in this case is that the scheme has the potential to be based more on how you think than a known fact. You might select a sequence of pictures containing boats, airplanes, and cars. The theme could be transportation, the colour blue or aluminium.

They aren't there yet - it seems their current scheme relies on you thinking they way they do. But it does strike me as an improvement over static images.

Sunday, April 08, 2007

Shared Items

If this works out right, you should now see a shared items feed near the top of this page. What are shared items? Well many sites produce a summary of their content in a form that can easily be syndicated - that is consumed by other sites, or applications. Often referred to as a 'feed' or an 'RSS feed'. RSS is a particular protocol used to implement this feature. ATOM is another such protocol, although they often get lumped together and called RSS.

So what can you do with it? There are many ways by which you can consume these feeds. Internet Explorer 7 has such a feature built in. So does my.yahoo. My favourite mechanism is Google Reader.

If you know somebody who reads a lot of internet content, you may wish you could 'read over their shoulder' every time they say 'Hmm. That's interesting!' Well that is exactly what Google reader shared items let us do. If I find something if interest, I can mark it 'shared', and it will then show up on my shared items feed. If you'd rather just look at a web page of the same items, you can do that instead.

There is one thing I would like to be able to do, that 'shared items' doesn't allow. And that is to make a comment on an item I share. Well maybe we don't need that. That's what this blog is for!

Have great day.

Thursday, March 22, 2007

Wednesday, February 07, 2007

Picking the Open Source Winners

This is a big topic for me at the moment. I am very interested in the OpenBRR and Navica scoring methods and the work undertaked by ohloh and optaros.

Wednesday, January 31, 2007

Authentication, Authorization, and Context

I was reading James McGovern's blog today and it reminded me of a conversation I had at work yesterday. James focus is on vendor product - something that is certainly of interest to us, but beyond that we also have to deal with our internal applications.

There are three issues that always pop-up when we try to integrate a new software product.
How do we authenticate?
How do we authorize?
What was the user try to doin the first place?

We are starting to get a good handle on #1 - but I would like to see us authenticate at fewer points and establish a trust network between applications. SPNEGO, SAML, WS-Federation, Liberty, and perhaps OpenID are all promising.

The James' blog speaks to the second part - authorization - The next big frontier. What roles can the principal hold (for this application)? The application part of that sentence is ways controversial. This is where XACML fits in.

The third piece - what were we doing is a tongue-in-check reminder that the user was actually trying to do a job before security "got in the way". The user likely had some kind of established context that should, ideally, be available to the next application. Although not always the case, it is a frequent requirement. For example, the user may have been working with a customer in the CRM application, and now needs to work on the customers Loan in the credit management application. This customer and loan context information needs to be carried forward. There are no good soutions for this that I know of. This is the undiscovered country.

Any thoughts out there?

Friday, January 19, 2007

Will we need a Fortress for our cores?

Two news theme over the last several months strikes me as mutally relevant.

First there is the information coming from Intel about future core densities. The Gulftown processor will likely make its debut in 2010 and will contain 32 cores. Intel research is also working on an 80-core research prototype. As the InformationWeek article discusses, software will need to change to accomodate.

Speaking of software, Java is a significant development language for the business world. What will business applications do with an 80-core engine? Multi-threaded applications - wherein the threads are explicit to the programmer are not the way to go. To the typically programmer I say "You can't handle the threads".

So what is the second news theme? Fortress - a research effort coming out of Sun Research. It is targeted at High Performance Computing; A replacement for Fortran. Am I suggesting that we rewrite all of our applications in Fortress? No. But Sun is implementing Fortress on the Java Virtual Machine. So there are some good possibilities that future Java language features will be able to drive out the same multi-threaded behaviour for which Fortress is striving.

Thursday, January 18, 2007

Bank loses data

Bloomberg.com: Canada: "CIBC's Talvest Mutual Fund Loses Client Data Files"! This cannot happen too many more times before companies will be forced to put the technology in place to ensure that data that leaves their premises is encrypted. That won't be cheap, but it seems that the cost is inevitable.