Friday, April 20, 2012

Preparing for Predictive Analytics - Data is the Key

The Basics

As I mentioned in my last post, I’ll be making a series of posts on some of the challenges that you will face when embarking on a predictive analytics project. In this post, I’m going to focus on what may be obvious to most, but frequently has proven to be a challenge for the customers we have worked with. Namely having ready access to the required data.

Your organization’s data is the key to a successful predictive analytics project. Quality historical data is required in order to build models that will let you make predictions about the likelihood of some event or behavior. Not all data is relevant in all modeling scenarios, but generally the more information you have the better. The modeling exercise will weed out the noise from the signal. Some modeling techniques are better than others at dealing with the noise as well. Familiarity with the capabilities of your tool is very important in this context.

Cataloging Your Data

How many of you know all the different kinds of data used in your company? My experience has been that most organizations have a lot of data in silos that are not well documented and certainly not well integrated with other corporate data. This data can range from duplicate customer data, to sales data, to communication data such as email logs, marketing data, or other data needed to GSD (Get Stuff Done). Master data management projects can help your organization centralize, de-duplicate, and cleanse that data, but the reality is that these kinds of projects can take years to complete and are very complex. Minimally, it would be helpful to begin with a data cataloging project to at least get your arms around the data that your organization has. Start with the basics. Identify the types of data you have and who is responsible for maintaining it. Make note of where the data lives; i.e., what tool/platform was the system developed in. Besides laying the ground work for a master data management project down the road, this will be extremely valuable in your predictive analytics projects because it will outline where you need to go to get the information you want to model.

Does It Matter How the Data is Persisted?

The answer to this is highly dependent upon the tool or platform you are using in order to create your models and then to subsequently score the data. If you are using a commercial tool, your choices are limited by what that tool needs. Generally, the best answer is to bring the data together into a consistent storage medium. Whether that’s a relational database, a data warehouse, flat files, or XML files, the important thing is that the data can be accessed and interrogated as a set. Efficient set based operations are critical to the performance of your analytics solution. Many of the modeling activities will involve slicing, dicing, counting, aggregating, and transforming your data set in numerous ways. This can be a very slow process if your data is not stored in a way that supports those kinds of activities.

Analytics Repositories

My recommendation is that whenever possible you should try to collect your data into a centralized analytics repository. With smaller data sets this is much more approachable and is a common way to do it. However, with very large scale enterprise data this can be expensive, time consuming, and impractical. The time to update the repository can make timely model analysis impossible especially if you are trying to model transactional data that is quickly being added or changed.

Wrapping Up

In my next post, I’ll be discussing two different approaches to designing analytics repositories that address these two scenarios. The first approach is to use the simpler relational repository to store the data set. The second approach is to use a virtualized metadata-driven repository, which can be extremely useful in the larger scale enterprise settings.

Tuesday, March 20, 2012

Preparing for Predictive Analytics

It’s been awhile since I have written here because I’ve been heads down on project work and have just had a chance to come up for air. I want to begin with some short posts to discuss a number of challenges that I’m currently working through. The major theme will focus on how companies can prepare for predictive analytics. If you aren’t familiar with the term, here’s a quick link to the Wikipedia topic to get you started.

About a year ago I joined a small company that is focused on delivering top-notch predictive analytics software to address some very specific commercial and educational market needs. Since joining the company as the solution architect, I have found that the single most complex and dynamic part of our engagement and solution delivery process is data acquisition, normalization, and access. This is largely because of the diversity in the platforms, technologies, and applications written for and used by our clients.

We work with our clients to analyze their data for the purposes of building predictive models. The requirements for model building are pretty straightforward. We need a clean and consistent view of the data - requirements that are not unlike any other analytical or reporting process needs. But the complexity of today’s enterprise environments make this more challenging. Additionally, we aren’t looking at a snapshot of data at a single point in time. We are looking at the data in real time or near real time in many cases. We are often working with both structured and unstructured data that come in a myriad of formats and accessed using many protocols. Everything from relational data in databases, to web services, to flat files, spreadsheets, etc. All of the information needs to be identified, cataloged, gathered, date/time stamped, and recorded for time based analysis.

Clients that understand master data management and have sound data governance policies are easier to work with because they understand the value of their data and most importantly how to get it. At the other end of the spectrum are those companies that have their data in many disparate systems, have no data governance or ownership policies, and don’t know the value of their data. Getting access to their data and getting it into a clean and consistent form can be quite a challenge.

Therefore, my next few posts will talk about the challenges we are facing and our approach to solving the data integration and normalization needs for a predictive analytics solution. My hope is that the information you find here will help you prepare for using predictive analytics in your organization to improve and optimize the decision making you do on a daily basis using one of your company’s most valuable assets – your data.

/imapcgeek

Tuesday, January 31, 2012

WCF Test Client Error

When starting to debug a new WCF service I received the following error message:

“The contract ‘IMetatdataExchange’ in client configuration does not match the name in service contract, or there is no valid method in this contract.”

image

I had added my MEX endpoint in the web.config and was able to review the service WSDL through a browser, so I was really not sure why I was getting this. As it turns out, this error message is caused by an entry in the machine.config for an endpoint declaration defined like this:

<endpoint address="" binding="netTcpRelayBinding" contract="IMetadataExchange" name="sb" />

I’ve commented it out for now and the WCF Test Client is now able to properly interrogate the service metadata without giving me an error.

After a bit of Google’ing this error, I found this explanation from Joel C on StackOverflow.com. Looks like at some point an installation for the .NET Service SDK (which I don’t remember installing) updated the machine.config with that information.

For now, it works having removed it.

Monday, January 30, 2012

You Are Dead to Me

OK, even if Silverlight isn’t completely dead, it certainly is starting to smell pretty bad. Mid last year we were looking at Silverlight as being a viable solution for our UI needs. It has a rich programming model.  The user experience is excellent and it has pretty decent market penetration – certainly not as good as Flash, but respectable.

However, as the majority of our development is greenfield and we are looking to build for the future, it just didn’t make long term sense for us to consider building on a product whose future was looking pretty iffy. When you consider that Microsoft recently cancelled Mix 2012, it’s clear to me that the future lies elsewhere.

Sure, you can still build on Silverlight. And sure it will be supported for quite some time to come. But I asked myself why build on a technology that is clearly questionable in its future? Certainly Microsoft is continuing to support XAML development for Metro style applications. But I questioned whether I was just delaying the inevitable.

In the end, I felt the best decision was to go with HTML5 for the applications we are building for the future. We have begun active development in ASP.NET MVC3 with the Razor view engine. We have achieved our goals of providing a rich user experience, excellent performance, and a testable loosely-coupled code base. It’s also a solution we can work with right now, which means we can deliver business value right away without sacrificing functionality.

v.Next of our applications will give us the opportunity to revisit the scene and reevaluate the options. I think the best news is that there is an excellent set of options out there now and the future looks very bright indeed.

Wednesday, January 18, 2012

Creating a Dependency Injected WCF Data Service

In this post, I’ll explain how to decouple your WCF data service from a specific data context. This is useful in many ways including changing out your context implementation at runtime or mocking out your context for testing purposes. The example code will use MEF (Managed Extensibility Framework), but any dependency injection framework, service locator implementation, or factory pattern could be used.

Getting Started

The first thing you will need to get started is an implementation of DbContext. I’ve chosen to go with a code first approach in this example because it gives me explicit control over the code in the context. Entity Framework 4.1, which is available as a Nuget package, provides a simple, purpose-driven API, which allows me to create a new DbContext and specify the entity types that it is responsible for in just a few lines of code. The base class and EF4.1 plumbing does all the hard work.

public class PersonContext : DbContext
{
public IDbSet<Person> People { get; set; }
public IDbSet<Address> Addresses { get; set; }
}

Figure 1: Basic PersonContext implementation

This implementation satisfies the most basic requirements of EF4.1 for a DbContext, but as you’ll see as we go along, we’ll want to flesh it out a bit more in order to support MEF’s requirements.


Give Your DbContext an Interface


Uncle Bob Martin’s SOLID object oriented principles is a must-read for all programmers. The ‘D’ in SOLID is for Dependency Inversion. In a nutshell, dependencies between objects should be based on abstractions not concretions.  Our next step is to create an interface that will represent our DbContext.

 

public interface IDbContext : IDisposable
{
IDbSet<TEntity> Set<TEntity>() where TEntity : class;
int SaveChanges();
}
Figure 2: IDbContext interface

 

Initializing Your Context


The EF4.1 DbContext class can take a connection string in as a constructor parameter. For simple cases, this might be enough for you. However, if your context requires additional information to operate properly, you may want create a configuration class that you can inject into your context through its constructor. I’ve take this approach because it allows for greater flexibility in initialization of the context.


[Export]
public class DbContextConfiguration
{
public string Name { get; set; }
public string ConnectionString { get; set; }
}

Figure 3: DbContextConfiguration needed to initialize your DbContext

 

[ImportingConstructor]
public PersonContext([Import("PersonContextConfiguration", typeof(DbContextConfiguration))] DbContextConfiguration configuration)
: base(configuration.ConnectionString)
{
_configuration = configuration;
}
Figure 4: PersonContext constructor

 

The first thing you’ll notice in the code above are the [Export], [ImportingConstructor], and [Import] attributes that are applied to the DbContextConfiguration class and PersonContext constructor. These are MEF attributes, which are used to support the dependency injection pattern. MEF will handle auto-magically wiring up the dependencies through a call to ComposeParts(). If you are new to MEF or unfamiliar with how it works, here are a couple of useful links to get you started.

 



 

Creating Your WCF Data Service


Writing a WCF data service couldn’t be easier. In a matter of a few mouse clicks, you can be serving up REST-based data. Microsoft has dramatically reduced the effort required to implement a service by providing a rich API that provides a lot of functionality under the covers.


image


Figure 5: Add New Item


Adding a new WCF Data Service through the Add New Item dialog results in an entry point to that API via the DataService<T> class.



   1: public class PersonDataService : DataService< /* TODO: put your data source class name here */ >
   2:     {
   3:         // This method is called only once to initialize service-wide policies.
   4:         public static void InitializeService(DataServiceConfiguration config)
   5:         {
   6:             // TODO: set rules to indicate which entity sets and service operations are visible, updatable, etc.
   7:             // Examples:
   8:             // config.SetEntitySetAccessRule("MyEntityset", EntitySetRights.AllRead);
   9:             // config.SetServiceOperationAccessRule("MyServiceOperation", ServiceOperationRights.All);
  10:             config.DataServiceBehavior.MaxProtocolVersion = DataServiceProtocolVersion.V3;
  11:         }
  12:     }

Figure 6: Boilerplate data service code


Simply replace the boilerplate TODO between the generic template brackets with a class that extends DbContext and you’re practically good to go. They even provide code snippets via comments that serve as placeholders for configuration changes that you can use to set access rules for the entity sets exposed by your DbContext.


*Note: in my example above, the DataServiceProtocolVersion.V3 is implemented in the October 2011 CTP release for data services. I began using it in order to integrate with the Entity Framework 4.1 DbContext API.


In our case, we will declare the data service this way:


public class PersonDataService : DataService<IDbContext> {…}


Override CreateDataSource


Next is the final piece of the puzzle. The DataService class provides a convenient way to control the instantiation of your service’s data source dependency. Namely - we will override the CreateDataSource method with an implementation such as this:


protected override IDbContext CreateDataSource()
{
context = compositionContainer.GetExportedValue<IDbContext>();
return context;
}

Figure 7: CreateDataSource implementation

 

Your implementation may differ, but essentially you want to use your MEF composition container, service locator, or other factory implementation in order to get an instance of your context based on the interface type. MEF has numerous ways for you to resolve this dependency.

 

A critical concern for your code in this area is to handle the multiplicity of dependencies you may find in your container. In a future post, I’ll demonstrate how to specify metadata to filter the results of your composition request.

 

Conclusion


Creating a loosely coupled design can sometimes be a challenge. In the case of WCF Data Services, Microsoft had the foresight to make this chore easier. Through abstraction of your data context, use of an extensibility framework like MEF, service locator or factory implementation, and a little glue code, you can decouple the service and context and reap the rewards of runtime composition and increased testability.

 

Friday, November 18, 2011

Duplicate MEF Exports When Export Has Metadata

I just discovered that the Managed Extensibility Framework will produce duplicate exports for parts that are exported with custom metadata. My scenario is pretty simple. I have created an Entity Framework DbContext subclass that is marked up with both [Export] and [DbContextMetadata] attributes. DbContextMetadataAttribute is my own custom metadata attribute.

   1: [Export(typeof(IDbContext))]
   2: [DbContextMetadata(ContextType = "Person")]
   3: public class PersonContext : DbContext, IDbContext
   4: {

The result is shown here:


image


I have verified that this is true by simply commenting out the custom metadata attribute. Interestingly, this behavior is not present when using the MEF ExportMetadataAttribute. I’m planning to dig into this a little more to see why it’s happening, but it certainly was unexpected.


 


Blogger Labels: Duplicate,Exports,Export,Metadata,Framework,custom,scenario,DbContext,DbContextMetadata,DbContextMetadataAttribute,IDbContext,ContextType,Person,PersonContext,behavior,ExportMetadataAttribute

Monday, November 14, 2011

How Much UML Modeling Is Right For Your Team Or Project?

This is a question I see asked a number of times in various forums. It is one of those “It Depends” sort of topics, but is one that I think is worth discussing here. In this post, I’ll give you my perspective on how I’ve used it in my projects.

Probably the most important thing you have to understand is, “Why are you using UML at all?” – assuming you are in the first place, or are considering using it in the near term. As the old saying goes, a picture is worth a thousand words. A UML diagram is simply a picture of part of a (potentially) complex system. And sometimes the best way to break down the complexity is through pictures.UML diagrams convey structure or behavior through pictures of classes and their interactions. It supports numerous diagram types to break down the system into various views, which can be combined in order to represent the system from different angles in order to help bring perspective and clarity.

image   image

Organized in a logical way, the UML diagrams you create help to tell the story of the system you are developing. The $64,000 question is, “Who is the story’s primary audience?”. Generally, the answer to that is the development team that is building the application. However, with the success of Agile, many teams feel that UML is unnecessary, or at best, it’s used in ad hoc ways with very lightweight and high level diagrams – maybe even temporarily on whiteboards during meetings. It’s simply viewed as a means to convey a concept or outline. I certainly agree with the spirit of that. UML can help to speed up the team because it helps to bring understanding and consensus through a shared view of what they’re building. How lightweight or detailed your diagrams are, or even if they are persisted, should be a decision you make based on the team’s appetite and ultimately the overall value add to the team and project. Flexibility is the key.

And no matter what the long term goals are, the short term benefit of an increase in velocity is probably the most valuable takeaway you will realize by using UML.

In order to gauge the right amount of diagraming to provide to the team, I have used the sprint retrospectives to reflect on this with the team and arrive at the right answer. The teams that I’ve worked on have all been comprised of a mixture of experience levels. Typically, the less experienced team members or newer team members gain the most benefit from the diagrams. The more experienced team members still benefit from them, though, as they almost always help to disambiguate the design details. The net effect is that it helps to level the playing field across the team and improves communication and ultimately productivity. I have almost always chosen to persist the diagrams in a modeling tool – even if for no one else but myself for future reference.

In a highly Agile team where practices like TDD are used, the use of UML may likely be perceived as an impediment. You can make a pretty well supported argument that this viewpoint is right. The design should emerge from the creation and iterative refactoring of code and unit tests to flesh out the details. Up front design using UML is the opposite of that approach. Unless you consider it from the perspective of general high-level architectural point of view that is. The UML diagrams you use can describe many higher level aspects of the system like: patterns; guidelines on organizing system components; describing the layers of abstraction;  deployment details; or other high-level details. Let TDD do what it’s best at – creating flexible and resilient designs at a low level of detail.

Most development teams have a certain amount of turnover throughout their lifetime. The reasons for this are myriad, but the bottom line is that you will need to bring new team members up to speed at various points in time throughout the life of your project. If the system that they are working on is reasonably complex, the UML diagrams that you created to convey your design concepts to the original team members can be an excellent way to help the new team members understand what they are working on. I am a firm believer that every team member should get the big picture. They shouldn’t be relegated to some dark corner of the application with little or no visibility into how it all fits together. Use of tools like UML helps to bring everyone up to speed, which is a good thing.

A common argument that I have heard against using UML is that like most documentation it will always be stale when compared to the code. I would say that this is generally true. There are many ways to deal with this, but probably the best way is to just accept that fact and understand that the main use of the tool is to help people understand the system. If the diagrams are too far out of alignment with reality, update them. Most tools support reverse engineering. Use it to refresh the details and update the diagrams. Your future team members will thank you. Just don’t get bogged down in the minutia of sync’ing code with models all the time. That’s one of the quickest killers of usefulness.

In the end, I will also yield with an “It Depends” answer for the original question. Since no two teams and no two projects are exactly alike, I’d say you will have to gauge the right answer based on your current circumstances. For me, I’ve always relied on UML as a means to reducing complexity and organize thought around the structure and design of complex systems. Let your team help you to decide the right answer.

Happy Modeling

Smile

Thursday, November 10, 2011

Fowler Is Right On the Mark About Premature Ramp-Up

Martin Fowler’s most recent post on “PrematureRampUp” http://martinfowler.com/bliki/PrematureRampUp.html is right on the mark. Our team has recently concluded work on two legacy projects, which have been running in parallel for some time. A strategic greenfield project for a custom analytics platform has been slowly gaining momentum. Rather than slam the custom platform project with all the devs from the previous two projects, we have started to migrate them over in pairs. This has proven to be very beneficial in that it has allowed us to control the team member learning curves and ensure that they are properly acclimated to the new project environment. We eliminated a great deal of hectic stress by bringing them over in this way and it also allows the new team members to assist with the knowledge sharing as well. Overall, this is a great way to grow a team to its optimal size.

Wednesday, November 2, 2011

Productivity Tip–Add Your FxCop Project to the Solution

image

I’m very interested in streamlining the workflow that our team uses when developing. That means I need to make the process get out of our way as much as possible. Sometimes, small changes can make a big difference in your overall productivity.

To some, code analysis is a waste of time, but our team sees it as a vital step in ensuring high quality code in the application. Obviously, I don’t want to forego code analysis, but I would like it to be as unobtrusive as possible.

As Martin Fowler says, “if it hurts, do it more often”. Committing changes that break the continuous build for something simple like code analysis violations is a waste of valuable cycles. So we do it earlier and more often. We’ve simply added the .fxcop file to our solution and added all of the project outputs to the project. Now, before committing our changes, we simply run the analysis prior to our commit and eliminate the breaks before they break the build.

image

This keeps our CruiseControl projects green longer and reduces the churn of having to rebuild to fix something we shouldn’t have committed in the first place.

I suggest that in your team retrospectives you reflect on your daily practices and look for small, quick wins like this. You’d be amazed at how much time throughout your day you can reclaim by eliminating friction points like this.

image

Friday, September 2, 2011

What’s It Like Working For A Startup?

OK, well the company I’m at technically isn’t a startup. It’s been around for over 10 years. But its focus has changed from being primarily a research organization to a commercial one that is productizing its successful research initiatives. For all intents and purposes it might as well be considered a startup.

Pros

Probably the greatest benefit to working in a company like this is that almost everything we do is greenfield – and I don’t mean just software development, I mean business process development too. We get the unique opportunity to form the way the business operates from the ground up. Find something that doesn’t work the way you like? Change it. Find something that works really well? Use it. Nothing is off the table.

Even though we aren’t a software development company per se, it does form a significant portion of what we are. And to that end we have the responsibility to deliver high-quality systems. In order to deliver a robust solution to the client we have to employ sound architectural and software engineering practices. And the beauty of it is that we can look to the successes of other companies that have used Agile and Lean methodologies to their advantage and incorporate the same principles here. No having to strip away layers of process in order to make a streamlined software development lifecycle.

Naturally, we’re a small company. That means that in order to get things done people are empowered. Not just empowered to do their day-to-day things, but to actually think and reflect on the way we do things and come together as a team to make real change. It takes the concept of team ownership of code to a whole new level – to the company level – where you are truly an owner in the products and the way they are delivered.

Small companies such as this also bring one very important benefit to the table – that you’re not just a number in a faceless corporate machine. People know you and you know them. It takes the benefits of an Agile team to the next level and brings a sense of camaraderie to the game that makes everyone feel that they have others that they can depend on to get the job done.

Cons

Not surprisingly, as a small company in this phase of its life, budgets are tight. Everything from server hardware costs, to developer tool costs, to travel budgets for conferences is looked at from a necessity perspective. Unlike some companies with deep pockets and seemingly bottomless pits of funds for infrastructure, we have to be vigilant and ruthlessly efficient in our expenditures.

Probably not unique to startups, but common to small companies, is the need to wear many hats. As the software architect, I often field common support needs on the network or server platforms. I participate in presales calls and lend guidance to the business development team to help bridge the gap between techies and non-techies.

Pressure to build up a sustainable and recurring revenue stream is high. A small company operating in a niche market in a tumultuous economy has a lot to worry about. The focus of each and every team member has to be like a laser to deliver a high-quality solution that the client simply can’t afford to live without. Frankly, that should be the focus anyway, so making that a negative probably isn’t fair. The point is that letting your focus slip a little can have dire effects on a small company in a fledgling state.

In Summary

Although there are a number of pros and cons to working in this sort of company the pros far outweigh the cons. So many of the day-to-day impediments that hamstring teams and projects can be eliminated through tight teamwork and the empowerment that comes from knowing that you can make important decisions to improve things.

Funny thing is that you don’t necessarily have to work in a small start-up to have your cake and eat it too. Larger, well established companies can achieve the same results by empowering their employees and fostering a culture of entrepreneurial spirit. The main difference is the depth of the sludge you have to dig through and how far can you go to clean it up.

Good luck!

Smile