Showing posts with label Who's data is it. Show all posts
Showing posts with label Who's data is it. Show all posts

Sunday, January 3, 2010

Who's data is it? (Part 4) - Counter-Point

The past three posts in this series have dealt with negative behaviors or attitudes about the responsibility of data ownership:
  1. It's the vendor's data
  2. It's MY data
  3. It's NOT my data
As I mentioned a the end of part 3, an important factor to consider when looking at any of those situations is the history behind how the team or organization came to be where they are.  Every organization forces behavioral pendulums to swing back and forth between extremes in response to negative outcomes.  What might appear to currently be a negative behavior may have once been a reasonable response to some other negative experience.

For instance, "it's the vendor's data" may well have developed in response to some historical data conversion or data integration project that well over schedule and over budget because data structures were poorly understood by the internal team and the use vendor services was discovered, too late, to be a huge benefit.


Likewise, the "it's MY data" crowd might come from an application support team that has come to understand the value of information integrity and master data, and has developed a sense of protectionism from those goals the hard way: by failing to meet past goals.  "We're responsible for maintaining CUSTOMER data for the company.  If someone else starts using that data without our close direction, then they'll misunderstand what they're looking at and misuse the data."

The philosophy behind this perspective is very reasonable, but the implementation becomes one of closing off access to information rather than increasing and easing accessibility and educating users and other teams on how data should be used.  Rather than maintaining institutional knowledge exclusively within a support team, how to use key information should be exported to the larger enterprise.

People who are focused on executing business processes tend to examine technology and information management from a localized context: what information does this  process need as input and what applications will this process use?  This siloed view lacks the concept that a great deal of value comes from examining the space in between business processes or applications.  In the space between various business applications, there exists an opportunity to gain a higher-level of understanding about information interrelates between those various silos.

Conclusion:
Thanks for joining me on this conversation about data ownership.  Personally, I detest the word "ownership."  It has negative historical connotations that somehow suggest to me "ownership of data" implies the "enslavement of data."  But I struggle to find a more appropriate way to refer the concept that someone has certain responsibilities when it comes to managing information.

Perhaps it's a bit trite, but to borrow from the famous Native American proverb about the earth:

We do not inherit data from upstream systems; we provide information to downstream ones.

Friday, January 1, 2010

Who's data is it? (Part 3)

In the first two parts of this series on negative attitudes on data ownership, I talked about protectionists who want to keep tight control over who has access to what data.  The most important negative consequence of that attitude is that it stifles both understanding and innovation. Fewer people looking at data means that the data is being examined from fewer perspectives.  Innovation comes from looking at existing situations or information from a new perspective and reaching new conclusions.  So, one of the best ways to confront data protectionists is to put the argument in terms of that business driver: innovation and resulting improvements.

It's NOT my data...

The final negative data ownership pattern I'll present in this series is the "it's NOT my data" attitude.  In this scenario, people within the information supply chain simply abdicate responsibility for anything regarding information, beyond their immediate operational duties.  This can manifest itself in several ways.  With respect to data quality concerns, the "not my data" crowd will simply point their fingers upstream:
  • "That's what's in the system."
  • "That's what the sales team entered."
  • "That's what the customer told me."
All of these are only excuses for poor data quality.  By themselves, they do nothing to address data quality problems.  These are reflective of a culture that doesn't look beyond the operational duty of data entry and service or product delivery. 

A more subtle manifestation of the same attitude might look something like "You can have the data, but don't ask me what it means.  You'd have to ask so-and-so about that."  And when you ask "so-and-so" he directs you to "what's-her-face," who tells you to ask her supervisor, who explains that they're just doing it the way the Standard Operating Procedure tells them to.  In this chain of inquiry, everyone is abdicating their own responsibility to understand how their work fits into a larger picture.  It's a form of willful ignorance.

In his motivational work, Christopher Avery uses a model for responsibility that describes several responses that all come before actually taking responsibility for a situation:
  • Denial
  • Lay Blame
  • Justify
  • Shame
  • Obligation
In the "not my data" culture, groups are stuck in one of those first three levels: denial, lay blame, justify, or shame.  They either deny that there is anything to even consider with regard to data ownership and data quality; they lay blame on other constituencies that are upstream from them (even all the way to a customer!); or they justify the situation, telling themselves that there simply isn't another way that things can be done.

The engineer in me sees a certain appeal to "not my data" situations.  In the extreme, they're a great challenge in reverse engineering.  Every steps yields more questions than it answers, and opens new places to explore for broken processes and more denial of responsibility.  It's like a software debugging exercise that leads you into deeper and deeper through twists and turns of function callbacks until suddenly you discover that critical nugget of information that finally allows you to describe the end to end flow of information, and the real impacts of poorly implemented processes or policies.

The real-world corporate director in me sees situations that need to be addressed, supervisors to be educated or replaced, policies and procedures to be changed -- all great places for change and growth -- but also, much less enthusiastically, politics to be navigated.

Counter-Points...
There are always multiple perspectives to any situation.  Teams and organizations don't usually evolve what may appear as negative attitudes out of ignorance or malice alone.  Often, other negative influences or behaviors lead to a culture that once served a valuable purpose, but may no longer yield more benefits than harm as the previously negative influence dissipate.  The final in this series will present situational counter-points to some of the arguments I've presented.

Wednesday, December 30, 2009

Who's data is it? (Part 2)

In the first part of this series on who's data is it, I talked about the black-box mentality that some teams have with regard to vendor applications.  In that scenario, the hurdle to overcome is primarily one of convincing the right authorities to grant you database access.  That's the biggest hurdle, assuming the application data is actually understandable once you get at the underlying database.  Unless the application team is more loyal to the vendor than to your company, you haven't burned any important bridges in that scenario.  The vendor won't support your efforts at extracting and integrating data from the application this way, but they're the one's who put out a ridiculously high $250,000 bid anyway.  The next scenario is more politically charged, however...

It's MY data!

The "it's my data" culture sees the information behind an application as something that needs to be closely held and funneled through a controlled group of experts.  Unlike the vendor-data perspective, the my-data group is interested in sharing information outside of the application interface, but often with point-to-point interfaces that they can tightly control (as opposed to any kind of pub-sub or hub-spoke model that creates a more open access model).  They're interaction often sounds like: "You provide exact specifications with which specific fields you want and we'll send you an extract with exactly that information."

Another common response from application teams to a request for access to an application database might be something along the lines of "what are you going to do with that data?  We have responsibility for that data and it can be very complicated to understand."  This kind of response is indicative of a belief that the application team has some kind of exclusive ownership of the data and should be the only group allowed to parcel information out to users individually through reports or extracts.  At the extreme, it's reminiscent of early reporting days where a user would call someone in IT for a report; that IT person would build the program to generate the report; print the entire report on green-bar paper; and then send it through interoffice mail to the requester.  If the application team isn't already on board with a service oriented approach to exposing data to other applications or integrated reporting through data warehousing, then their natural tendency will be to create a plethora of individual point-to-point solutions across the application landscape.

This kind of architecture may lead to well defined standards and definitions of data, but either only in pockets across the organization or with an ever increasing cost of maintenance as the number of interfaces expands and complexity of that interaction increases.  There are some positive aspects to a culture that wants to protect the purity of its data and have a strong understanding of how other groups are using the same data.  Part 3 of this series will touch on that data purist counter-point.

Who's data is it? (Part 1)

The problem of data ownership:
I want to thank Jim Harris for his great post about The First Law of Data Quality from earlier this month. It's certainly an excellent read.  One of the things it reminded me of is the experiences that I've had working with various other business and IT teams when trying to get access to and understand data from particular systems.  The issue of "data ownership" always seems to play an antagonistic role in any desire to acquire and integrate information from multiple systems.  Generalizing, I think there are three different negative responses you can receive when looking to a new source system for access to their data for integration into a data warehouse:
  • That's the vendor's data.  We wouldn't be able to understand it.
  • That's my data.  I'll tell you what you need.
  • That's not my data.  I don't know anything about it.
This series of blog posts will explore each of those in more detail:

It's the vendor's data...
The "it's the vendor's data" culture has a black-box view of applications: there is no separate understanding of application and information.  So, there's a belief that the only way to interact with the data in the application is have a vendor or partner create new application features or vendor supported interfaces.

Individuals who have experience in supporting the end users of applications and sometimes the administrative configuration of applications often put significant weight in the features and functionality of the application itself.  From an integration perspective, this significantly limits the kind of integration that can be done with many applications.  For instance, most applications provide users with various types of "report writing" functionality.  In many applications, this is merely an extract-generation feature that allows users to select the fields they're interested in and exporting that data to a CSV or Excel file.  Application support analysts who are used to supporting users in this way and have, maybe, not recently been involved in technical integration work may default to a view that the only way to get information out that particular application is through these manual extract tools.  Of course, having individuals manually run extracts from a system on a daily basis is not an ideal pattern for sourcing data into a data movement process.

In situations of heavy application reliance, application teams might never think of doing integration directly from the database; or the team might feel that any such endeavor would require many hours of training from the vendor.  Of course, the vendor would appreciate the professional services or training for that.  In most applications, however, that is typically unnecessary.  Reverse engineering a data model is an easy exercise and, depending on the complexity of system.  Reverse engineering the underlying data typically takes more analysis, but having access to both the application and database (in several environments) provides a powerful way to understand both the application and the data.

Enterprise scale applications typically provide some time of standard interfaces, whether those are X12, EDI, HL7, or a proprietary way of moving data in and out of the system.  Many modern applications provide a service oriented API to allow external applications to do read or write operations to the application.  Standard interfaces and services are an excellent way to retrieve data for a data warehouse.  In lieu of sufficient interfaces, though, database to data warehouse integration is still a common, powerful, and appropriate integration model.

Convincing application teams that connecting directly to an application database without engaging and paying for professional services from the vendor can be challenging.  One effective way to break through part of that cultural difference is to do a proof of concept and show how the technology and analysis can work to create custom integration solutions without necessarily needing vendor intervention. A quick win to prove that the mysterious data behind some application really is just letters and numbers goes a long way to gaining confidence more substantial integration work.