Showing posts with label Software. Show all posts
Showing posts with label Software. Show all posts

Thursday, April 04, 2019

Mindful Design: How and Why to Make Design Decisions for the Good of Those Using Your Product by Scott Riley

Mindful Design: How and Why to Make Design Decisions for the Good of Those Using Your ProductMindful Design: How and Why to Make Design Decisions for the Good of Those Using Your Product by Scott Riley
My rating: 5 of 5 stars

A wonderful book touching many aspects of design. One should not expect any technical suggestions or any process suggestions.

The key message that the author is giving in this is that instead of following a linear design where the user is guided through a path to achieve a goal one should design systems which allow the users to explore and learn the system and let them use the system in ways not anticipated by the designer. The second message that the author gives is that we need to move way from the penchant to make applications addictive and instead make applications which are assistive. I.e. applications that help the users solve a real problem in their life. One needs to be more empathetic and ethical in one's design. To be more empathetic one needs to consider a diverse set of persona and not just "white males" who can afford a $1000+ mobile phone and keep upgrading.

Here are some excerpts (some extracts and some summaries) from the book.

General Points
===========
Think of your structure and categories from an emotional perspective as well as a thematic one. What makes sense to you or your internal team is often a source of frustration when asked in the real world.

By grouping objects and content astutely, we allow the brain to disregard an entire “block” rather than having to selectively filter every item.

Structure and hierarchy form the backbone of a good UI. By surfacing what’s important and filtering out what isn’t, we can replicate the mind’s attentional system to a degree.

The state of an interface is changing all the time. Our goal isn’t to present a static structure of information, but to have our interfaces adhere to the perceived goals of the person using it at any given time.

Change and importance are the two key factors we can utilize to guide someone’s attention around our interface. With good use of contrast, hierarchy, and animation, we can ease people through tasks that would otherwise be taxing.

Attention is precious and limited. We operate under a genuine constraint in a zero-sum game, and a huge part of good design lies in the simplification of the dense and the difficult.

A very valid point "As I mentioned in Chapter 2, the acceptance among a population of a particular pattern or signifier portrays an extremely valid argument for convention; however, countless innovations, including the touch interfaces that revolutionized modern technology, would have never occurred if we always relied on or settled upon such conventions. "

"Often our work is best directed toward creating effective tools of cognitive easing."

Furthermore, we can and should embrace the avoidance of attempted forced motivation. Great design can lie simply in the acceptance that many times our products plainly do not need any explicit plan for reward-motivated behavior—whether that’s because they’re autotelic and provide their own intrinsic fun or value, or simply because the value they do provide is in their invisibility and essentialism."

To make a product attractive rather than just talk about the features
1. Let the user experience the features by running them through a sample.
2. Solve a user's problem rather than give a technical mechanism. E.g. Adobe photoshop allows the users to "remove red eyes", "darken the photograph", "lighten the photograph" etc.

Linearity
=======
We can and should embrace the avoidance of attempted forced motivation. Great design can lie simply in the acceptance that many times our products plainly do not need any explicit plan for reward-motivated behavior—whether that’s because they’re autotelic and provide their own intrinsic fun or value, or simply because the value they do provide is in their invisibility and essentialism.
To make a product attractive rather than just talk about the features
1. Let the user experience the features by running them through a sample.
2. Solve a user's problem rather than give a technical mechanism. E.g. Adobe photoshop allows the users to "remove red eyes", "darken the photograph", "lighten the photograph" etc.
Gauging intent is globally important, but especially when ushering someone into a linear flow from a more open, explorative environment. This comes purely down to the parallel between increased linearity and reduced control. By limiting options and directing someone through a constrained path, you’re essentially taking control away from them. Do this for long enough, and you run the risk of people becoming disinterested and losing motivation. Once you’re sure that a linear process is necessary and you’ve determined the essential elements and features for this linear flow, keep the following optimizations in mind:
1. Limit Options. The first step to linearity is to limit the available options and actions to the bare minimum required. When entering into a linear flow, we can reasonably assume that we must cater for a period of focused attention. That means being ruthless with the options we present and removing any distractions, clearly grouping and spacing related elements, and ensuring that copy is terse and clear.
2. Understand the Path to Completion. By their nature, linear paths have an ideal endpoint. Keep this endpoint in mind and constantly question how any decision you make allows for progression toward this.
3. Display Progression. Similarly, find ways to communicate someone’s progression through this process. “Linear” doesn’t always mean short and doesn’t ever have to mean compressed. In many cases, our linear flows might be broken down into multiple smaller steps. Whether our linear flows are explicitly stepped through or progress is a little more ambiguous, we should look to communicate how far toward this aforementioned end goal someone is at any point. This can be done through explicit progress indicators, such as numbered steps or progress bars, or through implicit indicators, such as Almost there! prompts.
4. Be Explicit and Timely. Linear moments in interfaces suffer immensely if ambiguity is introduced into the equation. Furthermore, delayed responses—such as waiting until the end state of a checkout flow to show an error in the first section of the address form—can be equally as catastrophic. Shallow processing and linear flows go hand-in-hand, so ensure that any signifiers and mental models used in communicating concepts are as universally recognizable as possible. A truly linear interface, although rarely encountered in real life, would rely exclusively on recognition and not require any learning or memorizing at all.
5. Design the End State. One of the most important aspects of any flow, but specifically a linear one, is how we present the success or completion state. If the flow involves some form of deletion, then take the learnings of empty states and apply them here. If it involves the creation of data, such as a new Tweet, then present an All done!’ message of some kind and, wherever possible, show their creation in context. If someone has given up control to follow a constrained and linear path to creating, removing, or editing something, the least we can do is let them see clearly that their changes were successful.

Non Linearity
==========
By thinking beyond the obvious visuals and offering multiple paths to the same solution, encouraging exploration and learning, and consistently providing timely feedback, we’re able to create environments that are open enough for self-expression and perceived mastery, but not so opaque or idiosyncratic that they’re too dense to be usable.
VIM and Trello are two examples which follow the non-linearity path. Trello balances between the two.
Research is what takes the amorphous, fuzzy blob that is the idea we have of the people who we’re designing for and turns it into a sharpened, well-documented set of findings, examples, hypotheses, and problem definitions. Most importantly though, everything we do at this stage is bubbling with potential sources of empathy and compassion. This is our chance to truly learn about the humans for whom we’re creating our products. To observe our audience, in all their fallibility and idiosyncrasy, while our project is still comparatively free of bias and guesswork is the best chance we get to truly understand why the thing we’re creating needs to exist.

About Personas
============
Customer Research must mean  "Getting to know people on an emotional and motivational level is absolutely critical at this phase. While data is an essential part of presenting research findings, it’ll never be more important than the intrinsic understanding we can get from delving into the emotional underpinnings of people’s problems, frustrations, excitement, and motivation."

Regardless of how deep into a project you’ll be involved, your understanding of the human aspects of the problem it intends to solve will be of much more important than an accessible and well-documented market analysis.

Now, deep breaths because here are my problems with user personas. First and foremost, they are manufactured empathy, an abstraction of a human being often distilled into a banal and mundane list of life experiences, skills, abilities, and preferences. Human beings do not fit into such a box. All of us have idiosyncrasies, bad days, things we enjoy, secrets, desires, dreams, and ambitions. It’s impossible to cram all of this into an entire book, never mind the single PowerPoint slide format that most personas occupy.
Second, personas can suffer greatly from a cascade of bias and assumption, depending on how many people there are between the people who’ve been observed and the weird, digital protohuman that sits in your research slide deck. It’s impossible, when creating these fantastical fictional characters, to avoid project-level, colleague-level, and personal-level biases and assumptions. So we often end up with personas that not only fail to represent even a fraction of the humanity that we’re really designing for, but that are essentially pre-judged and pre-victimized by the unconscious and systemic biases that live within our teams and ourselves. The overwhelming majority of manifestations of user personas I have witnessed in my career have served both as placation for executives that we feel are too busy, too important, or too removed from the project to display true empathy and as a convenient dictionary of the biases of our teams and companies.
Once we’ve traded away people’s frustrations, eccentricities, and self-expression for nicely formatted human-shaped pigeon holes, these artifacts often become the single point of reference throughout our work until we finally get some working prototypes and ideas in front of real people. This is a dangerous and insidious precedent in an industry that is already insufficient in the empathy department.

Persona Identification/Creation
Research and planning should provide you with answers to the following key questions:
How do I succinctly define the problem we are solving?
Who are we solving this problem for and what makes them tick as human beings?
What frustrates these people?
What motivates them?
What distracts them?
What kind of environments are these problems encountered in?
What mental models already exist in this problem space?
How could our product hurt, or otherwise negatively impact, someone who uses it?

Research and planning usually result in some of the following artifacts:
Competitor and customer research reports
Stakeholder interviews
User personas
Empathy maps
Problem definitions
Research videos and documentation
Usability analysis (for existing products or for your competitors)

Design Brief should answer the following questions
Who are we designing for?
What is the proposed solution?
What features are we going to include?

Design Testing
THE IMPORTANT QUESTIONS
Design testing should answer the following questions:
What false assumptions have I made?
What biases have I introduced in my work?
Does our product “work” in real-world situations?
Am I still solving the correct problems?

THE COMMON DELIVERABLES
Design testing will usually result in the following:
User testing reports and videos
Usability analyses
Accessibility analyses
UX reviews
A/B test reports
Snap-judgment test reports
Revised versions of other deliverables (new wireframes, prototypes, or designs to accommodate test findings, for example)


On Storyboarding
While storyboarding an important aspect is to capture the emotions associated with the different actions in addition to the friction that the user may encounter in using the interface.
Storyboarding allows us to present a condensed version of the moment story we’re trying to tell while still communicating the fundamental goals, motivations, points of friction, climax, and ending that constitute our moments.
A simple approach to storyboarding is to stick to nine panels, have your beginning panel set the scene—preferably with goals and motivations—spend the next six panels incorporating your potential app steps, and finally have panel seven as the moment of climax, panel eight as your ending, and panel nine as a moment of reflection. You can (and should) produce multiple storyboards per moment if you need to communicate more points of friction, different scenarios, and contrasting or conflicting goals and motivations. My advice for a storyboard is to follow a simple framework, starting with a template and following some set steps
Like any group brainstorming session, you’ll want to go away after the meeting and prune the ideas and suggestions that arose. While nobody wants to see their ideas discarded (and this is why we do it after the meeting, as opposed to looking a developer straight in the eye and tearing their Post-it up in front of them), an integral part of our jobs is ensuring that our work is representative and grounded. If the CEO thinks that someone will feel “delighted and in complete awe” at the “seeing-pricing-for-the-first-time” panel in a story, we’ll need to find a way to ease that out in favor of the more realistic “concerned-about-financial-impact” and “wondering-if-this-product-is-actually-worth-it” suggestions we might have. (In this case, feel free to do the whole staring-while-tearing-a-Post-it-up power play—it’s only the CEO.)
An obvious red flag is the kind of unbridled optimism that anyone who might describe a company as “my baby” (founders, CEOs, investors who like to overstep boundaries, etc.) might display. Being overly attached to a product’s success to the point where one believes it to be a flawless solution, devoid of any potential malaise, represents a special kind of myopia. In fact, one of the most important skills you can learn as a designer (that, unfortunately, tends to only come with experience) is the ability to quell this kind of optimism-ad-nauseum without appearing as if you’re a prodigal spoilsport who only exists to ruin dreams and scare children.
Other red flags include opinions from people who have lived a life of sheltered, abundant privilege, who have yet to prove capable of genuine empathy (apologies, but unless you’re working on an app to dodge mansion tax or a croquet-court finder, then coddling the opinions of ever-comfortable people can be one of the fastest ways to sink a product), suggestions that revolve completely around revenue, ideas that rely on manipulation, and proposals that appear completely rooted in “competitor X does this.” (Tread lightly with that last one, though, as competitor research is a hugely valuable insight into convention and mental model consistency.)

Conclusions
I feel that, in calling ourselves designers (of any kind), we accept a certain implicit responsibility to leave the world a little better off than we found it. Furthermore, I believe that the privileged among us should caveat that responsibility—to leave the world a little better than we found it—for people more vulnerable than us.
The ability of technology to augment human existence is an undoubtedly exciting concept and, with the amount of money and innovation in the technology industry right now, the impact of digital products on the world stands to be monumental. Yet, we still operate in an industry that’s rife with naivety, bias, and systemic oppression. As designers, this means that we’re often presented with ideas and problem definitions that are rarely harmonious with humans outside of the affluent world of tech entrepreneurship.
To achieve these goals, especially in an industry that seems to operate by its own rules, eschewing ethics in the name of profit under the cover of the mystique of programming and technological innovation, we must ensure that our process serves to elevate those who are far too often underrepresented in our work.
When an industry decides that plucky upstart founders who solve problems and get it done belong on a pedestal, when the distribution of venture capital is based on an almost instinctive appraisal of the usefulness of a product, when a gross misinterpretation of the satire that is meritocracy becomes the prevalent religion to its core demographic of white 20-something Californians, then the problems that get deemed worthy of solving tend to have a very noticeable, very white, very upper-middle-class feel to them.
While tech’s diversity struggles aren’t the only reason we seem so content to disregard self-expression and creativity in the name of control and dominance, I don’t believe it’s a coincidence that a predominantly white, predominantly male, predominantly affluent cohort presents as a hegemony with these apparent values.
Somewhere in the pursuit for unfettered growth and profit, design’s perceived role has shifted from democratization and empowerment through technology to invasive and manipulative vehicles of pop science at the behest of founders and investors. At a point in time when democracy is failing its most vulnerable, where racism and jingoism permeate every crack in the facade of social media, where the voices of those calling for a white ethnostate are amplified exponentially on a platform that is grossly incapable of suppressing hate speech, where online tools are used in months-long hate campaigns, I feel it’s an important time to revisit our ethical responsibilities. While it would be absurd to lay the blame of such sociopolitical malaise squarely at the feet of design and technology, I feel that, somewhere, somehow, our industry has dehumanized people at a time where we need most to see the vulnerabilities we design for. By breaking away from the common trend of self-serving, narrowly defined problem spaces, by seeing people—all types of people—for who they are, we can look to bring humanity and compassion back into our process.
People are messed up. They do weird things to themselves that fly in the face of survival instinct. Some of us harm ourselves because of chemical imbalances in our brains. Some of us have eating disorders. Some of us have drug addictions. Some of us can’t answer the phone or hold a conversation with a stranger due to anxiety and panic. We’re human beings and we grieve and we screw up and we get excited about silly things and we’re sometimes just as likely to sabotage ourselves as we are to celebrate or reflect positively on our life.
Until we factor this into our work by showing compassion, researching properly, educating ourselves on our privileges and biases, and until we earn our ridiculous salaries, day rates, egregious company perks, and the pedestals our industry so adores by actively designing and building for vulnerable, oppressed, and divergent human beings—how can we possibly say we’re making this world a better place? By releasing yet another product that is only usable by those who meet our ridiculous salaries, day rates, egregious company perks, and the pedestals our industry so adores by actively designing and building for vulnerable, oppressed, and divergent human beings—how can we possibly say we’re making this world a better place? By releasing yet another product that is only usable by those who meet our ridiculously narrow defaults, by those who are capable of the cognitive cost of “normal” attentional faculty—whether through not experiencing poverty, anxiety, sociological impact, or neurodivergence—or by those who simply fit nicely into our abstracted assumption of what makes for a normal human being, what are we really contributing?
Homogeneity’s grip on the technology industry is something that needs to be tackled from multiple angles—and design is but one of these. Many of the tech darling products we aggrandize to the point of hero worship boil down, in essence, into a category of “things white dudes’ parents used to do for them until they moved to silicon valley”. These products make it easy and cheap to get rides to work, while capitalizing on systemic oppression and non-unionized labor. They make it simple to get dinner on the table, while heaping more pressure on an already overworked and underpaid industry of service and delivery workers. They make it easier to get your laundry done, easier to find a place for your fourth vacation of the year. Some of them even make that grueling last half mile of your trip to work slightly more bearable by polluting the streets with electric micro-scooters.
There are many trade-offs to this incessant pursuit of lukewarm innovation, though. An astonishing number of products that we hold up as trailblazing innovators (or unicorns if you want to be as insufferable as the tech press) exist because they’ve found technological solutions to exploiting loopholes. Whether that’s “technology” companies that are able to drive down prices due to dodging government regulations, CEOs with the charismatic depth of an STD claiming anti-union rhetoric as the key to success, the tone-deaf social media platform that refuses to do anything about its Nazi infestation lest it damage vanity metrics, or just your average everyday monolithic company paying 1% of the tax it’s supposed to—the apparent need for tech to subvert humanity and democracy to succeed is a vile byproduct of untempered attempts at innovation—using nothing but a framework of rhetoric, blame deflection, and a tone-deaf agreement that certain, arbitrary metrics are somehow directly translatable into fiscal value.
Modern-day tech capitalism relies on an onion-skin model of subversion. Everything from where the money comes from to pay salaries to the Rosetta-Stone-needed levels of legalese in terms and conditions to the general public’s willingness to gloss over the systemic oppression that is as integral to the bedrock of tech innovations as the server stacks and app stores they exist on (“Well, most of my clothes are made in sweatshops, so why should I care that my taxi driver gets paid an unlivable wage?”) feeds into this implicit, societal permission we give 20-something tech dudes to “go off and innovate.”
All of this leads us to a landscape where social media platforms suddenly become arbiters of perceived truth and conduits of political clout—where the companies that make the devices, operating systems, and browsers on which we live incrementally more of our lives are responsible for everything from waking us up in the morning to helping us maintain our mental health. The more this perpetuates, the more of a responsibility design becomes, and the harder we must work to serve the needs of real, fallible people.
Design, then, leaves us with a choice. As the interface between an underlying system and the model of that system in someone’s head, we can either introspect our system through the lens of humanity, or we can attempt to contort and convince and persuade humanity to interact with our system in ways we desire. The former is, simply, our job; but it’s hard, and we’ll face battles and make enemies, and we’ll likely get pushed back from many sides. The latter perpetuates the rather dystopian notion that humanity—including our cognitive faculty and even our very concept of self—is subservient to technology. That being hooked on apps, feeling the dopamine squirt of notifications and generally being the good user is simply the price of admission into a world of possibilities. At least, a world of possibilities if you’re a well-off white person in or around a major city with a decent tech hub and own a £1,000 phone.
Design should exist to serve humanity. Real humanity. Not some essence or abstraction of humanity, not the privileged, amalgamated assumptions of four white people sitting down to solve the next mundane problem in their lives. But real, fallible, vulnerable human beings. Humans whose daily struggles and sacrifices in the face of a world and an industry that shows them nothing but apathy and disdain deserve more than to just being a casual footnote on a persona, or a box-checking exercise in our research.

View all my reviews

Thursday, October 04, 2018

Site Reliability Engineering: How Google Runs Production SystemsSite Reliability Engineering: How Google Runs Production Systems by Betsy Beyer
My rating: 5 of 5 stars

A wonderful book to learn how to manage websites so that they are reliable.

Some good random extracts from the book.

Site Reliability Engineering
1. Operations personnel should spend 50% of their time in writing automation scripts and programs.
2. the decision to stop releases for the remainder of the quarter once an error budget is depleted
3. an SRE team is responsible for the availability, latency, performance, efficiency, change management, monitoring, emergency response, and capacity planning of their service(s).
4. codified rules of engagement and principles for how SRE teams interact with their environment—not only the production environment, but also the product development teams, the testing teams, the users, and so on
5. operates under a blame-free postmortem culture, with the goal of exposing faults and applying engineering to fix these faults, rather than avoiding or minimizing them.
6. There are three kinds of valid monitoring output:
Alerts: Signify that a human needs to take action immediately in response to something that is either happening or about to happen, in order to improve the situation.
Tickets: Signify that a human needs to take action, but not immediately. The system cannot automatically handle the situation, but if a human takes action in a few days, no damage will result.
Logging: No one needs to look at this information, but it is recorded for diagnostic or forensic purposes. The expectation is that no one reads logs unless something else prompts them to do so.
7. Resource use is a function of demand (load), capacity, and software efficiency. SREs predict demand, provision capacity, and can modify the software. These three factors are a large part (though not the entirety) of a service’s efficiency.

SLI - Service Level Indicator - Indicators used to measure the health of a service. Used to determine the SLO and SLA.
SLO - Service Level Objective - The objective that must be met by the service.
SLA - Service Level Agreement - The Agreement with the client with respect to the services rendered to them.

Don’t overachieve

Users build on the reality of what you offer, rather than what you say you’ll supply, particularly for infrastructure services. If your service’s actual performance is much better than its stated SLO, users will come to rely on its current performance. You can avoid over-dependence by deliberately taking the system offline occasionally (Google’s Chubby service introduced planned outages in response to being overly available),18 throttling some requests, or designing the system so that it isn’t faster under light loads.

"If a human operator needs to touch your system during normal operations, you have a bug. The definition of normal changes as your systems grow."

Four Golden Signals of Monitoring
The four golden signals of monitoring are latency, traffic, errors, and saturation. If you can only measure four metrics of your user-facing system, focus on these four.

Latency: The time it takes to service a request. It’s important to distinguish between the latency of successful requests and the latency of failed requests. For example, an HTTP 500 error triggered due to loss of connection to a database or other critical backend might be served very quickly; however, as an HTTP 500 error indicates a failed request, factoring 500s into your overall latency might result in misleading calculations. On the other hand, a slow error is even worse than a fast error! Therefore, it’s important to track error latency, as opposed to just filtering out errors.
Traffic: A measure of how much demand is being placed on your system, measured in a high-level system-specific metric. For a web service, this measurement is usually HTTP requests per second, perhaps broken out by the nature of the requests (e.g., static versus dynamic content). For an audio streaming system, this measurement might focus on network I/O rate or concurrent sessions. For a key-value storage system, this measurement might be transactions and retrievals per second.
Errors: The rate of requests that fail, either explicitly (e.g., HTTP 500s), implicitly (for example, an HTTP 200 success response, but coupled with the wrong content), or by policy (for example, "If you committed to one-second response times, any request over one second is an error"). Where protocol response codes are insufficient to express all failure conditions, secondary (internal) protocols may be necessary to track partial failure modes. Monitoring these cases can be drastically different: catching HTTP 500s at your load balancer can do a decent job of catching all completely failed requests, while only end-to-end system tests can detect that you’re serving the wrong content.
Saturation: How "full" your service is. A measure of your system fraction, emphasizing the resources that are most constrained (e.g., in a memory-constrained system, show memory; in an I/O-constrained system, show I/O). Note that many systems degrade in performance before they achieve 100% utilization, so having a utilization target is essential.
In complex systems, saturation can be supplemented with higher-level load measurement: can your service properly handle double the traffic, handle only 10% more traffic, or handle even less traffic than it currently receives? For very simple services that have no parameters that alter the complexity of the request (e.g., "Give me a nonce" or "I need a globally unique monotonic integer") that rarely change configuration, a static value from a load test might be adequate. As discussed in the previous paragraph, however, most services need to use indirect signals like CPU utilization or network bandwidth that have a known upper bound. Latency increases are often a leading indicator of saturation. Measuring your 99th percentile response time over some small window (e.g., one minute) can give a very early signal of saturation.
Finally, saturation is also concerned with predictions of impending saturation, such as "It looks like your database will fill its hard drive in 4 hours."

If you measure all four golden signals and page a human when one signal is problematic (or, in the case of saturation, nearly problematic), your service will be at least decently covered by monitoring.

Why it is important to have control over the software that one is using? Why and when it makes sense to roll out one's own framework and/or platform?
Another argument in favor of automation, particularly in the case of Google, is our complicated yet surprisingly uniform production environment, described in The Production Environment at Google, from the Viewpoint of an SRE. While other organizations might have an important piece of equipment without a readily accessible API, software for which no source code is available, or another impediment to complete control over production operations, Google generally avoids such scenarios. We have built APIs for systems when no API was available from the vendor. Even though purchasing software for a particular task would have been much cheaper in the short term, we chose to write our own solutions, because doing so produced APIs with the potential for much greater long-term benefits. We spent a lot of time overcoming obstacles to automatic system management, and then resolutely developed that automatic system management itself. Given how Google manages its source code, the availability of that code for more or less any system that SRE touches also means that our mission to “own the product in production” is much easier because we control the entirety of the stack.
When developed in-house the platform/framework can be designed to manage any failures automatically. There is no external observer required to manage this.
One of the negatives of automation is that humans forget how to do a task when required. This may not be always good.

Google Cherry Picks features for release. Should we do the same?
"All code is checked into the main branch of the source code tree (mainline). However, most major projects don’t release directly from the mainline. Instead, we branch from the mainline at a specific revision and never merge changes from the branch back into the mainline. Bug fixes are submitted to the mainline and then cherry picked into the branch for inclusion in the release. This practice avoids inadvertently picking up unrelated changes submitted to the mainline since the original build occurred. Using this branch and cherry pick method, we know the exact contents of each release."
Note that cherry picking is of specific release branches and not changes in specific branch.

Surprises vs. boring
"Unlike just about everything else in life, "boring" is actually a positive attribute when it comes to software! We don’t want our programs to be spontaneous and interesting; we want them to stick to the script and predictably accomplish their business goals. In the words of Google engineer Robert Muth, "Unlike a detective story, the lack of excitement, suspense, and puzzles is actually a desirable property of source code." Surprises in production are the nemeses of SRE."

Commenting or flagging code
"Because engineers are human beings who often form an emotional attachment to their creations, confrontations over large-scale purges of the source tree are not uncommon. Some might protest, "What if we need that code later?" "Why don’t we just comment the code out so we can easily add it again later?" or "Why don’t we gate the code with a flag instead of deleting it?" These are all terrible suggestions. Source control systems make it easy to reverse changes, whereas hundreds of lines of commented code create distractions and confusion (especially as the source files continue to evolve), and code that is never executed, gated by a flag that is always disabled, is a metaphorical time bomb waiting to explode, as painfully experienced by Knight Capital, for example (see "Order In the Matter of Knight Capital Americas LLC" [Sec13])."

Writing blameless RCA
Pointing fingers: "We need to rewrite the entire complicated backend system! It’s been breaking weekly for the last three quarters and I’m sure we’re all tired of fixing things onesy-twosy. Seriously, if I get paged one more time I’ll rewrite it myself…"
Blameless: "An action item to rewrite the entire backend system might actually prevent these annoying pages from continuing to happen, and the maintenance manual for this version is quite long and really difficult to be fully trained up on. I’m sure our future on-callers will thank us!"

Establishing a strong testing culture
One way to establish a strong testing culture is to start documenting all reported bugs as test cases. If every bug is converted into a test, each test is supposed to initially fail because the bug hasn’t yet been fixed. As engineers fix the bugs, the software passes testing and you’re on the road to developing a comprehensive regression test suite.

Project Vs. Support
Dedicated, noninterrupted, project work time is essential to any software development effort. Dedicated project time is necessary to enable progress on a project, because it’s nearly impossible to write code—much less to concentrate on larger, more impactful projects—when you’re thrashing between several tasks in the course of an hour. Therefore, the ability to work on a software project without interrupts is often an attractive reason for engineers to begin working on a development project. Such time must be aggressively defended.

Managing Loads
Round Robin Vs. Weighted Round Robin (Round Robin, but taking into consideration the number of tasks pending at the server)
Overload of the system has to be avoided by usage of load testing. If despite this the system is overloaded then any retries have to be well controlled. A retry at a higher level can cascade the retries at the lower level. Use jitter retries (retry at random intervals) and exponential retry (exponentially increase the time between the retries) and fail quickly to prevent overload on the already overloaded system.
If queuing is used to prevent overloading of server then sometimes FIFO may not be a good option as the user waiting for the tasks at the head of the queue may have left the system not expecting a response.
If task is split into multiple pipelined tasks then it will be good to check at each stage if there is sufficient time for performing the rest of the tasks based on the expected time that will be taken by the remaining tasks in the pipeline. Implement a deadline propagation.

Safeguarding the data
Three levels of guard against data loss
1. Soft Delete (Visible to user in the recycle bin)
2. Back up (incremental and full) before actual deletion and test ability to restore. Replicate live and backed up data.
3. Purge data (Can be recovered only from backup now)
4. Out of Band data validation to prevent surprising data loss.

Important to
1. Continuously test the recovery process as part of your normal operations
2. Set up alerts that fire when a recovery process fails to provide a heartbeat indication of its success

Launch Coordination Checklist
This is Google’s original Launch Coordination Checklist, circa 2005, slightly abridged for brevity:
1. Architecture: Architecture sketch, types of servers, types of requests from clients
2. Programmatic client requests
3, Machines and datacenters
4, Machines and bandwidth, datacenters, N+2 redundancy, network QoS
5. New domain names, DNS load balancing
6. Volume estimates, capacity, and performance
7. HTTP traffic and bandwidth estimates, launch “spike,” traffic mix, 6 months out
8. Load test, end-to-end test, capacity per datacenter at max latency
9. Impact on other services we care most about
10. Storage capacity
11. System reliability and failover

What happens when:
Machine dies, rack fails, or cluster goes offline
Network fails between two datacenters
For each type of server that talks to other servers (its backends):
How to detect when backends die, and what to do when they die
How to terminate or restart without affecting clients or users
Load balancing, rate-limiting, timeout, retry and error handling behavior
Data backup/restore, disaster recovery
12. Monitoring and server management
    Monitoring internal state, monitoring end-to-end behavior, managing alerts
    Monitoring the monitoring
    Financially important alerts and logs
    Tips for running servers within cluster environment
    Don’t crash mail servers by sending yourself email alerts in your own server code
13. Security
Security design review, security code audit, spam risk, authentication, SSL
Prelaunch visibility/access control, various types of blacklists
14. Automation and manual tasks
    Methods and change control to update servers, data, and configs
    Release process, repeatable builds, canaries under live traffic, staged rollouts
15. Growth issues
    Spare capacity, 10x growth, growth alerts
    Scalability bottlenecks, linear scaling, scaling with hardware, changes needed
    Caching, data sharding/resharding
16. External dependencies
    Third-party systems, monitoring, networking, traffic volume, launch spikes
    Graceful degradation, how to avoid accidentally overrunning third-party services
    Playing nice with syndicated partners, mail systems, services within Google
17. Schedule and rollout planning
    Hard deadlines, external events, Mondays or Fridays
    Standard operating procedures for this service, for other services

As mentioned, you might encounter responses such as "Why me?" This response is especially likely when a team believes that the postmortem process is retaliatory. This attitude comes from subscribing to the Bad Apple Theory: the system is working fine, and if we get rid of all the bad apples and their mistakes, the system will continue to be fine. The Bad Apple Theory is demonstrably false, as shown by evidence [Dek14] from several disciplines, including airline safety. You should point out this falsity. The most effective phrasing for a postmortem is to say, "Mistakes are inevitable in any system with multiple subtle interactions. You were on-call, and I trust you to make the right decisions with the right information. I'd like you to write down what you were thinking at each point in time, so that we can find out where the system misled you, and where the cognitive demands were too high."

"The best designs and the best implementations result from the joint concerns of production and the product being met in an atmosphere of mutual respect."

Postmortem Culture

Corrective and preventative action (CAPA) is a well-known concept for improving reliability that focuses on the systematic investigation of root causes of identified issues or risks in order to prevent recurrence. This principle is embodied by SRE's strong culture of blameless postmortems. When something goes wrong (and given the scale, complexity, and rapid rate of change at Google, something inevitably will go wrong), it's important to evaluate all of the following:

What happened
The effectiveness of the response
What we would do differently next time
What actions will be taken to make sure a particular incident doesn't happen again

This exercise is undertaken without pointing fingers at any individual. Instead of assigning blame, it is far more important to figure out what went wrong, and how, as an organization, we will rally to ensure it doesn't happen again. Dwelling on who might have caused the outage is counterproductive. Postmortems are conducted after incidents and published across SRE teams so that all can benefit from the lessons learned.

Decisions should be informed rather than prescriptive, and are made without deference to personal opinions—even that of the most-senior person in the room, who Eric Schmidt and Jonathan Rosenberg dub the "HiPPO," for "Highest-Paid Person's Opinion"


View all my reviews

Friday, June 29, 2018

Software Requirement Patterns by Stephen Withall

Software Requirement PatternsSoftware Requirement Patterns by Stephen Withall
My rating: 4 of 5 stars

A very nice book for all software business analysts, product owners, developers and architects. It provides very good guidelines on how different types of requirements should be specified or captured what aspects need to be considered when specifying or capturing the requirements.

The list of patterns with a brief explanation where not clear are as under:

Fundamental Requirement Patterns
a) Inter-System Interface Requirement Pattern - This pattern covers how to capture the details of data to be exchanged between various applications in that will be part of the system.
b) Inter-System Interaction Requirement Pattern - This pattern covers how to capture the details of how data will be passed between the various applications that will part of the system.
c) Technology Requirement Pattern - If there is a requirement of usage of specific technologies this pattern tells how these requirements of technology can be specified.
d) Comply-with-standard Requirement Pattern - This pattern shows how one should specify the regulartory requirements. It may not be sufficient to say adhere to this standards or guidelines. It will be necessary to be specific in many scenarions.
e) Refer-to-Requirements Requirement Pattern - This is reference to other requirements gathered previously.
f) Documentation Requirement Pattern - This specifies the kinds of documenation that need to be prepared as part of the development process.
Information Requirement Patterns
a) Data Type Requirement Pattern - At a fundamental level this indicates what kind of data types need to be support. This becomes important as once can specify how amount or date time fields should be captured, stored and displayed in the application.
b) Data Structure Requirement Pattern - This specifies the functional structures that need to be created in the system using the ID pattern and the Data Type Requirement.
c) ID Requirement Pattern - This pattern indicates the requirements of creation, maintenance and format of ID fields in the system.
d) Calculation Formula Requirement Pattern - This pattern covers on how to specify requirements of various formulae to be used in the application.
e) Data Longevity Requirement Pattern - This pattern covers how the requirements of retention of data in the system needs to be specified.
f) Data Archiving Requirement Pattern - This pattern covers how the requirements how the data in the system needs to be archived is to be specified.
Data Entity Requirements Pattern
a) Living Entity Requirement Pattern - This patterns covers how details of entities which have a life in the system need to be specified. Some examples are customers, accounts etc.
b) Transaction Requirement Pattern - This pattern covers how requirements should be specified for events attached to one or more Living Entity in the system.
c) Configuration Requirement Pattern - This pattern covers how requirements for the different business configuration parameters should be specified.
d) Chronicle Requirement Pattern - This patterns covers how requirements for tracking of changes to the different Living Entity or Transaction should be captured and stored.
e) Information Storage Infrastructure - This pattern specifies
User Function Requirement Patterns
a) Inquiry Requirement Pattern - This pattern gives details of how to specify requirements for inquiries to be made in the system. These are interactive inquiries as opposed to static report requirements which is the next pattern.
b) Report Requirement Pattern - This specifies the reports that need to be generated in the system along with the filter criteria, export requirement, and synchronous and asynchronous requirement.
c) Accessibility Requirement Pattern - This could be regulatory requirement or something specific to the type of users or requirements due to the environmental conditions.
d) User Interface Infrastructure - This pattern covers how to specify the infrastructures on which the User Interface should be supported.
e) Reporting Infrastructure - How configurable reports need to be and how they need to be generated and how they need to be delivered.
Performance Requirement Patterns
a) Response Time Requirement Pattern
b) Throughput Requirement Pattern
c) Dynamic Capacity Requirement Pattern - This pattern covers specification of dynamic load requirement like increased number of users, transactions etc.
d) Static Capacity Requirement Pattern - This pattern covers the specification of static capacity requirements such as storage space.
e) Availability Requirement Pattern - This pattern covers how the availability of the system should be specified and what factors should be considered when specifying this.
Flexibility Requirement Patterns
a) Scalability Requirement Pattern - This pattern covers specifying the requirements of how much data/numbers that needs to be supported by the system.
b) Extendability Requirement Pattern - This pattern covers specifying the requirements of how extendible the system must be to support different ways of doing something particular. E.g. adding different types of payment channels or sending notification to user etc.
c) Unparochialness Requirement Pattern - The system should be able to support working with multiple currencies, or with a flavour local to the geographic location or department of the organization. This applies when the application is installed at different locations.
d) Multiness Requirement Pattern - This is similar to unparochialness, but applicable when all of this has to be achieved in a single installation of the application. It may also mean the way in which the ID values need to be allocated in the system.
e) Multi-Lingual Requirement Pattern -
f) Installability Requirement Pattern - This pattern covers how the installation requirements of the application should be specified if it is necessary.
Access Control Requirement
a) User Registration Requirement Pattern - This pattern addresses how a user should be made known to the system.
b) User Authentication Requirement Pattern - This pattern addresses how a user should be validated by the system.
c) User Authorization Requirement Pattern - This pattern addresses how the privileges of the users in the system be controlled.
d) Specific Authorization Requirement Pattern - This pattern addresses how generic requirements like denial-by-default rule etc. should be specified.
e) Configurable Authorization Requirement Pattern - This pattern addresses how to specify requirements when it is required to let the administrators or other authorized personnel to manage this.
f) Approval Requirement Pattern - This pattern addresses how to specify any approval process required in the system.
Commercial Requirement Patterns
a) Multi-Organization Unit Requirement Pattern - This pattern addresses how requirements of Organizations with Multiple units needs to be specified.
b) Fee/Tax Requirement Pattern - This pattern specifies the fee/tax requirement that the system must calculate, report on, or levy. This requirement pattern can also be used (with only slight variations) to specify a discount on an amount a party is charged.

View all my reviews

Thursday, January 18, 2018

Mastering Docker by Russ McKendrick

Mastering Docker - Second Edition: Master this widely used containerization toolMastering Docker - Second Edition: Master this widely used containerization tool by Russ McKendrick
My rating: 3 of 5 stars

A good book to read to understand Docker and to start off with it. Definitely not one which will give one a mastery over Docker. In depth coverage required for mastering Docker is missing.


View all my reviews

Friday, September 01, 2017

Infrastructure as Code: Managing Servces in Cloud by Kief Morris

Infrastructure as Code: Managing Servers in the CloudInfrastructure as Code: Managing Servers in the Cloud by Kief Morris
My rating: 4 of 5 stars

A very nice book to start managing infrastructure as a code. Has very helpful tips on the strategies to be adopted when doing this.

The book does not cover any one tool (Chef, Puppet or any other tool), it merely gives the techniques to be followed to ensure that this is a success.

Definitely a very good read.

View all my reviews

Friday, April 21, 2017

8th standard Mathematcis to estimate software development efforts

Reading this line "“Well, Feature X took 75 KLOC to implement, and we think that Feature Y will take twice as much effort. So, we’re currently 15 KLOC in after a week, and figure it will take 10 total weeks.” Or something like that." at https://dzone.com/articles/alternatives-to-lines-of-code-1 reminded me of the day, about 20 years ago arguing with my manager on why I will not use the Function Point Analysis to determine the effort that will be required to execute a project.

The situation was that there was a team in our organization which was writing an application in Visual Basic (good old days of VB glory). Another client wanted to develop another application (no way related to the application that was being developed in VB) in Java. The timesheets of the team members working on the VB project were being maintained. The function points being addressed by the VB project was being maintained. Given these inputs my manager wanted me to
1. Calculate the productivity of the VB team in terms of Function Points
2. Estimate the function points for the new application to be developed in Java
3. Translate the productivity computed in VB to Java (using industry standards ahem..)
4. Use this computed productivity to calculate the effort required to execute Java project.

The points to be noted are as under
1. The VB team members were filling in timesheets only because there was a mandate to fill them in. A sample timesheet would have shown
     Developer A - 8 hours of Development
     Developer B - 8 hours of Design
     Developer C - 8 hours of Meeting (just joking)
     And the productivity of the VB developers was based on these timesheet numbers
2. There were no benchmarks of productivity of the team members in the Java team in any technology.
3. The application being developed in VB was no way similar in functionality to the application to be developed in Java.
4. One can think of an productivity translation from VB to Java only if it were the same or at least similar application and similar teams.

Despite all of these facts my manager refused to accept that what I could give from my gut feeling based on the understanding of the functionality and its complexities would be better than what can be done in terms of function points.

I will not be surprised if it still happens today.

Sunday, October 02, 2016

Your code as a Crime Scene by Adam Tornhill

Your Code As a Crime Scene: Use Forensic Techniques to Arrest Defects, Bottlenecks, and Bad Design in Your ProgramsYour Code As a Crime Scene: Use Forensic Techniques to Arrest Defects, Bottlenecks, and Bad Design in Your Programs by Adam Tornhill
My rating: 4 of 5 stars

A very different way of looking at problems in the code. The author suggests various non-traditional means of identifying problems in software development.

Number of times a code has changed, code churn (number of lines added, removed), number of developers working on a single piece of code, code changing together.

The primary requirement for all of this that the version control should be adhered to correctly. Every developer should have a personal login Id and should use it to checkin at regular intervals.

Many of the techniques provided in the book are available in the form the tool from https://github.com/adamtornhill/code-.... These are only starting tools. These need to be coupled with other tools like d3.js https://d3js.org/ to get a good representation of the status of the code.

A must read for all software developers.

View all my reviews

Friday, June 03, 2016

How to Stop Sucking and be Awesome Instead by Jeff Atwood

How to Stop Sucking and Be Awesome InsteadHow to Stop Sucking and Be Awesome Instead by Jeff Atwood
My rating: 4 of 5 stars

Another nice book of blog entries from Jeff Atwood.

First few blogs are about how one should determine at a very early age if one can program or not and should drop out of a programming career if one is not. He speaks about "sheep that can program and goats that cannot program" should be separated out early in the career so that software can become better.

Some of the key observations that I liked are "You have to truly believe, as a company, and as peers, that crucial innovations and improvements can come from everyone at the company at any time, in bottom-up fashion - they aren't delivered from on high at scheduled release intervals in the almighty master plan.

In another blog he speaks about how important it is to persuade others to do something. He refers to a set of dialog from the movie based on Idi Amin. Idi Amin is speaking to his trusted aide a Scottish Doctor.
Idi Amin: I want you to tell me what to do!
Garrigan: "You want me to tell you what to do?
Amin: Yes You are my advisor. You are the only one I can trust here. You should have told me not to throw the Asians out in the first place!
Garrigan: I did!
Amit: But you did not persuade me, Nicholas. You did not persuade me!

Not a very atypical dialog one is likely to have with either one's manager or client. :-)

Another important advice with which I cannot agree more since I have given the same advice to many others who have asked me for my opinion. "Whatever project you are working on, consider it an opportunity to learn and practice your craft. It is worth doing because, well it is worth doing. The journey of the project should be its own reward regardless of whatever happens to lie at the end of that journey."
The corollary is another thing that I keep stressing on; "Never get attached to a project. Execute the project to the best of the abilities, learn along the way and if for some reason beyond your control the project fails or does not see the light of the day, so be it."

Speaking on merit based growth in an organization Jeff has to say the following "Remove barriers that rob people in management and in engineering of their right to pride of workmanship. This means [among other things] abolishment of annual or merit rating and management by objectives. "Even people who think of themselves as Deming-ites have trouble with this one. They are left gasping. What the hell are we supposed to do instead? Deming's point is that MBO and its ilk are copouts. By using simplistic extrinsic motivators to goad performance, managers excuse themselves from harder matters such as investment, direct personal motivation, thoughtful team formation, staff retention, and ongoing analysis and redesign of work procedures. Our point here is somewhat more limited: Any action that rewards team members differentially is likely to foster competition. Managers need to take steps to decrease or counteract this effect.

In one blog Jeff compares F-86 and MIG-15s. The latter was far more superior to the former, but the fighter pilots preferred the former. The difference was that F-86 had Hydraulic flight controller compared to the manual flight controller of MIG-15. This meant that each maneuver increased the fatigue of the MIG-15 pilot even though he might have out-maneuvered the F-86 pilot. The F-86 could maneuver quicker as compare to the MIG-15 as he was less fatigued and this tilted the pilots to favour F-86 despite its limited abilities. Jeff calls this the Boyd's Law of Iteration which states "Speed of iteration beats quality of iteration". Jeff argues the same is true for software development. Although he says in other places that quality cannot be sacrificed beyond a point.

There is a whole set of blogs in User Interface and Usability. One book the author highly recommends is Don't make me think by Steve Krug https://www.amazon.in/Dont-Make-Think... and another one is Rocket Surgery made easy again by Steve Krug https://www.amazon.in/Rocket-Surgery-....
He speaks about the Fitts law which states "Put all commonly accessed UI elements on the edges of screen. Because the cursor automatically stops at the edges, they will be easier to click on. Make clickable areas as large as you can. Larger targets are easier to click on". One should not ignore the corollary of rule which would read "Make all the clicks that the user must be kept safe from as difficult as possible". Jeff refers to this principle as the "seat ejector" button. This button should be easy to find in an emergency, but should not be place such that the pilot ends up turning this on instead of the navigation lights. Buttons like delete all my mails and such should be available, but should be placed such that the no user would click it by mistake.

Speaking on importance of saying not to demands, Jeff says "It is easy to dismiss Just say No as a negative mindset, but I think it is a healthy and natural reaction to observation that optimism is an occupational hazard of programming". Cannot agree with him more as I have been pulled up for saying no many a time in my career.

Speaking about usability Jeff argues that most users of the applications do not progress beyond the intermediary stage. He argues that most move from the novice to Intermediary stage quite quickly or drop off as an user of the system if they find it too difficult to use. Once they reach this stage they remain in this stage for a long time and only a very few move to the expert level. Given this he argues that software should be targeted towards these users rather than at the novices or the experts. He states that most marketing people would advocate making software for the novice as these would be the ones that the marketing people would encounter most of the times, while the software developer would want to address the experts as it is likely that they would be geeks themselves and would want maximum flexibility and most features.

Writing on security and hacking the author says that today to hack into a site one needs social skills and not technical skills. Technology has advanced to a level where hacking into a website has been made difficult enough, but people are not inure to social engineering and that it is much easier way of hacking into a system.

On the whole a very good read.




View all my reviews

Friday, May 20, 2016

Eloquent JavaScript by Marijn Haverbeke

Eloquent JavaScriptEloquent JavaScript by Marijn Haverbeke
My rating: 3 of 5 stars

A good introductory book for beginners in JavaScript. Any beginner can to through the book and start writing real good JavaScript.

The author introduces modules, regex, DOM, the HTTP Protocol and also gives direction to write two different fun program with JavaScript.

Some of the chapters could be seen as outdated given the changes that have come into JavaScript, but nevertheless still a relevant book.

After this has been read one can read "JavaScript - The Good Parts" by Douglas Crockford.

View all my reviews

Thursday, December 31, 2015

Release It!: Design and Deploy Production-Ready Software by Micheal T. Nygard.

Release It!: Design and Deploy Production-Ready Software (Pragmatic Programmers)Release It!: Design and Deploy Production-Ready Software by Michael T. Nygard
My rating: 4 of 5 stars

Introduction
The author asserts that software of today is built for passing the tests of the QA and not for the rigours of the Production environment. The author provides tips to design systems which will withstand the assaults it will have to face in the Production environments.
The author states that most decisions made upfront are the decisions that are the ones that impact the system the most, are most difficult to reverse or change and ironically these are the ones that are taken when the knowledge about the required system is minimal.
The author ironically states decrees such as "Use EJB container-managed persistence!", "All UIs shall be constructed with JSF!", "All that is, all that was and all that shall ever be lives in Oracle!" are given by architects in ivory towers.

In the stability section the author speaks about how to create and maintain stable systems in this section. The first example the author gives is of an airline company which had the following code:

package com.example.cf.flightsearch;
. . .
public class FlightSearch implements SessionBean {
    private MonitoredDataSource connectionPool;
    public List lookupByCity(. . .) throws SQLException, RemoteException {
        Connection conn = null;
        Statement stmt = null;
        try {
            conn = connectionPool.getConnection();
            stmt = conn.createStatement();
            // Do the lookup logic
            // return a list of results
        } finally {
            if (stmt != null) {
                stmt.close();
            }
            if (conn != null) {
                conn.close();
            }
        }
    }
}
Which looks and feels good. But if the stmt.close() ever throws an exception the conn.close() will never be called, resulting in connections leaking from the connection pool leading to all the connections in the pool being used up.

The author suggests that one should be prepared for as many points of breakages as possible. Tight coupling between systems leads to cascading failures. To avoid this there should be loose coupling between systems. As a corollary calls across systems should be asynchronous. This is not always possible and where possible complicates communication. So one needs to take a proper call on where to have asynchronous processing and where not to.

In chapter 4 the author discusses anti-patterns that lead to failures. The first anti-pattern is that all points of integration are fragile and can lead to failures. It is highlighted that most connections are based on TCP/IP. In TCP IP the first step is a three way handshake to setup the connection between the two systems that need to communicate. The first step is for the requestor to send a SYN packet. This has to be acknowledged by a SYN-ACK packet from the listener and finally the requestor sends a SYN packet to complete the three way handshake and establish the connection. If there is no listener then the failure is quick as the OS responds with RESET packet telling the requestor that its request does not have listener. This is a manageable situation. But if the listener is slow then the request will languish in the listen queue till it is timed out. The typical timeout is in minutes. This means that the requestor can wait for a long time before realising the problem.
The classic example of firewall killing the TCP connection between the application server and database server due to long time idle connection is quoted as an example.
The next example is about how it is difficult to timeout HTTP Connections in Java.
It is stressed again and again that it is better to be cynical than optimistic when developing software. Be prepared for the worst.
It is suggested that circuit breakers, i.e. stopping to retry a transaction after a particular number of failed attempts and/or maintaining the status of the underlying layer and deciding to not invoke the layer if there is a problem and timing out after waiting for a reasonable amount of time for the underlying layer to respond are two key mechanisms to avoid cascading problems from one layer to another.
The storing of large datasets in the session is highlighted as one of the more frequent ways of running out of memory. It is suggested that either the session be kept light or Softreference be used for storing large datasets to prevent out of memory errors.
The author rightly points out that the usage of synchronized keyword can be dangerous in a highly concurrent environment.
The advice is also to test third party libraries for breakability.
The author coins a new word called "Attack of Self Denial" where an event is published which leads to a flood of requests to the specified application. E.g. the news of a deep discount on a product for a retailer could be a cause for "Attack of Self Denial". One needs to be prepared beforehand to handle such situations better.
A very good suggestion is that if it is not possible to build a shared nothing architecture then limit the number of systems sharing the resources. E.g. instead of sharing the sessions with all the application servers sharing it amongst two application servers so that the replication factor is limited.
One key point that the author brings about is most systems "treat the database with far too much trust." and this is the major cause of problems in most systems. The author illustrates this with an example of how making an unbounded query resulted in continuous crashes at a retailer. The author suggests that always limit the results fetched from the database as a precaution.

The Stability Patterns
----------------------
In this the author lists the patterns that will help the system be more stable.
1. Timeout: Use timeouts whenever interacting with a third party, especially when this involves some form of network, even though it may be within a LAN. It is a fail fast pattern to be used along with a circuit breaker, where if a few requests timeout then the resource is marked as down till it is found to be good again. The retry to check if the resource is good can be done at a regular interval suitable for the resource, i.e. delay the retry.
2. Circuit Breaker: Akin to the electrical circuit breaker a software circuit breaker prevents the entire software from collapsing under stress by stopping requests to the faulty interface. The users may see errors if this happens to be a crucial interface, but this is better than the user not being able to use the whole system. Typically if an interface frequently times out or fails frequently then the circuit breaker can mark this interface as broken for sometime. After a suitable amount of time it can retry the interface and if found functional it can close the circuit once again enabling the execution of the specific interface. All opening and closing of circuit breaker should be logged and made visible to the operations team so that they are aware of the change in the status.
3. Bulkheads: Bulkheads are compartments in a ship which prevent the ship from sinking if there is a damage to the hull. Each bulkhead stops the water from entering beyond it. Similarly use of multiple servers to deploy applications is one form of bulkhead. If the application is compartmentalized so that impact to one compartment does not impact the other is creating bulkheads in application. One example quoted is that of the airlines where ticketing system, flight status systems, flight search system, checkin system could all be deployed separately so that one does not interfere with the other.
Another example is if there are two systems which require the same service and if both the systems are critical, it makes sense to have separate setup of the common service for the two systems. Problem access of the common service in one system will not impact the service access of the other system.
Grouping pool of thread for specific purpose in a single process will ensure that problem in one thread pool does not prevent the process from servicing other types of requests.
The negative side of bulkheads is that it can make optimization of resource usage difficult. One would potentially have to provide more capacity than actually required.
4. Steady State: Maintaining a steady state of the systems is very important. Any kind of fiddling with the system for any reason can lead to instability. At the same time to maintain steady state some cleaning up is required. Log files will be generated by the applications and it is important to have a process that will keep removing the log files at the same rate or greater than the rate of generation. Similarly archival of records in a database is important to ensure that the queries on the database continue run consistently.
It is important to ensure that one has a finite, controlled number of entries in the in memory cache. Use an LRU or LFU mechanism to keep clearing the cache if it is expected to keep growing beyond known values.
5. Fail Fast: Quickly failing a request is very important to the health of the transaction. One should upfront have the statuses of all the external systems before beginning to process a transaction and if any of the external system is in a state which will mean that the transaction will fail then it is better to fail the transaction immediately. This will ensure that no compute power is wasted in processing doomed transactions.
6. Handshaking: It is important to have handshaking between any two systems so that the server process has the ability to state that it has its hands full and cannot respond and the client does not waste time trying to make a request which is going to take the server a long time to respond. This helps in failing fast.
7. Test Harness: A test harness should be able to emulate bizarre problems, like accepting a connection, but not sending any response, resetting the connection without ever accepting it and so on, not responding for a very very long time, send out large amounts of data as response. Testing the against such a test harness will help test how the system will behave under unexpected conditions.
8. Decoupling Middleware: A middleware typically helps shielding the requestor from the nitty-gritties of the server and also from the failures of the server. It helps decoupling two systems while integrating them.

The author very rightly concludes that "Sadly, the absence of a problem is not usually noted. You might be salvaging a badly botched implementation in which case you now have an opportunity to look like a hero. On the other hand, if you’ve done a great job of designing a stable system from the beginning, it’s unlikely that anyone will notice your system’s lack of downtime. That’s just the way it is. Deliver an unbreakable system, and users will surely go on to
complain about something else. That’s just what users do. In fact, with a system that never goes down, the users will most likely complain that it’s slow. Next, you’ll look at capacity and performance and how to get the most out of your resources."

In a case study it is illustrated how usage of sessions killed the application. The bots and the regular users increased the number of sessions far beyond what the system could handle and site crashed. This was later resolved by supporting session through URL rewriting so that no new sessions are created by the bots and also by creating a throttling mechanism to control the total number of sessions in the system. The key learning is that the performance test only tested for happy paths and never for situations like bots hitting the site.
When planning for capacity it is important to ensure that the software written is optimal and has minimum wastage. If this is not done it would lead to increasing costs of resources required to run the application. As an example if an HTML page has 1K of junk data, this will translate into 1GB of extra bandwidth usage if there are a million requests to this page. The cost of resources multiplies as the usage of the application increases.

Some good patterns to follow are:
1. Pool resources, size them properly and monitor them.
2. Use caching, limit the maximum memory that can be used by the cached objects and monitor the hit ratio.
3. Precompute whatever is possible and recompute only when absolutely necessary.
4. Tune Garbage Collection

Some Network points
1. Servers in production tend to be multi-homed and it is important to bind the applications to the right home to prevent security issues.
2. Given the above scenario it becomes important to correctly make the network routing scenarios.
3. Use Virtual IPs where native clustering of applications is not possible. Applications need to be written keeping in mind that this will be the case in production systems.

Some Security aspects:
1. Follow the principle of "least privilege". This states that every action should done with the least privilege required to execute the action. Rnu each application with its own user so if one application is compromised it is only that application and none of the others.
2. Ensure that the passwords use to access other services are secured properly. Ensure that the memory dumps of the processes will not reveal the passwords. Keep the passwords away from the installation directory.

Some Availability Aspects
1. The cost of the a system grows exponentially with the required availability. Availability should be defined realistically, not idealistically.
2. The SLAs should be well defined and measurable. SLAs should be defined by features and dependent on 3rd party SLAs available. The location from where the application is accessed also matters.
3. Load Balancing and Reverse Proxies should be used to balance the load across the multiple servers and across the various tiers.
4. Clustering will be required in scenarios where the servers need to communicate with each other to exchange some data.

To ensure reliability of the system the topology of the QA environment should be same as that of the Production although the capacity may be far lower.
Configuration of the application and environment related configuration should be separated out.
Application should be able to announce if it has not started properly.
Provide command line options to configure the systems. GUI can be used when sufficient time is at hand and automation is not required.

Every system needs to be transparent, i.e. it needs to show what it is using and what it is doing. Without this information it is very difficult to manage the system. While it is necessary to know the status of the individual parts, it is important to also know the status across all the parts of the system. This helps in analysing any problem that is manifesting in the system.
It is not necessary to log the stack trace of a business exception like a validation error which states a mandatory parameter was not entered. It is vital to log the stack trace in case a non business exception occurred.
It is important to have a network separate from the production data network for monitoring traffic.
A good monitoring system provide visibility to to business outcome and not just technical parameters.

A very good comparison between crystals and tight coupling in software design.
"A cluster of objects that can exist together only in a tight collaboration resembles a crystal in a metal. The objects stay together in a tightly bound relationship, just as the atoms in a crystal are tightly bound. In metal, small crystals mean greater malleability. More malleable metals recover from stress better. Large crystals encourage crack formation. In software, large “crystals” make it harder to change the software. When objects in one grain participate in multiple collaboration patterns, they bridge two crystals, forming a larger grained crystal—further reducing the malleability of the software.
There is no limit to how far this region of tightly bound crystals can spread. In the extreme case, the crystal grows until it is the boundary of the application. When that happens, every object suits exactly one purpose to which it is supremely adapted. It fits perfectly into place and ultimately relates to every other object. These crystal palaces might even be beautiful in a baroque sort of way. They admit no improvement, in part, because no incremental change is possible and, in part, because nothing can be moved without moving every other object. These tend to be dead structures. Developers tiptoe through crystal palaces, speaking in hushed tones and trying not to touch anything."


View all my reviews

Friday, December 11, 2015

Ship It! by Jared Ricardson and William Gwaltney Jr.

Ship It!Ship It! by Jared Richardson
My rating: 3 of 5 stars


A collection of lessons learned by various developers in the trenches. The book starts off with a quote of Aristotle "We are what we repeatedly do. Excellence, then, is not an act, but a habit.". The book strengthens this argument by stating "Extraordinary products are merely side effects of good habits.". So the first tip of the book is "Choose your habits". Do not follow something just because it is popular or well known or is practised by others around you.

The author says that there are three aspects that one needs to pay attention to:
  1. Techniques: How the project is developed? I.e. Daily meetings, Code Reviews, Maintaining a To Do List etc.
  2. Infrastructure: Tools used to develop the project. I.e. Version Control, Build Scripts, Running Tests, Continuous Build etc.
  3. Process: The process followed in developing the applications. Propose Objects, Propose Interfaces, Connection Interfaces, Add Functions, Refactor Refine Repeat.

Tools and Infrastructure

The author highlights the need for a proper tool for Source Control Management. The author also issues a warning that the right tool should be chosen. A tool should not be chosen because it is backed by a big ticket organization. Vendors would push for "supertools", but one needs to exercise discretion when choosing between the tools.

Good Development Practices

  1. Develop in a Sandbox, i.e. changes of one developer should not impact the other until the changes are ready.
  2. Each developer should have a copy of everything they need for development, this includes web server, application server, database server, most importantly source code and anything else.
  3. Once all the changes by the developer are finished they should check it in to the Source Control so that the others can pick up and integrate it with their code and make any changes they need to make to integrate.
  4. The checked in changes should be fine grained.

Tools Required for ensuring Good Development Practices

  1. SCM
  2. Build Scripts
  3. Track Issues

What to keep in SCM?

  1. While it can be debated whether runtimes like Java need to be kept in the SCM, it is important that all the third party libraries (jars, dlls) and configuration templates be available in the SCM. Note that configuration templates need to be available as the contents itself can change from environment to environment.
  2. Anything that is generated as part of the build process (jars, dlls, exes, war) should not be stored in the SCM.

What a Good SCM should offer

  1. Ensure that the usage of SCM is painless to the developers. The interactions with the SCM should be fast enough to ensure that the developers do not hesitate to use it.
  2. A minimal set of activities that should be supported by the SCM are
  • Check out the entire project.
  • Look at the differences between your edits and the latest code in the SCM.
  • View the history for a specific file—who changed this file and when did they do it?
  • Update your local copy with other developers’ changes.
  • Push (or commit) your changes to the SCM.
  • Remove (or back out) the last changes you pushed into the SCM.
  • Retrieve a copy of the code tree as it existed last Tuesday.

Script the Build


Once the required artefacts are checked out from the SCM it should be possible for any developer to run a script and have a working system (sandbox) of her own to work on. For this one needs a Build Script. This should be a completely automated build requiring no manual intervention or steps. This build script should be outside of the IDE so that it can be used irrespective of the IDE being used. The IDE could use the same script for local builds.
Once the one step/command build script is ready, automate the build. Ideally everytime a code is checked in the following should be done.
  1. Checkout the latest code and build
  2. Run a set of smoke tests to ensure that the basic functionality is not broken.
  3. Configure the build system to notify the stakeholders of new code checked, the build and the test results.

This is Continuous Integration

Tracking the Issues

It is important to track the issues that are reported for the application so that they can be tracked and fixed.
At a bare minimum one needs to know the following about an issue:
  • What version of the product has the issue?
  • Which customer encountered the issue?
  • How severe is it?
  • Was the problem reproduced in-house (and by whom, so they can help you if you’re unable to reproduce the problem)?
  • What was the customer’s environment (operating system, database, etc.)?
  • In what version of your product did the issue first occur?
  • In what version of your product was it fixed?
  • Who fixed it?
  • Who verified the fix?
Some more that will help in the long term
  • During what phase of the project was the bug introduced?
  • The root cause of the bug
  • The sources that were changed to fix the problem. If the checkin policy demands that the checkin comment indicate the reason for the fixes, then it should be possible to correlate the checkin with the issue that they fixed or requirement that they addressed.
  • How long did it take to fix the error? (Time to analyze, Fix, Test)

Some warning signs that things are not OK with the issue system
  • The system isn't being used.
  • Too many small issues have been logged in the system
  • Issue-related metrics are used to evaluate team member performance.

Tracking Features

Just as it is important to track the issues, it is important to track the features that have been planned for the application.
The system used to track issues may also be used to track the features as long as it provides the ability to identify them separately.

Test Harness

Have a good Test Harness which can be used to run automated tests on the system.
  1. Use a standard Test Harness which can generate all the required reports.
  2. Ensure that every team member uses the same tool.
  3. Ensure that the tool can be run from the command line. This will enable driving it from an external script or a tool.
  4. Ensure that the tool is flexible to test multiple types of applications and not specific to a particular type.

Different types of testing needs to be planned for
  1. Unit Testing - Testing small pieces of code. This forces the developers to break up the code into smaller pieces. This makes is easier to maintain and understand, reduces copy paste, ensures that overall functionality is, if at all, minimally impacted by refactoring.
  2. Functional Testing - Testing all the functions of the application.
  3. Performance Testing - Testing the application to ensure that the application is performing within acceptable limits and meets the SLAs.
  4. Load Testing - This is similar to the Performance Testing. The goal of this is to ensure that the application does not collapse under load.
  5. Smoke Testing - This is a light-weight testing which will test the key functionality of the application. This should be included as part of Continuous Integration so that any breakage in key functionality comes to light very quickly.
  6. Integration Testing - This ensures that the integration of the modules within the application and the integration of the application with the external systems is functioning correctly.
  7. Mock Client Testing - This mocks the client requests and ensures that the client get the right response and within the expected time period.

Pragmatic Project Techniques


Some of the good practices to follow when working in projects are as follows:
  1. Maintaining a list of activities to do. This should be visible and accessible to everybody on the project. Even the client should have visibility to the list so that they are check the speed and prioritize the items in the list. Each item should have a target time. The list should reflect the current status and should not be out of date.
  2. Having Tech Leads in the project is important. The Tech lead should guide the team in the selection and utilization of the technology. Tech lead should be responsible to ensure that the deadlines are realistic. The Tech lead should act as the bridge between the developers and the management. It is an important role to be played by a person with the right temperament.
  3. Coordinating and Communicating on daily basis is very important. Meetings need to be setup on a daily basis. These meetings should be short and to the point, with everybody sharing details of what they are doing and what they plan to do. Team should highlight any problem they are facing. The solutions for these problems should not be part of this meeting, but should happen separately.
  4. Code review is a very crucial part of the project and every piece of code should be reviewed. Some good practices of code review are
    1. Review only a small amount of code at any time
    2. A code should not be reviewed by more than two people
    3. Code should be reviewed frequently, possibly several times a day
    4. Consider pair programming as a continuous code review process.

Tracer Bullet Development


Just like it is possible to fire a Tracer Bullet in the night to track the path before aiming the real bullet, it should be possible to predict the path of the project using the process opted for.

Process

Have a process to follow.
The process followed should not claim exclusivity in success of projects. If it does so, then suspect it.
Follow a process that embraces periodic reevaluation and inclusion of whatever practices work well for the projects.

Executing

  • Define the layers that will exist in the application.
  • Define the interfaces between the layers.
  • Let each layer be developed by a separate team, relying on the interface promised by the adjacent layers.
  • Keep it flexible so that the interface can be changed as it is hard to get the interfaces perfect the first time around.
  • First create the large classes like the Database Connection Manager, Log Manager etc required for each layer, then write the fine grained classes.
  • Collaboration between the teams developing the different layers is key to the success. These collaborations will Trace the Path that the project will take.
  • Do not let an architect sitting in an ivory tower dictate the architecture.
  • It is dangerous to have one person driving the whole project. If this person leaves, the project will come to a standstill.
  • Create stubs, or mock the interfaces of the adjacent layers so that it becomes easy to test.
  • Code the tough and key pieces first and test them before addressing the simpler ones. It may take time to show progress, but when the progress happens it will be very quick.


Common Problems and How to fix Them

What to do when legacy code is inherited?

  1. Build it - Learn to build it and script the build.
  2. Automate it - Automate the build.
  3. Test it - Test to understand what the system does and write automated test cases.
Don't change legacy code unless you can test it.

Some other tips from the chapter

  1. If a code is found unsuitable for automated test, then refactor the code slowly so that it becomes amenable to automated testing.
  2. If a project keeps breaking repeatedly, automated test cases, emulating the user actions will help reduce the incidents.
  3. Ensure that the automated tests are updated with change in code/logic whenever required, otherwise these would become useless.
  4. It is important to have a Continuous Intergration so that the automated tests can be run regularly.
  5. Early checkins (in fact daily or more than once a day) and quick updates by the developers is important so that the integration problems are detected as early as possible.
  6. It is important to communicating with the customers and getting regular feedback.
  7. Best way to show the customer the progress of the project is to show them a working demo of the application.
  8. Introduce a process change when the team is not under pressure. Point out the benefit the stakeholders will have with the new process. Show them the benefit of the process/practice rather than talk and preach about it.

A wonderful Dilbert quote from the book
"I love deadlines. I especially love the swooshing sound they make as they go flying by." — Scott Adams

Some Excerpts from the book

View all my reviews