Why is it so slow?

That’s THE question! But if it “crawls,” it still means it “sort of works”… And if it works but the user complains (always rightly so, obviously! Unless you’re dealing with a PEBCAK…), then the answer is going to be a bit more complicated than if it didn’t work at all! […]

RRobin/December 08, 2010/9 min read

That’s THE question!

But if it “crawls,” it still means it “sort of works”…

And if it works but the user complains (always rightly so, obviously! Unless you’re dealing with a PEBCAK…), then the answer is going to be a bit more complicated than if it didn’t work at all! Naturally.

We’ll still try to give a general answer to the question “Why?”.

And just before saying why, it’s worth saying What? and How much? is “crawling.”

How much is it crawling?

Saying that “it crawls” means that the application is slow. More precisely, some or all of an application’s unit operations (we call that an application transaction) are slower than normal. Ah! So there can be a norm! Interesting.

And so we’re talking about a metric related to time… So we have a unit of measure, the second, and the metric that carries it has a name, it’s theApplication Response Time as seen by the user. And we can even consider that there are several types of Response Time:

  • an overall response time for the application, although that doesn’t necessarily make sense since transactions can vary so much in form and content,

  • a Response Time for a particular application transaction, that’s already more precise, but there can be a great diversity of them,

  • finally, a Response Time for an application scenario, that is, a typical and representative sequence of chosen application transactions. Here we get closer to the real experience and therefore to the user’s perception.

Unfortunately, it’s rare for IT organizations to define a typical application transaction scenario and to have a “baseline,” that is, that famous norm of what the application response time as seen by the user should be for that transaction scenario.

And yet a user knows when “it crawls”! In the end, this baseline exists “in their head” and depends on their average experience. So the user must be believed! Even if the feedback is sometimes subjective, it’s valid.

And of course, there are tools to get technical and therefore objective visibility into this application Response Time as seen by the user. There are even ways to measure baselines manually or automatically, for later comparison. But that would deserve a dedicated article on Application Response Time and how to measure it.

So for now let’s focus on the why.

CNS breakdown: Client-Network-Server ventilation

To start with, we should distinguish several origins “horizontally.” There is indeed a path through the use of the application, each part of which can introduce delays and therefore increase the famous Application Response Time as seen by the user.

Roughly speaking, we ultimately have in play, along the application’s “path”:

  • the client workstation (C), the user depends on it, and so does their perception. Who has never experienced that, just to do a fully local operation on their machine, already, “it crawls”! So the client workstation can be a cause, not always the simplest one, of an application Response Time problem.

  • the whole network (N), from the site’s LAN, through routers, a wide area network, known as WAN with its local loops, its backbone and its carriers, then the security infrastructure allowing or not allowing traffic to the servers, and finally the LAN with its VLANs and all the Datacenter paraphernalia where the application servers are hosted. That makes for a lot of possible causes of slowdowns!

  • “the” server (S), which very often and increasingly is not just 1 server but several, one behind the other (multi-tier application) and/or beside each other (cluster, to share the load and cover the risk of failure), and what’s more, it is also now often “virtual,” that is, purely software alongside other virtual servers sharing the resources of a big physical machine whose only job is to hand out its resources. In short, quite a complexity of causes in store here too.

We clearly understand that this “CNS” ventilation of application response time is an important basis for dichotomizing the cause of the performance problem. And there is still room to distinguish, within each of these 3 components, different sub-parts involved.

It’s “throughput,” of course!

Aaaah throughput! Bandwidth! That one at least, everyone knows: if it’s slow, it’s because the network isn’t fast enough, not enough throughput. Just increase the bandwidth and that’s it!

Yes, but no.

First, the network has more characteristics than just its throughput: its latency, its packet loss rate, its jitter, etc.

Fine, but let’s focus nonetheless on bandwidth alone. The application may or may not be bandwidth-hungry (file transfers for example), it may use small (Citrix) or large packets (HTTP, FTP), it may be fragile in the face of competition (real-time, transactional, voice over IP) or robust. And the user may see the effect of the lack of bandwidth (on voice, it’s audible! on Citrix, it stalls) or not at all (an email can arrive 30 seconds later, what difference does it make!).

And bandwidth can be lacking with different subtleties and different effects:

  • Saturation: bandwidth is completely and continuously lacking, potentially under the effect of a single voracious application! And classically, you can no longer work, everything is slow. This case is often the easiest to diagnose and fix once the source is identified.

  • Congestion: the nuance of the term is more a matter of convention, but the case is concretely different. The lack of resource is episodic, it appears under the effect of competition between applications with different protocols and behaviors. The effect is also variable depending on the application: congestion is fatal for voice over IP, rendering it unusable, painful under Citrix, but potentially barely noticeable on basic web and totally invisible on email.Congestion can be difficult to detect because it can disappear under the effect of averaging in graphs! It can also be tricky to fix because it is the result of several throughput-consuming sources, so you have to subtly control the applications by prioritizing flows for example, which requires analysis and special means.

  • Inflation: bandwidth can also be lacking “whatever happens,” even if you increase it, the thirst of applications (some of them, shall we say) is not quenched! It’s a very perverse effect of TCP at work here in applications transferring large volumes. Naturally, TCP seeks to occupy all the available space as long as there is no congestion, and so… it is quite normal to constantly reach congestion levels when transferring a lot of data! The other inflationary effect is not technical but psychological and comes from users: give them more resources, better comfort, and hop, usage increases! In short, there can almost never be enough bandwidth. What can you do? Analyze and control, again.

Thus, the single question of bandwidth, the one that seems most obvious, can still turn out to be subtle. And nobody wants to pay for a bigger network with a yearly or longer commitment only to realize that it doesn’t improve the situation… Even though analysis and good advice could have avoided it with a prioritization solution or a simple usage policy on a firewall!

Visibility and Control, once again, for this “simple” question of bandwidth are necessities.

And protocols, they’re everywhere, protocols!

And yes, protocols are everywhere! On every floor… of the OSI layers, of course.

And there the debate reaches levels of complexity. The network is generally the layers up to IP, that is, OSI 1, 2 and 3, namely Physical, Data Link and Network. The application, for its part, with its APIs does Sessions, Presentation and finally Application, i.e. OSI 5, 6 and 7. And that’s already quite a few protocols for each, when you have to analyze a problem on the network side and on the application side.

But… what about layer 4? Transport… TCP!

TCP, that unknown known to everyone! From the application, which doesn’t see it, since it talks to APIs or at best opens Sockets, period. The transport, well, the OS handles it. And from the Network side, which is transparent for everything IP, it’s promised! So TCP, well, it manages on its own, doesn’t it?

And that’s just the thing! TCP manages on its own with a whole barrage of internal mechanisms and algorithms: Three Way Handshake, Slow Start, Congestion Avoidance, Selective ACK, Triple Duplicate ACK, Delayed ACK, Nagle’s Algorithm, Maximum Transmit and Receive Window, Frozen Window…

So protocol effects that can harm performance, I won’t draw you a picture… there can be quite a number of them! There are those from the network, there are those from the applications, and the worst part is that there are also those that are neither one nor the other!

Hell, these protocols, hell!

What can you do? Know them, that’s all. To diagnose, it’s better. Ask your doctor how they go about treating you! They learned all the body’s mechanisms, and their diseases…

Application chattiness and network latency, who to blame?

If at last you have been able to identify that the application, through its infernal protocol layers, was doing ch-at-te-ry through the network (who said CIFS?), and that the said network had an in-sup-por-ta-ble latency (who said 3G?), and that this is why “It crawls!,” then bravo, you have a fine diagnosis!

Example, taken entirely at random… You browse a directory on drive Z: over the network through your 3G key: 200 round-trips of application transactions x 200 milliseconds = 40 seconds… that can be a long wait just to see the list of folders. No?

But there you have it, you also have a System or Application manager telling you “It’s not my app, it’s the network that has lousy latency, I can’t do anything about it!”; you also have a Network manager telling you “The application has to work over the network? Well then it just shouldn’t be so chatty, it’s badly written, that’s all! I’m not going to move France closer to Asia, nor change the speed of light!”

So, even when we know why it crawls… who is right? What do we do?

Obviously, I’m tempted to take a stand in favor of the world’s geography and the speed of light, which give a few solid foundations to the normal operation of a network… and of the application that must run on it. But when the application uses protocols or APIs that aren’t easy to overhaul… you have to be a bit more subtle.

Solutions exist, in pretty much every case. Most of the time, I even manage to offer my clients a choice! But one thing is certain: you have to analyze very, very precisely Who, What, Where and Why it crawls!

So, Why does it crawl? Application or Network?

In the end, the real cause, the answer to “Why?,” is rarely just “the network”!

Network technicians and engineers, the heads of “Telecom and Networks” departments will understand me: it feels good to talk about it. And to talk about it a bit more, that is the whole purpose of this Blog: to address these aspects of performance that are poorly mastered or little known, which are not just caused by “the network.”

It’s also a key topic for CIOs, who very often find themselves facing users (through the hierarchical channel of a site or business department manager!), with their own Network, Systems or Applications departments passing the buck and the incident ticket. For the more organized ones, in ITIL mode of course, it’s the Problem Manager and their expert committees who get bogged down, very often over a simple problem of methodology and cross-functional expertise in Application Performance Analysis (APA).

And finally, keep in mind that, Murphy’s law obliges, if a problem comes up to you, there’s a not negligible chance that it’s the accumulation of several sources, for “maximum aggravation”…

So, Why does it crawl? Not so obvious to say.

But otherwise, it wouldn’t be fun. And I wouldn’t be here…

R

Robin

Author