Contents
TABLE OF CONTENTS
Anshuman Praharaj16 min read

Understanding HTTP from First Principles

I recently gave an interview, and the feedback I got from the interviewer at the end was that even though I have a good breadth of knowledge in backend engineering, I need to go deeper into certain concepts and contexts to strengthen my fundamentals.

The first thing I did was go back to a playlist called Backend from First Principles . It's not like I didn't know about this playlist. I had actually bookmarked it a few months ago, but I was learning different things at the time. I also went through his Go playlist about a month ago and built a few projects around it. You can read about those in another blog.

But now that I had this specific feedback, I knew this was something I needed to pursue immediately. I don't want to keep making the same mistake and miss opportunities simply because I wasn't focusing enough on the things that actually matter.

So I picked up the playlist.

The first two or three videos were more ad hoc and weren't really about a specific technical concept. They were about learning, how to learn, and why learning from first principles is useful. The first proper conceptual topic in the playlist is HTTP.

This blog is about that video. The video itself is around an hour long, and I'm going to try to compress what I learned into a few minutes, hopefully under 20 minutes.

We are going to go through a few concepts around HTTP that are particularly relevant for backend engineers. This isn't going to be a deep dive into networking layers, the data link layer, or the entire OSI model. There are a lot of things that happen underneath HTTP, but for this blog, we're going to stay at the application layer and focus specifically on HTTP.

1. The property of HTTP

HTTP and HTTPS are stateless protocols by nature.

What does stateless actually mean?

It means that the server does not inherently remember previous interactions with a client. Each HTTP request is treated independently. So if the server needs to identify a user or understand their context, the required information has to be provided with the request.

This is where things like cookies, tokens, session identifiers, and other user-related data come into the picture. They allow the server to understand who is making the request and what context that request belongs to.

This is also one of the fundamental properties of HTTP that becomes important when building backend systems.

2. HTTP Headers

An HTTP request is made up of different parts, and one of the important parts is HTTP headers.

HTTP headers are name-value pairs that allow the client and server to pass additional information and metadata along with an HTTP request or response.

Headers exist on both sides of the request-response cycle.

When a client sends a request to a server, the request contains headers. When the server sends a response back, that response can also contain headers.

Headers and the actual data being transferred are two different things. You can think of headers as metadata , additional information that provides more context about what is happening during the request-response cycle.

Headers can contain information about things like the type of request, the host, the client or user agent, connection behavior, caching policies, security policies, content type, and much more.

We'll look at some of these as we go through HTTP.

3. HTTP Methods

The next important concept is HTTP methods.

The methods we commonly work with as backend engineers are:

  • GET

  • POST

  • PUT

  • PATCH

  • DELETE

So what are these methods actually used for?

GET

GET is primarily used to fetch data from the server.

Conceptually, a GET request should be used for retrieving data and should not modify the state of the server. A GET request typically does not have a request body, although the HTTP specification does not simply define "GET can never have a body" as a universal prohibition.

Most of the time, when we're building APIs, GET is what we use when we want to retrieve a resource.

POST

POST is primarily used when we want to create or submit data to the server.

The data can be passed in different ways depending on how the server is designed. It can be present in the request body, and information can also be provided through query parameters or other parts of the request.

As a backend engineer, it's important to follow the intended semantics of HTTP methods instead of using POST for everything.

For example, POST shouldn't become the default method for updating data just because it works. HTTP already provides methods specifically intended for updating resources.

Those are PUT and PATCH.

PUT vs PATCH

PUT and PATCH are very similar because both can be used to update a resource, but they differ in how the update is represented.

PUT is generally used to replace a resource with a new representation.

For example, if you have a user resource containing:

{
  "name": "Anshuman",
  "email": "anshuman@example.com",
  "age": 23
}

A PUT request can represent the complete new version of that resource.

PATCH, on the other hand, is used for a partial modification of a resource.

If you only want to change the user's age, you don't necessarily need to send the entire resource again. You can send just the field that needs to be modified.

So the simple distinction is:

PUT → replace the resource

PATCH → partially modify the resource

DELETE

DELETE is pretty self-explanatory. It is used to request the removal of a resource.

For example, if you have an endpoint like:

DELETE /users/123

the intention is to remove the user resource identified by 123.

Idempotent vs Non-Idempotent Methods

HTTP methods can also be classified based on whether they are idempotent.

An operation is idempotent if performing it multiple times has the same intended effect on the server as performing it once.

For example, if you make the same PUT request multiple times, the resource should end up in the same state as it would after making that request once.

GET, PUT, and DELETE are considered idempotent methods.

This doesn't necessarily mean that the response will be byte-for-byte identical every single time. It means that repeating the same operation should not keep producing additional changes to the resource's state.

POST is generally non-idempotent because sending the same POST request multiple times can result in multiple resources being created.

PATCH requires a little more nuance. PATCH is not inherently non-idempotent — whether a particular PATCH operation is idempotent depends on what the operation actually does.

Options

There is also another HTTP method called OPTIONS.

This isn't something we usually use directly for normal application-level data operations. It is used to discover what communication options are available for a server or a particular resource.

For example, a client or browser can use OPTIONS to find out which HTTP methods are supported and what other request-related options are available.

You'll commonly encounter this when dealing with CORS and browser preflight requests.

In a browser-based application, the browser may send an OPTIONS request before the actual request to check whether the server allows the intended cross-origin communication.

So, in a simplified way, you can think of an OPTIONS request as the client asking the server:

"Before I send the actual request, what am I allowed to do here?"

Once the server responds with the appropriate information, the browser can decide whether it can proceed with the actual request.

HTTP, CORS, and Types of Requests

Now let's talk about CORS, or Cross-Origin Resource Sharing.

Before sending certain cross-origin requests, the browser sends what is called a pre-flight request. This is the OPTIONS request we discussed earlier. The purpose of this request is to find out what communication options, methods, headers, and origins the server allows.

The important thing to understand here is that CORS is enforced by the browser.

Browsers follow something called the Same-Origin Policy, which basically restricts a web page from freely accessing resources from a different origin.

An origin is determined by things like the scheme, domain, and port.

For example, your frontend application might be running on:

https://myapp.com

while your API might be running on:

https://api.myapp.com

These are different origins, even though they belong to the same overall domain.

Similarly, during local development, your frontend might run on:

http://localhost:3000

and your backend might run on:

http://localhost:8080

The host is the same, but the ports are different, so these are different origins as well.

This is why you often need to configure CORS on your backend.

If the browser sees that your frontend is trying to communicate with an API from a different origin, it checks whether that communication is allowed by the server. For requests that require a pre-flight, the browser first sends an OPTIONS request.

The server can respond with headers such as:

Access-Control-Allow-Origin

along with other CORS-related headers that specify which methods and headers are allowed.

If the response indicates that the origin and requested operation are allowed, the browser proceeds with the actual request.

If the server doesn't allow it, the browser blocks the frontend from accessing the response.

One important distinction here is that CORS is primarily a browser-side security mechanism. The server can still receive requests from other clients such as curl, Postman, or another backend service. CORS doesn't act as a general firewall for your API.

HTTP Response Codes

The next thing is HTTP response codes.

HTTP status codes are standard three-digit numbers sent by the server to communicate the result or state of a request to the client.

One useful thing about HTTP is that these status codes are standardized. It doesn't matter whether your web server is written in Go, JavaScript, Java, Ruby, or any other language. The HTTP semantics remain the same.

For example, a successful request can return 200, a resource that doesn't exist can return 404, and an internal server error can return 500.

The important thing is to use the status code that actually represents what happened.

For example, you shouldn't return 401 Unauthorized when the actual situation is 403 Forbidden.

Here are some commonly used HTTP status codes:

  • 200 OK: the request was successful.

  • 201 Created: a new resource was successfully created.

  • 204 No Content: the request was successful, but there is no response body.

  • 301 Moved Permanently: the resource has permanently moved to another location.

  • 304 Not Modified: the cached version of the resource can still be used.

  • 400 Bad Request: the request is invalid or malformed.

  • 401 Unauthorized: authentication is required or the provided authentication credentials are invalid.

  • 403 Forbidden: the server understood the request but refuses to authorize it.

  • 404 Not Found: the requested resource doesn't exist.

  • 405 Method Not Allowed: the HTTP method isn't supported for that resource.

  • 409 Conflict: the request conflicts with the current state of the resource.

  • 429 Too Many Requests: the client has sent too many requests in a given period.

  • 500 Internal Server Error: something went wrong on the server.

  • 502 Bad Gateway: a server acting as a gateway or proxy received an invalid response from an upstream server.

  • 503 Service Unavailable: the server is currently unable to handle the request, often because it is overloaded or temporarily unavailable.

There are many more status codes, but these are some of the ones you'll come across regularly while building backend systems.

HTTP Caching

The next concept is HTTP caching.

HTTP caching allows the client or an intermediary cache to reuse a response that was previously sent by the server instead of requesting the same resource again every time.

The basic idea is pretty simple. If a resource hasn't changed, there is no reason to download the same data again.

This helps reduce unnecessary requests, save bandwidth, and improve response times.

Two important headers to understand here are Cache-Control and ETag.

Cache-Control

Cache-Control tells the client or intermediary caches how a response should be cached and for how long.

It can contain directives such as:

  • max-age: how long the response can be considered fresh.

  • no-cache: the cached response can be stored, but it needs to be revalidated before being reused.

  • no-store: don't store the response in the cache.

  • public: the response can be cached by shared caches.

  • private: the response is intended for a private cache, such as the user's browser.

So instead of blindly downloading the same resource again and again, the client can use the caching rules provided by the server.

ETag

Another important caching mechanism is ETag, or Entity Tag.

An ETag is a value that identifies a particular version of a resource. It is commonly generated from the contents or state of that resource, although it doesn't necessarily have to literally be a hash.

For example, imagine the server sends a resource with:

ETag: "abc123"

The browser can cache that response.

Later, instead of downloading the entire resource again, the browser can send a conditional request containing:

If-None-Match: "abc123"

The server then checks whether the resource has changed.

If the resource is still the same, the server can respond with:

304 Not Modified

This basically tells the client, "the version you already have is still valid, you can use your cached copy."

The actual resource doesn't need to be sent again, which saves bandwidth.

If the resource has changed, its ETag will generally be different. The server can then return the new representation along with the new ETag, and the client can update its cached version.

So at a high level:

Cache-Control tells us how caching should work.

ETag helps us determine whether the cached representation is still valid.

This is one of those HTTP features that looks simple on the surface, but becomes pretty useful when you're building systems where reducing unnecessary network requests and bandwidth actually matters.

HTTP Content Negotiation

Now let's talk about HTTP Content Negotiation.

Basically, the client and server negotiate how the data should be shared and used.

The client can send its preferences for things like encoding, compression, language, and the format it can understand.

The server then tries to respond according to those preferences.

Of course, the server may not always be able to follow the exact format requested by the client.

In those cases, it can choose a compatible format that both sides can work with.

Language is one common example of this.

There are two important headers here:

Accept-Language is sent by the client.

Content-Language is sent by the server.

They work together as part of Content Negotiation.

For example, a browser can tell the server that the user prefers English. The server can then respond with the English version of the page if it supports it.

Compression is another important part of this.

It can significantly reduce the amount of data that needs to be transferred between the client and server.

For some types of documents, compression can reduce the size by a lot. In some cases, the reduction can be around 70%.

That means less bandwidth is required to transfer the same content.

Some commonly used compression algorithms are gzip and Brotli.

There are other compression algorithms as well, but these are two that you'll commonly come across while working with HTTP.

Persistent Connections and Keep-Alive

Now let's talk about persistent connections.

One of the problems with HTTP/1.0 was that connections were generally short-lived.

A connection would be created for a request, the response would be sent, and then the connection would be closed.

If the client needed to make another request, it had to create another connection.

This obviously isn't very efficient when a webpage needs to load a lot of resources.

So, around the mid-1990s, developers started adding an unofficial extension called Keep-Alive to HTTP/1.0 implementations.

The idea was pretty simple.

Instead of closing the connection immediately after sending the response, the connection could be kept open and reused for multiple request-response cycles.

This eventually became part of HTTP/1.1.

The benefit is that multiple requests and responses can use the same underlying TCP connection.

You don't need to establish a completely new TCP connection every time you want to send some data.

This reduces the overhead of creating new connections and makes communication more efficient.

Multipart and Streaming

Another interesting part of HTTP is how it handles larger or different types of data.

Multipart is used when a single HTTP request or response needs to carry multiple parts of data.

A common example is uploading a form with both text fields and a file.

The request can contain multiple parts, and each part can have its own content type and other metadata. The server can then read each part separately.

This is commonly done using multipart/form-data.

Streaming is a little different.

Instead of sending the entire response at once, the server can send the data in smaller chunks over time.

For example, imagine you're reading a PDF from a server. The server doesn't necessarily have to prepare the entire file and then send everything at once. It can start sending the data as it becomes available, and the client can start receiving and processing it while the rest is still being sent.

The important thing to understand is that the underlying TCP connection is still carrying the data. HTTP is the application layer that defines how that data is represented and transferred.

So you can have one HTTP request and response, while the actual data is split across many smaller pieces underneath.

This is useful for things like large file downloads, video playback, server-sent data, and other cases where waiting for the entire response before starting to process it would be inefficient.

TLS, SSL and HTTPS

Now let's talk about TLS, SSL, and HTTPS.

This isn't something you need to go extremely deep into just to understand HTTP.

The important thing to know is that TLS is the successor to SSL.

SSL was originally used to provide encrypted communication between the client and server.

But older versions of SSL had security issues and were eventually deprecated.

TLS was introduced as its successor and is the protocol used today for securing HTTPS connections.

So when you see:

https://example.com

the S basically means that HTTP is being used over a secure TLS connection.

HTTPS provides encrypted communication between the client and server.

So at this point, the simple thing to remember is:

HTTP is the protocol used for communication.

HTTPS is HTTP secured using TLS.

And TLS is the modern successor to SSL.

thats all about this blog . keep visiting though i will be posting a lot of blogs on backend concepts.