MAJOR=13
MINOR=65
DEVNAME=input/event1
  #   ˆl–äY>”Je¶¥8 ¸Û7ú˜´oÁ„Á}¯À¿ ?÷     <!--
Copyright (C) Daniel Stenberg, <daniel@haxx.se>, et al.

SPDX-License-Identifier: curl
-->

# The Art Of Scripting HTTP Requests Using curl

## Background

This document assumes that you are familiar with HTML and general networking.

The increasing amount of applications moving to the web has made "HTTP
Scripting" more frequently requested and wanted. To be able to automatically
extract information from the web, to fake users, to post or upload data to
web servers are all important tasks today.

curl is a command line tool for doing all sorts of URL manipulations and
transfers, but this particular document focuses on how to use it when doing
HTTP requests for fun and profit. This documents assumes that you know how to
invoke `curl --help` or `curl --manual` to get basic information about it.

curl is not written to do everything for you. It makes the requests, it gets
the data, it sends data and it retrieves the information. You probably need
to glue everything together using some kind of script language or repeated
manual invokes.

## The HTTP Protocol

HTTP is the protocol used to fetch data from web servers. It is a simple
protocol that is built upon TCP/IP. The protocol also allows information to
get sent to the server from the client using a few different methods, as is
shown here.

HTTP is plain ASCII text lines being sent by the client to a server to
request a particular action, and then the server replies a few text lines
before the actual requested content is sent to the client.

The client, curl, sends an HTTP request. The request contains a method (like
GET, POST, HEAD etc), a number of request headers and sometimes a request
body. The HTTP server responds with a status line (indicating if things went
well), response headers and most often also a response body. The "body" part
is the plain data you requested, like the actual HTML or the image etc.

## See the Protocol

Using curl's option [`--verbose`](https://curl.se/docs/manpage.html#-v) (`-v`
as a short option) displays what kind of commands curl sends to the server,
as well as a few other informational texts.

`--verbose` is the single most useful option when it comes to debug or even
understand the curl<->server interaction.

Sometimes even `--verbose` is not enough. Then
[`--trace`](https://curl.se/docs/manpage.html#-trace) and
[`--trace-ascii`](https://curl.se/docs/manpage.html#--trace-ascii)
offer even more details as they show **everything** curl sends and
receives. Use it like this:

    curl --trace-ascii debugdump.txt https://www.example.com/

## See the Timing

Many times you may wonder what exactly is taking all the time, or you want to
know the amount of milliseconds between two points in a transfer. For those,
and other similar situations, the
[`--trace-time`](https://curl.se/docs/manpage.html#--trace-time) option is
what you need. It prepends the time to each trace output line:

    curl --trace-ascii d.txt --trace-time https://example.com/

## See which Transfer

When doing parallel transfers, it is relevant to see which transfer is doing
what. When response headers are received (and logged) you need to know which
transfer these are for.
[`--trace-ids`](https://curl.se/docs/manpage.html#--trace-ids) option is what
you need. It prepends the transfer and connection identifier to each trace
output line:

    curl --trace-ascii d.txt --trace-ids https://example.com/

## See the Response

By default curl sends the response to stdout. You need to redirect it
somewhere to avoid that, most often that is done with `-o` or `-O`.

# URL

## Spec

The Uniform Resource Locator format is how you specify the address of a
particular resource on the Internet. You know these, you have seen URLs like
https://curl.se/ or https://example.com/ a million times. RFC 3986 is the
canonical spec. The formal name is not URL, it is **URI**.

## Host

The hostname is usually resolved using DNS or your /etc/hosts file to an IP
address and that is what curl communicates with. A