LV
Back to writing

// Desenvolvimento Web · Linux · Debian · Desenvolvimento · Linux Mint · Linux Mint Debian Edition · Ubuntu

Copy / Download an Entire Website With WGET [rip]

To copy an entire website, including photos, JS and CSS files, using wget on Linux, run the following command:

$ wget \
     --user-agent="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/50.0.2661.94 Safari/537.36" \
     --recursive \
     --no-clobber \
     --page-requisites \
     --html-extension \
     --convert-links \
     --restrict-file-names=windows \
     --domains exemplo.com \
     --no-parent \
     -e robots=off \
     --load-cookies=cookies.txt \
         exemplo.com/blog

The –domains parameter tells wget to only download files from that domain.
The last one says which root to copy from. In this case I would only copy the files, recursively, from /blog of the site exemplo.com.

The explanation of the parameters follows below:

–recursive: download the entire Web site.

–domains website.org: don’t follow links outside website.org.

–no-parent: don’t follow links outside the directory tutorials/html/.

–page-requisites: get all the elements that compose the page (images, CSS and so on).

–html-extension: save files with the .html extension.

–convert-links: convert links so that they work locally, off-line.

–restrict-file-names=windows: modify filenames so that they will work in Windows as well.

–no-clobber: don’t overwrite any existing files (used in case the download is interrupted and
resumed).

Comments 0

No comments yet.