To copy an entire website, including photos, JS and CSS files, using wget on Linux, run the following command:
$ wget \
--user-agent="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/50.0.2661.94 Safari/537.36" \
--recursive \
--no-clobber \
--page-requisites \
--html-extension \
--convert-links \
--restrict-file-names=windows \
--domains exemplo.com \
--no-parent \
-e robots=off \
--load-cookies=cookies.txt \
exemplo.com/blog
The –domains parameter tells wget to only download files from that domain.
The last one says which root to copy from. In this case I would only copy the files, recursively, from /blog of the site exemplo.com.
The explanation of the parameters follows below:
–recursive: download the entire Web site.
–domains website.org: don’t follow links outside website.org.
–no-parent: don’t follow links outside the directory tutorials/html/.
–page-requisites: get all the elements that compose the page (images, CSS and so on).
–html-extension: save files with the .html extension.
–convert-links: convert links so that they work locally, off-line.
–restrict-file-names=windows: modify filenames so that they will work in Windows as well.
–no-clobber: don’t overwrite any existing files (used in case the download is interrupted and
resumed).

Comments 0
No comments yet.