gallery-dl

Commit Graph

Author	SHA1	Message	Date
Mike Fährmann	421a9740a3	[tumblr] add 'tumblr:' to force Tumblr extractor (#71 )	7 years ago
Mike Fährmann	40d35c87bc	[paheal] add tag- and post-extractors (closes #69 )	7 years ago
Mike Fährmann	cc0c2cca57	[reddit] add extractor for reddit-hosted images (closes #68 )	7 years ago
Mike Fährmann	f10ffc0839	update extractor blacklist to also allow classes	7 years ago
Mike Fährmann	35e09869d1	[mangapark] fix image URLs and use HTTPS	7 years ago
Mike Fährmann	9a049bdf51	[tumblr] add 'likes' extractor (#65 )	7 years ago
Mike Fährmann	67d4462d26	[batoto] rudimentary Cloudflare bypass	7 years ago
Mike Fährmann	29d75fc3fa	[tumblr] add support for OAuth authentication (#65 )	7 years ago
Mike Fährmann	4edb25346e	[slideshare] support mobile URLs (closes #67 )	7 years ago
Mike Fährmann	e420a28bbc	fix cookie tests	7 years ago
Mike Fährmann	b33efc99a4	[idolcomplex] add support for idol.sankakucomplex.com	7 years ago
Mike Fährmann	75b2e84b6d	[tumblr] use s3.amazonaws.com for image URLs (#64 )	7 years ago
Mike Fährmann	5b094328b5	[puremashiro] add chapter- and manga-extractor (closes #66 ) Also adds support for region subtags in language codes (e.g. en-us)	7 years ago
Mike Fährmann	974e73bdbb	[booru] smaller code adjustments	7 years ago
Mike Fährmann	03b8a548cb	[tumblr] change `reblogs` default value to `true` (#61 )	7 years ago
Mike Fährmann	d235f68f59	[tumblr] add option to filter reblogged posts (#61 ) Reblogs are ignored by default, but can be included by setting 'extractor.tumblr.reblogs' to 'true'.	7 years ago
Mike Fährmann	a794fffc6d	[batoto] extend chapter-string regex (closes #60 ) Non-numeric chapter indices exist after all ...	7 years ago
Mike Fährmann	1219ebb7f5	[danbooru] use alternate subdomains; support safebooru	7 years ago
Mike Fährmann	9e8a84ab6c	[booru] rewrite using Mixin classes (#59 ) - improved code structure - improved URL patterns - better pagination to work around page limits on - Danbooru - e621 - 3dbooru	7 years ago
Mike Fährmann	0876541e43	[seiga] update tests	7 years ago
Mike Fährmann	88bb0798fd	delay initialization of PathFormat objects This allows the DeviantArt group-check to be moved inside the Extractor.items() method which in turn allows for better exception handling. As a new general rule: Never raise exceptions during extractor initialization.	7 years ago
Mike Fährmann	c24e0e70a7	[pixiv] simplify main loop	7 years ago
Mike Fährmann	c1e331edbb	[mangapark] replace manga test	7 years ago
Mike Fährmann	28cd78aae0	[kissmanga] extend chapter-string regex (closes #58 )	7 years ago
Mike Fährmann	a3e9b51bea	[imgbox] update test results Image URLs of older galleries have been updated to the new format. https://i.imgbox.com/qHhw7lpG.png --> https://images3.imgbox.com/6d/9a/qHhw7lpG_o.png	7 years ago
Mike Fährmann	d0886f411e	[gelbooru] re-enable API use (closes #56 ) Gelbooru's API allows access to all images and is not restricted to the first 20000. This also adds an option to select between API use and manual information extraction in case their API gets disabled again.	7 years ago
Mike Fährmann	8102aae311	[mangahere] support ".cc" TLD and mobile URLs	7 years ago
Mike Fährmann	676602056c	[reddit] unescape output URLs	7 years ago
Mike Fährmann	2eedbaaaf9	[deviantart] use cache to store new refresh_tokens The 'refresh_token' set in a user's config file gets used once to get a new 'access_token' and 'refresh_token', which is then stored in gallery-dl's cache and gets used the next time the 'access_token' needs to be refreshed. This means deleting the cache file invalidates the refresh_token- chain and requires the user to re-authenticate.	7 years ago
Mike Fährmann	fc7d165c97	[deviantart] add support for OAuth2 authentication Some user galleries [] require you to be either logged in or authenticated via OAuth2 to access their deviations. [] e.g. https://polinaegorussia.deviantart.com/gallery/ -------------- known issue: A deviantart 'refresh_token' can only be used once and gets updated whenever it is used to request a new 'access_token', so storing its initial value in a config file and reusing it again and again is not possible.	7 years ago
Mike Fährmann	91c2aed077	[nhentai] fix JSON extraction	7 years ago
Mike Fährmann	444008a14a	[khinsider] use urljoin() to complete page URLs	7 years ago
Mike Fährmann	263741d243	[luscious] update URL pattern (closes #55 )	7 years ago
Mike Fährmann	0a9a07a6e1	[slideshare] improve metadata; flake8 - added 'views' and 'published' keywords - fixed longer titles and descriptions	7 years ago
Leonardo Taccari	a8d2dde8b2	[slideshare] Add a new extractor for slideshare.net (#54 )	7 years ago
Mike Fährmann	19a6ae57b2	[sankaku] add pool extractor	7 years ago
Mike Fährmann	e52f0cc1ed	[sankaku] add post extractor	7 years ago
Mike Fährmann	595593a35e	[sankaku] rewrite - better code structure and extensibility - better metadata	7 years ago
Mike Fährmann	a3924d2072	[sankaku] fix swf extraction (closes #52 )	7 years ago
Mike Fährmann	291369eab2	various smaller changes/additions	7 years ago
Mike Fährmann	300346ecdf	[mangazuki] remove extractors This site has been in "rebuild"-mode for a fairly long time and the current extractor code isn't going to work for the new version either.	7 years ago
Mike Fährmann	d275b1d9a3	[khinsider] fix extraction ... again	7 years ago
Mike Fährmann	6b8e3003df	[hentai2read] ensure consistent extraction results	7 years ago
Mike Fährmann	a1980b16f3	[gelbooru] various improvements - better metadata for pools - map ratings to s/q/e like other boorus do - skip() support	7 years ago
Mike Fährmann	93482a1f88	implement 'util.advance()'	7 years ago
Mike Fährmann	038e3b3369	[kissmanga] handle "AreYouHuman" redirects (#51 )	7 years ago
Mike Fährmann	2b9a783fc7	[khinsider] fix extraction	7 years ago
Mike Fährmann	214972bc9a	[gelbooru] use manual extraction ... to compensate for their disabled API. (https://gelbooru.com/index.php?page=forum&s=view&id=3875) This also adds an extractor for image-pools.	7 years ago
Mike Fährmann	55c64cad4b	[khinsider] fix filename extension and test-pattern	7 years ago
Mike Fährmann	b14de6ffc2	[tumblr] small improvements - don't transform inline GIF URLs - set 'type' parameter for API calls if there is only one post type selected	7 years ago
Mike Fährmann	9296a26eae	[tumblr] add warning messages	7 years ago
Mike Fährmann	65c1c53eb8	[khinsider] fix extraction	7 years ago
Mike Fährmann	12de658937	[tumblr] add options to control extraction behavior (#48 ) - posts : list of post-types to inspect - inline : scan post bodies for inline images - external: follow external links	7 years ago
Mike Fährmann	077f8c12be	[tumblr] original video URLs + continuous offset	7 years ago
Mike Fährmann	8eb12ebeae	[tumblr] support more post/media types (#48 ) This adds support for audio and video posts (most videos are shared from youtube/instagram which isn't supported -> youtube-dl), as well as link posts and image-search inside of text posts. Most of this is just WIP and will need some sort of improvement and options to enable/disable different media types etc.	7 years ago
Mike Fährmann	b8cdd42cab	[senmanga] fix extraction (again) this is basically a re-revert of `2ace5c7`	7 years ago
Mike Fährmann	e6814aebe2	add 'extractor.*.user-agent' config option	7 years ago
Mike Fährmann	6913eeaa40	[powermanga] replace manga extractor unit test My Hero Academia is gone	7 years ago
Mike Fährmann	7e0d9257a7	[hbrowse] fix manga extraction	7 years ago
Mike Fährmann	3c576d10c0	[seiga] better metadata + 'skip()' support	7 years ago
Mike Fährmann	f72318e593	[seiga] support more than 200 images Due to API restrictions and/or missing knowledge about and documentation of API usage, it was only possible to retrieve the latest 200 images of a niconico seiga user with said API. The new approach manually visits each HTML page and gets its information from there.	7 years ago
Mike Fährmann	baf8094868	improve Extractor.request()'s retry behavior	7 years ago
Mike Fährmann	7e7b64162b	[batoto] handle error 10031	7 years ago
Mike Fährmann	92027f67f9	use consistent names for URL constants root := <scheme>://<host> base_url := <root>/<common path>	7 years ago
Mike Fährmann	69cbc0619f	[mangastream] fix 'next-page' URLs (fixes #49 )	7 years ago
Mike Fährmann	980fd3616d	[tumblr] use API v2 (#48 )	7 years ago
Mike Fährmann	d6bed9f36f	[tumblr] prevent premature exit to get all images (fixes #48 )	7 years ago
Mike Fährmann	305da540c3	[mangahere] fix metadata extraction	7 years ago
Mike Fährmann	2d0cfb33e1	[xvideos] add user profile extractor (#45 )	7 years ago
Mike Fährmann	a393e6e538	[xvideos] add gallery extractor (#45 )	7 years ago
Mike Fährmann	3a8a0c1f35	[imgbox] rewrite / fix extraction (closes #47 )	7 years ago
Mike Fährmann	035ef655f1	[imagefap] update unit tests old gallery/image has been deleted	7 years ago
Mike Fährmann	239d7afea7	[hosturimage] fix extraction of larger images	7 years ago
Mike Fährmann	158e60ee89	[3dbooru] enable download continuation behoimi.org doesn't respect 'Range' headers and doesn't report 'Content-Length' for compressed content encodings.	7 years ago
Mike Fährmann	c4fcdf2691	Revert "[senmanga] fix extraction and download" This reverts commit `2ace5c7b3c`.	7 years ago
Mike Fährmann	81a7788b40	replace space characters in unit test URLs	7 years ago
Mike Fährmann	bf82181359	[jaiminisbox] fix extraction	7 years ago
Mike Fährmann	16783e327f	[common] fix UnboundLocalError in Extractor.request()	7 years ago
Mike Fährmann	2ace5c7b3c	[senmanga] fix extraction and download	7 years ago
Mike Fährmann	4d8387f93b	[pixiv] support mobile URLs (https://touch.pixiv.net/)	7 years ago
Mike Fährmann	ab2bf0b0dd	[deviantart] replace collection unittest	7 years ago
Mike Fährmann	289d6b65d2	[danbooru] extend and improve URL regex - add support for danbooru mirrors: - hijiribe.donmai.us - sonohara.donmai.us - todo: actually use these domains instead of redirecting everything to danbooru itself - improve handling of query string parameters	7 years ago
Mike Fährmann	5fa42336a2	[sankaku] add warning for unauthenticated users also improve URL pattern and add missing options to default config file	7 years ago
Mike Fährmann	6af921a952	[sankaku] rewrite/improve (fixes #44 ) - add wait-time between HTTP requests similar to exhentai - add 'wait-min' and 'wait-max' options - increase retry-count for HTTP requests to 10 - implement user authentication (non-authenticated users can only view images up to page 25) - implement 'skip()' functionality (only works up to page 50) - implement image-retrieval for pages >= 51 - fix issue with multiple tags	7 years ago
Mike Fährmann	9aecc67841	[common] explicitly handle HTTP status code 429	7 years ago
Mike Fährmann	d68a24aa70	[kissmanga] fix extraction site changed '\n' to '\r\n' for newlines	7 years ago
Mike Fährmann	864a63ed33	fix typo [skip ci]	7 years ago
Mike Fährmann	f3fbaa5c3e	[reddit] allow users to override the API User-Agent Only overriding the Client-ID is not enough if you want to follow Reddit's API access rules [1]. [1] https://github.com/reddit/reddit/wiki/API#rules	7 years ago
Mike Fährmann	31ea6001e8	[dynastyscans] improve metadata and filename formats	7 years ago
Mike Fährmann	2ef3c35c98	smaller textual changes - swapped doc for deviantart.mature and .original - updated gallery-dl.conf - "transferred" -> "delegated"	7 years ago
Mike Fährmann	68a0a7579c	fix/improve some regular expressions	7 years ago
Mike Fährmann	393755ee94	[tumblr] update tests	7 years ago
Mike Fährmann	75d3a1f72f	[deviantart] always download original images Deviation-objects returned by the DeviantArt API don't always contain the URL and metadata of the original image ([1]). Getting this information requires an additional API call [2], which is indicated by the 'is_downloadable' and 'download_filesize' metadata within a deviation-object. [1] https://myria-moon.deviantart.com/art/Aime-Moi-part-en-vadrouille-261986576 [2] https://www.deviantart.com/developers/http/v1/20160316/deviation_download/bed6982b88949bdb08b52cd6763fcafd	7 years ago
Mike Fährmann	a1c8b21cfd	[senmanga] improve metadata	7 years ago
Mike Fährmann	994b2fc1e7	[deviantart] replace 'author[urlname]' keyword author[urlname] has always only been the lowercase version of author[username], which can now be directly converted to lowercase using the 'l' conversion: '{author[username]!l}'	7 years ago
Mike Fährmann	633b376f35	improve/adjust default filename formats for manga sites	7 years ago
Mike Fährmann	41adb99e9c	[pawoo] fix extraction - changed access_token - use account-search instead of general search	7 years ago
Mike Fährmann	b319f4bab3	smaller code and text changes	7 years ago
Mike Fährmann	ad4580800c	[pixiv] add support for more URL patterns - https://www.pixiv.net/mypage.php#id=USERID - https://www.pixiv.net/#id=USERID	7 years ago
Mike Fährmann	82ea6c0cd3	adjust format strings with optional titles ... except for anything manga/comic related	7 years ago
Mike Fährmann	85a2b2ae59	[khinsider] fix extraction	7 years ago
Mike Fährmann	26a866e7d8	implement (sub)category-transfer between extractors (#41 ) ImageFap- and all Manga-Extractors will transfer their (sub)category values to other extractors instantiated by them, which will in turn allow those to use options set for their parents. Example: ImagefapGalleryExtractors will use options set under extractor.imagefap.user, if (and only if) they have been instantiated by a ImagefapUserExtractor; and options from extractor.imagefap.gallery otherwise.	7 years ago
Mike Fährmann	1ab4c7986f	[mangahere] fix extraction would switch to HTTPS, but there seem to be certificate issues	7 years ago
Mike Fährmann	8e14714c2b	[imgspice] fix extraction	7 years ago
Mike Fährmann	9c138dfc1f	[common] detect empty HTTP response bodies	7 years ago
Mike Fährmann	c51616f8d8	[foolslide] fix minor chapter number	7 years ago
H R X N	77bf923c56	Update imgur.py to include 'title' of single image (#40 ) Add {title} keyword.. Images on Imgur don't necessarily have a title, but I think most of them do, and since this should not break anything else..	7 years ago
Mike Fährmann	a85f06d2d1	[foolslide] restructure; convert suitable values to int	7 years ago
Mike Fährmann	deb2e803ba	simplify MangaExtractor class	7 years ago
Mike Fährmann	9fc1d0c901	implement and use 'util.safe_int()' same as Python's 'int()', except it doesn't raise any exceptions and accepts a default value	7 years ago
Mike Fährmann	8963da8fd8	[spectrumnexus] extract manga metadata	7 years ago
Mike Fährmann	a3e40734d1	[mangareader] extract manga metadata	7 years ago
Mike Fährmann	9196005a4d	[mangazuki] extract manga metadata	7 years ago
Mike Fährmann	543ba245eb	[deviantart] update test results thumbnail URLs changed from //tXX.… to //t00.…	7 years ago
Mike Fährmann	b7a54a51d0	[mangapark] extract manga metadata + code improvements	7 years ago
Mike Fährmann	d39b8779af	[mangahere] extract manga metadata	7 years ago
Mike Fährmann	c265cc074a	[hbrowse] fix syntax for Python3.3 and 3.4	7 years ago
Mike Fährmann	a9e7145651	[hbrowse] extract hmanga metadata & general maintenance	7 years ago
Mike Fährmann	92c8a6cb01	[hentai2read] extract hmanga metadata	7 years ago
Mike Fährmann	de174b40d6	[hentaihere] extract hmanga metadata	7 years ago
Mike Fährmann	04cc1ffe34	[kissmanga] extract manga metadata	7 years ago
Mike Fährmann	885bd4cbe2	[readcomiconline] extract comic metadata	7 years ago
Mike Fährmann	cebf800a7f	[foolfuuka] add support for more sites (#18 ) - https://arch.b4k.co - https://archive.whatisthisimnotgoodwithcomputers.com - https://archive.yeet.net Notes: - The name "whatisthisimnotgoodwithcomputers" is way too long ... - archive.yeet.net is out of date and also blocked by 4chan servers - newest threads are 2 weeks old - using "https://archive.yeet.net" as Referer header results in "403 Forbidden" when accessing 4chan	7 years ago
Mike Fährmann	84d4450410	[fallenangels] extract manga metadata	7 years ago
Mike Fährmann	f32b1a0292	[imgyt] fix extraction	7 years ago
Mike Fährmann	4ad903b797	[warosu] fix extraction	7 years ago
Mike Fährmann	b84f48dfa5	[batoto] extract manga metadata	7 years ago
Mike Fährmann	4ceb176c6b	[foolslide] extract manga metadata enables chapter filtering for - https://kobato.hologfx.com/ - https://jaiminisbox.com/ - https://reader.kireicake.com/ - https://powermanga.org/ - https://reader.seaotterscans.com/ - http://sensescans.com/ - http://www.slide.world-three.org/	7 years ago
Mike Fährmann	24e5f154a4	[deviantart] update test results API responses now contain proper https:// URLs and their image download server is now "orig00.deviantart.net" for all images.	7 years ago
Mike Fährmann	0dedbe759c	enable '--chapter-filter' The same filter infrastructure that can be applied to image URLS now also works for manga chapters and other delegated URLs. TODO: actually provide any metadata (currently supported is only deviantart and imagefap).	7 years ago
Mike Fährmann	31cd5b1c1d	[luscious] detect high-load responses	7 years ago
Mike Fährmann	470bbe9d8c	fix smaller stuff - change filename option in example config file - adapt default filename format for mangafox - remove unnecessary newline [skip ci]	7 years ago
Mike Fährmann	6f30cf4c64	change keyword names to valid Python identifiers This commit mostly replaces all minus-signs ('-') in keyword names with underscores ('_') to allow them to be used in filter-expressions. For example 'gallery-id' got renamed to 'gallery_id'. (It is theoretically possible to access any variable, regardless of its name, with 'locals()["NAME"]', but that seems a bit too convoluted if just 'NAME' could be enough)	7 years ago
Mike Fährmann	54c0715135	allow users to set their own API access_tokens/client_ids	7 years ago
Mike Fährmann	49c7e70c10	[acidimg] add image extractor	7 years ago
Mike Fährmann	9b21d3f13c	add '--filter' command-line option This allows for image filtering via Python expressions by the same metadata that is also used to build filenames (--list-keywords). The usually shunned eval() function is used to evaluate filter-expressions, but it seemed quite appropriate in this case and shouldn't introduce any new security issues, as any attacker that could do > gallery-dl --filter "delete-everything()" ... could as well do > python -c "delete-everything()"	7 years ago
Mike Fährmann	00420ff202	[booru] consistent order for "popular" results	7 years ago
Mike Fährmann	83cf1e1d6d	[sankaku] unescape image URLs	7 years ago
Mike Fährmann	f98e3e8002	[luscious] fix tag extraction	7 years ago
Mike Fährmann	65997d835b	replace popular/ranking tests with older ones Metadata of several year old lists shouldn't change as much as it would for newer ones, which makes metadata-comparisons of the output of build_testresult_db.oy easier.	7 years ago
Mike Fährmann	be30fb2f98	add common config category for boorus and foolslide	7 years ago
Mike Fährmann	c0755a4d5e	[exhentai] revert login-method to its old version (#37 ) Additional cookies don't seem to help and have to be manually set anyway. The older method is more likely to succeed, so I'd rather use this one.	7 years ago
Mike Fährmann	3ee39ffd93	[exhentai] update login procedure (#37 ) This new version behaves pretty much exactly like a browser would and caches all cookies sent to it and not just "ipb_member_id" and "ipb_pass_hash".	7 years ago
Mike Fährmann	88a386977e	[booru] add "popular" extractors for more sites - konachan.com - behoimi.org - e621.net	7 years ago
Mike Fährmann	07214f4007	[booru] place subcategories into base classes	7 years ago
Mike Fährmann	60a888a1e4	[foolfuuka] add common config category All FoolFuuka based 4chan-archive extractors can now be configured using their own config keys (extractor.<category>) as well as a common shared one (extractor.foolfuuka).	7 years ago
Mike Fährmann	47bcf53ec1	implement support for additional unit test result types - "pattern" matches all resulting URLs against the given regex - "count" allows to specify the amount of returned URLs	7 years ago
Mike Fährmann	2d0dfe9d56	[exhenai] init headers before login and detect sadpanda - also debug-logs html after failed login - #37	7 years ago
Mike Fährmann	c7ec103e15	[batoto] fix extraction of chapter URLs	7 years ago
Mike Fährmann	18e6ed1c7e	[booru] add extractors for "Popular" images	7 years ago
Mike Fährmann	f7cdfd4c25	add a simplified version of 'parse_qs' This version only returns a dict of plain string to string key-value pairs and ignores multiple values for the same query variable.	7 years ago
Mike Fährmann	3b21e0703c	[deviantart] allow distinction between users and groups (#26 ) This is done by prepending "group-" to an extractor's subcategory if the URL belongs to a group ("folder" becomes "group-folder" and so on). This changes the configuration-path being used and is also reflected in the output of '--list-keywords'.	7 years ago
Mike Fährmann	e61a3a56d1	[hentai2read] fix and update keywords Added the "author" keyword and changed the name of a few others to be consistent with other manga/chapter extractors.	7 years ago
Mike Fährmann	c45770331a	use 'str.partition()' The (r)partition method is always faster then split() or any other method that has been replaced in this commit.	7 years ago
Mike Fährmann	017a72f448	[pixiv] improve input validation	7 years ago
Mike Fährmann	dcf42c5e89	[pixiv] add extractor for ranking lists	7 years ago
Mike Fährmann	4ea82ea556	[warosu] add thread extractor	7 years ago
Mike Fährmann	9aa95fba8c	[deviantart] adapt download URLs to use https Even though DeviantArt is "completely switching over to HTTPS"[1], every URL contained in an API response is still using HTTP [1] https://danlev.deviantart.com/journal/DeviantArt-Is-Switching-To-HTTPS-697996906	7 years ago
Mike Fährmann	02e89700fc	[foolfuuka] ensure sorted posts	7 years ago
Mike Fährmann	8bcf88bff7	[flickr] fix extraction This issue was only noticeable with older Python versions, as these don't exhibit a consistent ordering of dict keys.	7 years ago
Mike Fährmann	004456d5d5	properly update the config-dictionary When using 2 or more config files, the values of the second would improperly overwrite nested dictionaries of the first one. The new method properly combines these nested dictionaries as well.	7 years ago
Mike Fährmann	cfa479fab5	update error message for unspecified exceptions - ask user to report unexpected errors, which usually indicate extractor failure - handle OSErrors separately (permissions, disk full, etc) - revert `30eef52`	7 years ago
Mike Fährmann	7e936e9c06	[luscious] simplify and remove dead code	7 years ago
Mike Fährmann	0245a0ba5f	fix extraction and update test results - fixes for hbrowse, imgyt, imgcandy, hosturimage - test updates for deviantart, gfycat	7 years ago
Mike Fährmann	abd7c559cd	[yonkouprod] remove module Every manga chapter on this site has been removed.	7 years ago
Mike Fährmann	da7219ba74	[kisscomic] remove module Image links on this site are dead.	7 years ago
Mike Fährmann	852e7acd31	[twitter] ignore "Promoted Tweets"	7 years ago
Mike Fährmann	915a0137de	improve 'extractor.request' - add 'fatal' argument - improve internal logic and flow - raise known exception on error - update exception hierarchy	7 years ago
rachmadani haryono	dcd573806e	chg: dev: fix error (#32 ) * fix: dev: error * fix: dev: AttributeError when getting artist * fix: dev: typo on luscious parser	7 years ago
Mike Fährmann	c4713404c8	[directlink] improve URL pattern	7 years ago
Mike Fährmann	d443822fdb	[luacious] get correct image URLs (fixes #33 ) Instead of using thumbnail URLs and modifying them the extractor now goes through every single image-page and gets its download URL from there.	7 years ago
Mike Fährmann	6950708e52	[hentaicdn] use HTTPS	7 years ago
Mike Fährmann	4f1e6c109f	[deviantart] remove 'invalid escape sequence' warning - use r"\w" or "\\w" instead of "\w"	7 years ago
Mike Fährmann	c864be479e	[directlink] update URL pattern & PEP 8 - combine some file extensions - don't match '.je' - line length < 80	7 years ago
H R X N	45f9d64c23	Update directlink.py with additional file exts. (#30 ) Add WebP, still not that common, but it's increasing. Add 3rd JPEG variant (https://en.wikipedia.org/wiki/JPEG#JPEG_filename_extensions) Never seen JFIF in the wild, would probably be overkill. Extend Ogg formats (https://en.wikipedia.org/wiki/Ogg; https://wiki.xiph.org/MIME_Types_and_File_Extensions)	7 years ago
Mike Fährmann	4357966a70	[kissmanga] make URL pattern case-insensitive (fixes 28)	7 years ago
Mike Fährmann	7aa9fa796a	code cleanup and fixes	7 years ago
Mike Fährmann	f08af03845	Merge branch 'cookies'	7 years ago
Mike Fährmann	55f048d02b	ignore case of cookiejar magic strings	7 years ago
Mike Fährmann	f53bf1a323	[thebarchive] add thread extractor	7 years ago
Mike Fährmann	b8cf434bb0	[rebeccablacktech] add thread extractor	7 years ago
Mike Fährmann	808f67ba7d	use 'cookiedomain' for cookies set by object-config-values otherwise these cookies would not be picked up by the _check_cookies() method.	7 years ago
Mike Fährmann	390eeded4c	[mangazuki] support 'raws.…' subdomain	7 years ago
Mike Fährmann	4a60f6068a	[mangazuki] add manga extractor	7 years ago
Mike Fährmann	394241cd6f	[2chan] fix extraction	7 years ago
Mike Fährmann	a13eb6010f	[fallenangels] fix extraction of chapter URLs	7 years ago
Mike Fährmann	1cb1d2e0a3	[mangazuki] add chapter extractor	7 years ago
Mike Fährmann	2f2e363c97	[imgur] use /a/<key>/all as album-url	7 years ago
Mike Fährmann	1cec03c9c6	[imgur] fix extraction of large albums	7 years ago
Mike Fährmann	0610ae5000	skip login if cookies are present	7 years ago
Mike Fährmann	f105782435	[fireden] add thread extractor	7 years ago
Mike Fährmann	c93f7d7496	[archiveofsins] add thread extractor	7 years ago
Mike Fährmann	96e13604da	[archivedmoe] add thread extractor	7 years ago
Mike Fährmann	30d3a5f9b2	support redirects on 4chan archives	7 years ago
Mike Fährmann	98464d1f1b	[loveisover] add thread extractor	7 years ago
Mike Fährmann	47692f28da	[2chan] add thread extractor	7 years ago
Mike Fährmann	3460dc8950	update gallery-dl.conf	7 years ago
Mike Fährmann	9be8f7e106	[deviantart] add "extractor.deviantart.flat" option Setting this to 'false' downloads images into individual subdirectories for each gallery-folder or favourite-collection, otherwise it is just creating a flat list of images.	7 years ago
Mike Fährmann	d075627fd9	[deviantart] support group galleries (#26 ) For groups the 'GalleryExtractor' collects all gallery-folder URLs and defers its work to the 'FolderExtractor'.	7 years ago
Mike Fährmann	b37a62501b	[pixiv] unquote tags	7 years ago
Mike Fährmann	fbd7dcdfdb	[desuarchive] add thread extractor	7 years ago
Mike Fährmann	af9bd17b19	[deviantart] adjust default paths - user.deviantart.com/(gallery\|favourites\|journal)/ images go into * <user>/ * <user>/Favourites/ * <user>/Journal/ (having an extra "Gallery" folder for a user's gallery-images seems a bit too much if these are all you want to download, which is probably the default use-case) - single "deviations" (user.deviantart.com/(art\|journal)/name-123) go into their owner's directory: * <user>/ (putting them into their own directory seems weird in practice)	7 years ago
Mike Fährmann	eb64fb267c	[nyafuu] add thread extractor (#18 )	7 years ago
Mike Fährmann	726c6f01ae	allow 'cookies' config option to be a dictionary	7 years ago
Mike Fährmann	4877ef6314	[deviantart] support '?catpath=/' URLs (#26 ) They previously weren't supported for galleries and journals. This also increases the 'limit' parameter for API calls to its respective maximum.	7 years ago
Mike Fährmann	8c16cbe7ea	fix tests	7 years ago
Mike Fährmann	a6f689e01a	[deviantart] add gallery-folder extractor (#26 ) The code for this and the available metadata is probably going to change again. This extractor is very similar to the favorite- extractor, so they might be "combined" or something like that.	7 years ago
Mike Fährmann	474e9c1aec	[4plebs] add thread extractor (#18 )	7 years ago
Mike Fährmann	a804a42e23	add '--cookies' command-line option	7 years ago
Mike Fährmann	dcc1d3b2ea	[hentaifoundry] fix infinite loop for multiple of 25 images	7 years ago
Mike Fährmann	34e6e1099e	[batoto] adapt to https chapter URLs	7 years ago
Mike Fährmann	85696d0b3b	[reddit] fix issue with datetime errors	7 years ago
Mike Fährmann	80c2e03aaa	[reddit] allow 'date-min/max' to be human readable dates If the date-min/max config value is a string, try parsing it using datetime.strptime [1] with 'date-format' as format string [2] (default: "%Y-%m-%dT%H:%M:%S") Example: get all submissions posted in 2016 $ gallery-dl reddit.com/r/... \ -o date-format=%Y \ -o date-min=\"2016\" \ -o date-max=\"2017\" [1] https://docs.python.org/3/library/datetime.html#datetime.datetime.strptime [2] https://docs.python.org/3/library/datetime.html#strftime-strptime-behavior	7 years ago
Mike Fährmann	58e95a7487	share extractor and downloader sessions There was never any "good" reason for the strict separation between extractors and downloaders. This change allows for reduced resource usage (probably unnoticeable) and less lines of code at the "cost" of tighter coupling.	7 years ago
Mike Fährmann	4414aefe97	small fix for aes_cbc_decrypt_text	7 years ago
Mike Fährmann	21064146c1	fix test	7 years ago
Mike Fährmann	f3d0373120	[reddit] add ability to filter by submission id 'extractor.reddit.id-min' and '….id-max' specify the lowest and highest submission-/post-id to consider, similar to 'date-min' and 'date-max'	7 years ago
Mike Fährmann	1dac76fd1c	update extractor docstrings	7 years ago
Mike Fährmann	92a11528d1	smaller changes	7 years ago
Mike Fährmann	44d98e562b	[pixiv] support pixiv.me URLs (#23 )	7 years ago
Mike Fährmann	b373fe0eea	[pixiv] support shortened URLs and other variants (#23 )	7 years ago
Mike Fährmann	c951d6276c	[imagetwist] use https	7 years ago
Mike Fährmann	d3b04076f7	add .netrc support (#22 ) Use the '--netrc' cmdline option or set the 'netrc' config option to 'true' to enable the use of .netrc authentication data. The 'machine' names for the .netrc info are the lowercase extractor names (or categories): batoto, exhentai, nijie, pixiv, seiga.	7 years ago
Mike Fährmann	e1d82af5e0	small fixes	7 years ago
Mike Fährmann	719d45f89e	[flickr] allow the use of Flickr's specifiers for format selection - renamed the 'width-max' option to 'size-max' - filter by both width and height	7 years ago
Mike Fährmann	b4c438c9ad	[oauth] add the 'extractor.oauth.browser' option enables/disables the use of webbrowser.open() during OAuth authorization	7 years ago
Mike Fährmann	2633337833	[kissmanga] update regex (fixes #20 )	7 years ago
Mike Fährmann	e68af4febe	[flickr] add 'width-max' option (#16 ) This option allows for simple format selection by specifying a maximum image width.	7 years ago
Mike Fährmann	2993206c4b	smaller fixes and "security" measures - move the OAuthSession class into util.py - block special extractors for reddit and recursive - ignore 'only matching' tests for testresults script	7 years ago
Mike Fährmann	8d5e92f641	resolve cyclic dependency between oauth and flickr	7 years ago
Mike Fährmann	d60781de7b	[oauth] workaround for ctrl+c on Windows	7 years ago
Mike Fährmann	9759fe8c6b	allow 'only_matching' tests	7 years ago
Mike Fährmann	56bec79e6a	[reddit] add ability to load more comments (#15 ) The 'extractor.reddit.morecomments' option enables the use of the '/api/morechildren' API endpoint (1) to load even more comments than the usual submission-request provides. Possible values are the booleans 'true' and 'false' (default). Note: this feature comes at the cost of 1 extra API call towards the rate limit for every 100 extra comments. (1) https://www.reddit.com/dev/api/#GET_api_morechildren	7 years ago
Mike Fährmann	05ed95e5b0	[flickr] add search extractor	7 years ago
Mike Fährmann	5f55c854b9	[flickr] replace getPublic... API call with regular ones	7 years ago
Mike Fährmann	9a620784f9	[flickr] add support for user authentication (#16 ) Call '$ gallery-dl oauth:flickr' to get an access_token and access_token_secret for your account.	7 years ago
Mike Fährmann	d5a70f2580	add simple progress indicator for multiple URLs (#19 ) The output can be configured via the 'output.progress' config value. Possible values: - true: Show the default progress indicator "[{current}/{total}] {url}" (default) - false: Never show the progress indicator - <string>: Show the progress indicator using this as a custom format string(1). Possible replacement keys are: - current: current URL index - total : total number of URLs - url : current URL (1) https://docs.python.org/3/library/string.html#formatstrings	7 years ago
Mike Fährmann	3ee77a0902	[oauth] print URL if webbrowser.open fails	7 years ago
Mike Fährmann	090e11b35d	[reddit] enable user authentication with OAuth2 (#15 ) Call '$ gallery-dl oauth:reddit' to get a refresh_token for your account.	7 years ago
Mike Fährmann	e682e06518	[flickr] add group extractor (#16 )	7 years ago
Mike Fährmann	8fd66ef0b3	[flickr] add gallery extractor (#16 )	7 years ago
Mike Fährmann	8456b84a12	fix tests and small stuff	7 years ago
Mike Fährmann	fbfc8d0f78	[reddit] ignore Authorization errors for subreddits - also made the limit for retrieved comments customizable via the 'extractor.reddit.comments' config value - default is 500; 0 ignores comments completely	7 years ago
Mike Fährmann	e365f1d799	[pixiv] rewrite - same functionality, better(?) code quality, easier to extend - added test for the user-tag functionality - removed the 'artist-id', 'artist-name' and 'artist-nick' keywords, which can be replaced with 'user[id]', 'user[name]' and 'user[account]' respectively	7 years ago
aiasdfd	338f79147f	[pixiv] support tag for user downloads (#17 ) [pixiv] support tag for user downloads	7 years ago
Mike Fährmann	5f05543f23	[reddit] support filtering by timestamp (#15 ) - Added the 'extractor.reddit.date-min' and '….date-max' config options. These values should be UTC timestamps. - All submissions not posted in date-min <= T <= date-max will be ignored. - Fixed the limit parameter for submission comments by setting it to its apparent max value (500).	7 years ago
Mike Fährmann	4e80e0c884	[flickr] add user extractor (#16 )	7 years ago
Mike Fährmann	b81d068a6d	[flickr] add favorites extractor (#16 )	7 years ago
Mike Fährmann	c921b4f32a	code cleanup and fixing tests	7 years ago
Mike Fährmann	72f1c6f87a	[flickr] add support for flic.kr/p/... URLs Example: https://flic.kr/p/FPVo9U	7 years ago
Mike Fährmann	93e5d8cba3	[flickr] add album extractor	7 years ago
Mike Fährmann	659c65dbb0	[flickr] add image extractor	7 years ago
Mike Fährmann	b6fffa9e26	[directlink] update filename format and metadata	7 years ago
Mike Fährmann	c184e47ee3	put common directory- and filename formats in base classes	7 years ago
Mike Fährmann	bce51e90a5	[reddit] support sorting options and sub-options (#15 ) Example: https://www.reddit.com/r/<subreddit>/top/?sort=top&t=month (the 'sort=top' parameter is irrelevant and can be omitted)	7 years ago
Mike Fährmann	5f45ce2930	[gfycat] add "format" config key to select a video format Possible values: - one of "mp4" (default), "webm", "gif", "webp", "mjpg" If the selected format is not available, "mp4", "webm" and "gif" (in that order) will be tried instead, until an available format is found.	7 years ago
Mike Fährmann	011659ced5	[imgur] add "mp4" config key to decide between GIF and MP4 possible values: - false : always use GIF - true : use MP4 if "prefer_video" flag is set, GIF otherwise (default) - "always": always use MP4	7 years ago
Mike Fährmann	48ccee2505	[gfycat] add image extractor	7 years ago
Mike Fährmann	bf452a8516	[imgur] choose .mp4 over .gif if available	7 years ago
Mike Fährmann	f79320e35b	fix tests	7 years ago
Mike Fährmann	67791e1b36	[imgur] improve and add image extractor	7 years ago
Mike Fährmann	99b72130ee	[reddit] enable recursion (#15 ) reddit extractors now recursively visit other submissions/posts linked to in the initial set of submissions. This behaviour can be configured via the 'extractor.reddit.recursion' key in the configuration file or by `-o recursion=<value>`. Example: {"extractor": { "reddit": { "recursion": <value> }}} Possible values: * -1 - infinite recursion (don't do this) * 0 - recursion is disabled (default) * 1 and higher - maximum recursion level	7 years ago
Mike Fährmann	691c4dd709	support direct image links	7 years ago
Mike Fährmann	d2dceb35b7	implement context-manager to blacklist extractors	7 years ago
Mike Fährmann	e425243b1e	[reddit] some small fixes - filter or complete some URLs - remove the 'nofollow:' scheme before printing URLs - (#15)	7 years ago
Mike Fährmann	a22892f494	[reddit] add subreddit- and submission-extractor - these extractors scan submissions and their comments for (external) URLs and defer them to other extractors - (#15)	7 years ago
Mike Fährmann	832a4a8ee9	[fallenangels] add manga extractor	7 years ago
Mike Fährmann	f226417420	simplify code by using a MangaExtractor base class	7 years ago
Mike Fährmann	2974d782a3	[yomanga] remove module site has been shut down	7 years ago
Mike Fährmann	cbb4323f66	add setup.cfg to configure flake8	7 years ago
Mike Fährmann	232fe2dd08	improve the test extractor	7 years ago
Mike Fährmann	b0131ea402	[fallenangels] support this site's Vietnamese version - https://truyen.fascans.com/	7 years ago
Mike Fährmann	b6b214f7e9	[deviantart] fix headers for custom-style journals example: http://shimoda7.deviantart.com/journal/Temporary-absence-231936282	7 years ago
Mike Fährmann	e9a2738257	[deviantart] support images on top of journal entries example: http://raxnae.deviantart.com/art/Kami-s-Journal-679482236	7 years ago
Mike Fährmann	92597f46d4	[deviantart] add title to journals	7 years ago
Mike Fährmann	107d29ad8a	improve handling of text:... URLs - don't require // after the colon - open output files in text mode	7 years ago
Mike Fährmann	677c8ced11	[deviantart] add "journal" extractor (#14)	7 years ago
Mike Fährmann	e5f79ae839	[deviantart] add support for all media types - this includes - images - videos - flash-animations - journals - also renamed some of the extractors - User -> Gallery - Image -> Deviation	7 years ago
Mike Fährmann	9f1c83297f	[pinterest] allow URLs with any TLD	7 years ago
Mike Fährmann	b3b92ac243	[deviantart] support "All" favorites and add "mature" option - since there is apparently no actual way to get the "All" favorites listing via API, corresponding URLs (.../favourites/?catpath=/) will be handled by yielding all deviations from all favorite collections of that user - the "mature" config key works on a per extractor basis (like "username" or "password"). values can be the strings "true" or "false", or the booleans true or false. - (#14)	7 years ago
Mike Fährmann	7376ad7f3d	[deviantart] turn the "Mature Content Filter" off (#14)	7 years ago
Mike Fährmann	cfbf79d788	[pixiv] fix login	7 years ago
Mike Fährmann	85a46ed700	[booru] fix issue with multiple tags	7 years ago
Mike Fährmann	fc9223c072	add '--abort-on-skip' option and ability to control skip behavior the 'skip' config option controls skipping behavior: true - skip download if file already exist (default) false - download and overwrite files even if it exists "abort" - abort extractor run if a download would be skipped (same as '--abort-on-skip')	7 years ago
Mike Fährmann	d948ba1322	[readcomics] remove module - site has been unavailable for two weeks - (#12)	7 years ago
Mike Fährmann	a610b35a0d	[mangashare] remove module this site has been unavailable for at least two months	7 years ago
Mike Fährmann	4e8587bad4	[pixiv] add support for https://i.pximg.net URLs	7 years ago
Mike Fährmann	e41efbd2d9	[kissmanga] fix edge-case	8 years ago
Mike Fährmann	ffd72424bf	[kissmanga] another attempt at getting the AES key	8 years ago
Mike Fährmann	af56887a47	[exhentai] fall back to e-hentai if no username is given	8 years ago
Mike Fährmann	4b967fa189	implement and use extractor.config() method	8 years ago
Mike Fährmann	82ab1fca07	[seiga] reduce cache maxage to one week	8 years ago
Mike Fährmann	ec48d25afc	[pawoo] fix extraction results	8 years ago
Mike Fährmann	244ab75cad	[kissmanga] update AES key retrieval	8 years ago
Chen John L	a5485a46cb	fixed the module for pixhost	8 years ago
Mike Fährmann	13dc5d72bc	update some extractors to use https	8 years ago
Mike Fährmann	342371086b	[pawoo] add extractors for accounts and statuses https://pawoo.net is a Mastodon[1] instance hosted by Pixiv [1] https://github.com/tootsuite/mastodon	8 years ago
Mike Fährmann	0770de0ea1	[deviantart:image] add support for sta.sh URLs	8 years ago
Mike Fährmann	f4aa452bd1	update unit test results	8 years ago
Mike Fährmann	71e08dc9c4	[tumblr] keyword consistency	8 years ago

... 4 5 6 7 8 ...

976 Commits (392a08165795b5579c5e8861a9b925520dad469b)