Mike Fährmann
48ef062867
fix issues with 'Extractor.finalize()'
...
- prevent crash in InstagramUserExtractor (#4359 )
- call it at the end of every DownloadJob
- add it to tests
1 year ago
Mike Fährmann
ed21908fda
initial support for child extractor options
...
Using "parent-category>child-category" as extractor category in a config
file allows to set options for a child extractor when it was spawned by
that parent.
For example "reddit>gfycat" to set gfycat options for when it was found
in a reddit post.
{
"extractor": {
"gfycat": {
"filename": "regular filename"
},
"reddit>gfycat": {
"filename": "reddit-specific filename"
}
}
}
Note: This does currently not work for most imgur links due to how its
extractor hierarchy is structured.
1 year ago
Mike Fährmann
255d08b79e
add test for 'Extractor.initialize()' ( #4359 )
1 year ago
Mike Fährmann
2bcf0a4c49
[instagram] fix initialization order ( #4359 )
...
regression caused by the changes in a383eca7
1 year ago
Mike Fährmann
7eab101144
[acidimg] fix extraction
...
swap ' and " again (2e309a13
)
and add a fallback in case this happens yet another time
1 year ago
Mike Fährmann
62fce6a75f
[imagehosts] adjust variable names ( #4358 )
...
prefix them with underscores to prevent a clash
with the new 'self.cookies' from d97b8c2f
1 year ago
Mike Fährmann
e8299b459a
[moebooru] match search URLs with empty 'tags' ( #4354 )
1 year ago
Mike Fährmann
7fbc304ae9
[twitter] fix crash on private user ( #4349 )
1 year ago
Mike Fährmann
1ece3b92ff
[mangadex] allow multiple values for 'lang' ( #4093 )
...
This was already possible by setting 'lang' to a list of strings,
but now it can also be done as a more command-line friendly string.
-o lang=fr,it
1 year ago
Mike Fährmann
52053b58f0
[lensdump] fix extraction ( #4352 )
1 year ago
Mike Fährmann
11f71a9cba
remove 'mememuseum' module
...
This was forgotten when adding generic Shimmie2 support in 7865067d
1 year ago
Mike Fährmann
a383eca7f6
decouple extractor initialization
...
Introduce an 'initialize()' function that does the actual init
(session, cookies, config options) and can called separately from
the constructor __init__().
This allows, for example, to adjust config access inside a Job
before most of it already happened when calling 'extractor.find()'.
1 year ago
Mike Fährmann
6c9432165e
add return value to 'PostProcessor._init_archive()'
1 year ago
Mike Fährmann
54d974deb0
add 'python' post processor
...
similar to 'exec' but calls a Python function
1 year ago
Mike Fährmann
1baf83a9e5
[hiperdex] fix for unicode titles ( #4325 )
1 year ago
Mike Fährmann
7da954f810
[flickr] update default API credentials ( #4332 )
...
and add a delay between API requests
1 year ago
Mike Fährmann
a45a17ddb7
[pixiv] ignore 'limit_sanity_level' images ( #4328 )
1 year ago
Mike Fährmann
088e8d5fcf
[pornhub] fix extraction ( #4301 )
1 year ago
Mike Fährmann
d97b8c2fba
consistent cookie-related names
...
- rename every cookie variable or method to 'cookies_*'
- simplify '.session.cookies' to just '.cookies'
- more consistent 'login()' structure
1 year ago
Mike Fährmann
ceebacc9e1
remove 'pyopenssl' option
1 year ago
Mike Fährmann
3c2c7e21dd
merge #4319 : [zerochan] fix 'tags' extraction
1 year ago
Mike Fährmann
0ba8d1f168
merge #4312 : [redgifs] add 'niches' extractor
1 year ago
Mike Fährmann
c5565f79f7
merge #4096 : [danbooru] add support for booru.borvar.art instance
1 year ago
Mike Fährmann
63326e3168
[danbooru] add tests for booruvar
1 year ago
Mike Fährmann
5171d8975c
[E621] support 'e6ai.net' ( #4320 )
1 year ago
Mike Fährmann
a996d936d2
[imagefap] fix pagination ( #3013 )
1 year ago
Mike Fährmann
22099422ca
[deviantart] fix shortened URLs ( #4316 )
1 year ago
Mike Fährmann
90231f2d5a
[twitter] add 'tweet-endpoint' option ( #4307 )
...
use the newer TweetResultByRestId only for guests by default
1 year ago
Mike Fährmann
20ed647f6f
[twitter] add 'user' extractor and 'include' option ( #4275 )
1 year ago
Mike Fährmann
86be197d11
[twitter] remove '/search/adaptive.json'
1 year ago
enduser420
d52ed2bc5a
[zerochan] fix 'tags' extraction
1 year ago
enduser420
12cd85658b
[redgifs] add 'niches' extractor
1 year ago
Mike Fährmann
248e8bc699
release version 1.25.8
1 year ago
Mike Fährmann
bc9123cfee
[naverwebtoon] fix 'comic' metadata extraction
1 year ago
Mike Fährmann
ab5dde7221
[mangaread] fix 'tags' extraction
1 year ago
Mike Fährmann
c9a82c9313
[erome] ignore duplicate album IDs
1 year ago
Mike Fährmann
c84397023a
[slideshare] fix extraction
1 year ago
Mike Fährmann
ffbbbd3baf
[gelbooru_v01] 'vidyart' -> 'vidyart2'
1 year ago
Mike Fährmann
e40b90e137
merge #4303 : [gelbooru_v01] fix 'source' ( #4302 )
1 year ago
Mike Fährmann
c6b31a2169
[reddit] set default 0.6s delay between requests ( #4292 )
...
to limit API requests to 100 per minute
https://www.reddit.com/r/redditdev/comments/14nbw6g/
1 year ago
Mike Fährmann
20da41018d
[pornhub] set 'accessAgeDisclaimerPH' cookie ( #4301 )
1 year ago
ncaat
75757c4ace
[gelbooru_v01] fix 'source' ( #4302 )
1 year ago
Mike Fährmann
2dd6942d1c
[jpgfish] update domain to 'jpeg.pet'
1 year ago
Mike Fährmann
1137b89ed4
[lineblog] remove module
...
"LINE BLOGは2023年6月29日をもちましてサービスを終了いたしました"
1 year ago
Mike Fährmann
86560fe0cd
[bcy] remove module
...
"The website was shut down on July 12, 2023"
https://danbooru.donmai.us/wiki_pages/bcy
1 year ago
Mike Fährmann
fceabee433
[philomena] use API interface class
...
handle 429 errors and retry after 10min (#4288 )
1 year ago
Mike Fährmann
f079d9a703
[reddit] notify users about registering an oauth application
...
(#4292 , #4253 , #3943 )
1 year ago
Mike Fährmann
fb3d1462b1
merge #4291 : [wikifeet] fix 'tag' extraction
1 year ago
Mike Fährmann
0b08e2e8a8
merge #4287 : [twitter] Fix following extractor not getting all users
1 year ago
Mike Fährmann
f6553ffd2f
[twitter] simplify '_pagination_users'
...
- remove 'stop' variable
- call 'cursor.startswith()' only once
1 year ago
Mike Fährmann
1590124aae
[twibooru] fix '--range'
1 year ago
enduser420
a2111dd025
[wikifeet] fix 'tag' extraction
1 year ago
Mike Fährmann
a1ffa1ff09
[philomena] fix '--range' ( #4288 )
1 year ago
Mike Fährmann
a27dbe8c82
[twitter] use 'TweetResultByRestId' endpoint ( #4250 )
...
allows accessing single Tweets without login
1 year ago
Mike Fährmann
d3d639a159
[twitter] don't treat missing 'TimelineAddEntries' as fatal ( #4278 )
1 year ago
ActuallyKit
c321c773f2
make the code less ugly
1 year ago
ActuallyKit
a437a34bcf
fix lint i guess?
1 year ago
ActuallyKit
6cbc434b54
Fix users pagination
1 year ago
Mike Fährmann
d5b6802774
[seiga] set 'skip_fetish_warning' cookie ( #4242 )
1 year ago
Mike Fährmann
88d1e29401
[bunkr] use '.la' TLD for 'media-files12' servers ( #4147 , #4276 )
1 year ago
Mike Fährmann
f0cb951566
[paheal] unescape 'source'
1 year ago
Mike Fährmann
b480b7076a
[paheal] fix a78f8ce5
for enabled 'metadata' ( #4262 )
1 year ago
Mike Fährmann
384337d3dd
[fantia] send 'X-Requested-With' header only for API requests ( #4273 )
1 year ago
Mike Fährmann
c2ac665ff7
[fantia] send 'X-Requested-With' header ( #4273 )
1 year ago
Mike Fährmann
7444fc125b
[gfycat] implement login support ( #3770 , #4271 )
...
For the record: '/webtoken' and '/weblogin' are not the same ...
1 year ago
Mike Fährmann
e9b9f751bf
[gfycat] support '@me' user ( #3770 , #4271 )
1 year ago
Mike Fährmann
5b59a0d143
update default User-Agent header to Firefox 115 ESR
1 year ago
Mike Fährmann
0556e1ad45
merge #4268 : [newgrounds] extract & pass auth token for login
1 year ago
Mike Fährmann
a16d7c59cb
[newgrounds] access 'response.text' only once
1 year ago
Mike Fährmann
1bf9f52c99
[twitter] add 'ratelimit' option ( #4251 )
1 year ago
Mike Fährmann
f86fdf64a6
[twitter] use GraphQL search by default ( #4264 )
1 year ago
Mike Fährmann
1d4db83d49
[weibo] fix end of cursor based pagination
1 year ago
Mike Fährmann
a78f8ce5b0
[paheal] fix extraction ( #4262 )
...
swap ' and "
1 year ago
FrostTheFox
9576652fa5
extract & pass auth token for newgrounds
1 year ago
Mike Fährmann
5457007dd3
release version 1.25.7
1 year ago
Mike Fährmann
3d8de383bf
[mangapark] extract 'source_id' for manga
...
forgot to add this to 6ae3101f
1 year ago
Mike Fährmann
6ae3101fd0
[mangapark] add 'source' option ( #3969 )
1 year ago
Mike Fährmann
c45a913bfd
[flickr] add 'exif' option
1 year ago
Mike Fährmann
3845c0256d
[sankaku] improve warnings for unavailable posts
1 year ago
Mike Fährmann
46cae04aa3
[piczel] update API server ( #4244 )
1 year ago
Mike Fährmann
3479646f65
[mangapark] update and fix 'manga' extractor ( #3969 )
...
TODO:
- non-English chapters
- 'source' option
1 year ago
Mike Fährmann
10786c657e
[mangapark] update and fix 'chapter' extractor ( #3969 )
1 year ago
Mike Fährmann
9c31c2daef
[poipiku] improve error detection ( #4206 )
1 year ago
Mike Fährmann
260ff55e19
[senmanga] ensure download URLs have a scheme ( #4235 )
1 year ago
Mike Fährmann
ccbc1a1d55
[flickr] add 'metadata' option ( #4227 )
1 year ago
Mike Fährmann
c1cce4a80b
[twitter] extend 'conversations' option ( #4211 )
1 year ago
Mike Fährmann
b6c959744d
[furaffinity] improve 'description' HTML ( #4224 )
...
- ignore header
- include footer and closing <div> if present
1 year ago
Mike Fährmann
8357acf359
[gelbooru_v01] replace 'extract_all()' with 'extract_from()'
...
It's even slightly faster, especially on Python before 3.11
1 year ago
Mike Fährmann
068aa26c3e
[gelbooru_v01] fix '--range' ( #4167 )
1 year ago
Mike Fährmann
2052e7ce59
[hentaifox] fix titles containing '@' ( #4201 )
1 year ago
Mike Fährmann
92d98697b2
[wallhaven] update API error message
1 year ago
Mike Fährmann
a673998b1e
release version 1.25.6
1 year ago
Mike Fährmann
339fcdb8ad
[wallhaven] handle '429 Too Many Requests' errors ( #4192 )
...
- set 1.4s delay between API requests
(WH allows 45 requests per minute)
- wait and retry on 429 errors
1 year ago
Mike Fährmann
ef9891ec9d
[fantia] extract 'plan' metadata ( #2477 , #4128 )
1 year ago
Mike Fährmann
f8452984fa
[fantia] emit warning for non-visible contents ( #4128 )
1 year ago
Mike Fährmann
dc7af00014
[fantia] refactor
...
- embed response data as hidden '_data' field
(instead of returning/passing 'resp')
- split _get_urls_from_post()
1 year ago
Mike Fährmann
6c8bf9a762
[pornhub] improve redirect handling ( #4188 )
1 year ago
Mike Fährmann
654267a335
[weibo] fix 'json' extension for some videos
1 year ago
Mike Fährmann
ce93c460a6
[formatter] implement 'H' conversion ( #4164 )
...
to remove HTML tags and unescape HTML entities
1 year ago
Mike Fährmann
deff3b434d
[vipergirls] implement login support ( #4166 )
1 year ago
Mike Fährmann
db20a645c5
[vipergirls] use API endpoints ( #4166 )
1 year ago
Mike Fährmann
0b34a444e0
[pixiv:novel] only detect Pixiv embeds ( #4175 )
1 year ago
Mike Fährmann
9f1aee3884
[vipergirls] limit number of requests per second ( #4166 )
1 year ago
Mike Fährmann
21c75d03a3
merge #4133 : [furaffinity] extract 'favorite_id' metadata
1 year ago
Mike Fährmann
5e3a1749c8
[furaffinity] simplify 'favorite_id' assignment
1 year ago
Mike Fährmann
ad882291d3
[instagram] fix retrieving '/tagged' posts ( #4122 )
...
reduce number of retrieved posts per API request from 50 to 20
1 year ago
Mike Fährmann
0a9aaa7a8d
[weibo] prevent fatal exception due to missing video ( #4150 )
1 year ago
Mike Fährmann
ac651c604c
[senmanga] fix and update ( #4160 )
1 year ago
Mike Fährmann
df106fb58b
[bunkr] fix video downloads
1 year ago
Mike Fährmann
aad5e6490c
merge #4159 : [bunkr] update domain to bunkrr.su
1 year ago
Mike Fährmann
e0522ffb3d
[bunkr] update
1 year ago
Mike Fährmann
e04796e04b
merge #3447 : [jschan] add generic extractors for jschan imageboards
1 year ago
Mike Fährmann
b9692341fe
[jschan] update
1 year ago
Stephan
a7c066cbac
Update bunkr.py
1 year ago
Stephan
72e697b8b5
Update bunkr.py
...
Support bunkrr.su
1 year ago
Mike Fährmann
4ae925c88f
[kemonoparty] support '.su' TLD ( #4139 )
1 year ago
Mike Fährmann
2d9e3093ca
merge #4134 : [postimage] add gallery support, update image extractor
1 year ago
Mike Fährmann
e64b521287
merge #4136 : [acidimg] fix extractor
1 year ago
Mike Fährmann
a90974178d
[jpgfish] update domain to 'jpg.pet' ( #4138 )
1 year ago
Mike Fährmann
ee959052ac
merge #4138 : add jpg.pet as alias for jpgfish
1 year ago
Mike Fährmann
0281cc7d08
[fanbox] skip 404ed fanbox embeds ( #4088 )
...
continuation of 4fc9675d
1 year ago
Prinz23
97c0d13cbb
add jpg.pet as alias for jpgfish
1 year ago
chio0hai
2e309a13a7
[acidimg] fix extractor
1 year ago
chio0hai
92178b369c
[postimage] add gallery support, update image extractor to download
...
original image instead of main image
1 year ago
Bad Manners
952c03bc9e
Add fav_id data to FuraffinityFavoriteExtractor
...
An extra field is collected when paginating favorites, and saved to
a temporary cache variable. This field is identical for both the old
and the new page layouts for FurAffinity, but can only be collected
during pagination, hence the cache variable. Other FurAffinity
extractors should be unaffected by this change.
1 year ago
Mike Fährmann
54cf1fa3e7
[twitter] use GraphQL search endpoint ( #3942 )
...
for guest users; selectable with 'search-endpoint' option.
adapted from 9c7b888ffa
1 year ago
Mike Fährmann
864a654b25
[twitter] update query hashes
1 year ago
Mike Fährmann
45cc7cee1a
[twitter] better error message for guest searches ( #3942 )
1 year ago
Mike Fährmann
271f23d971
[twitter] extract 'conversation_id' metadata ( #3839 )
1 year ago
Mike Fährmann
94b6a67666
[reddit] fix crash with empty 'crosspost_parent_lists' ( #4120 )
1 year ago
Mike Fährmann
0cf7282fa0
[pixiv] add 'full-series' option for novels ( #4111 )
1 year ago
Mike Fährmann
bab13402df
[redgifs] update 'search' URL pattern ( #4115 )
1 year ago
Mike Fährmann
5a6fd8027d
[redgifs] support galleries ( #4021 )
1 year ago
Mike Fährmann
0ad59c92b1
[blogger] download files from 'lh*.googleusercontent.com' (4070)
1 year ago
Mike Fährmann
ffed7efb6f
[pixiv] use BASE_PATTERN
1 year ago
Mike Fährmann
b286efefcc
[pixiv] add 'novel-bookmark' extractor ( #4111 )
1 year ago
Mike Fährmann
5283db1aae
release version 1.25.5
1 year ago
Mike Fährmann
28f6487c64
[instagram] add 'metadata' option ( #3107 )
1 year ago
Mike Fährmann
8cf13f8696
merge #4104 : [lensdump] add lensdump.com extractors
1 year ago
Mike Fährmann
58f7480d46
[lensdump] update
...
- update docs/supportedsites.md
- add GPL2 header
- use BASE_PATTERN
- improve LensdumpImageExtractor
1 year ago
Mike Fährmann
3516fdae74
[kemonoparty] fix kemono and coomer logins using the same cache
...
(#4098 )
1 year ago
chio0hai
d5300cf381
[lensdump] subcategory
1 year ago
chio0hai
82ba6bfdc0
[lensdump] f-string fix
1 year ago
chio0hai
9b2326e4e1
[lensdump] add lensdump.com extractor
1 year ago
Mike Fährmann
a5d0b03bde
[ytdl] fix crash due to removed 'no_color' attribute
...
8417f26b8a
1 year ago
Mike Fährmann
148bdc04a4
merge #2719 : [jpgfish] add 'jpgfish' extractors
1 year ago
Mike Fährmann
609c4f3e07
[jpgfish] simplify and improve
1 year ago
Mike Fährmann
2b1f875ef4
[jpgchurch] update to 'jpgfish'
1 year ago
Mike Fährmann
3d29c42142
[mangaread] fix 'tags' extraction
1 year ago
Mike Fährmann
5f86527cbe
merge #2781 : [mangaread] Add Mangaread extractor
1 year ago