ApifyScheduler
Index
Methods
__init__
Parameters
optionalasync_thread_timeout: timedelta = timedelta(seconds=60)
optionalcrawler: Crawler | None = None
Returns None
close
Close the scheduler.
Resolve the requests Scrapy still holds, then shut down the event loop and its thread gracefully.
Parameters
reason: str
The reason for closing the spider.
Returns None
enqueue_request
Add a request to the scheduler.
This could be called from either from a spider or a downloader middleware (e.g. redirect, retry, ...).
Parameters
request: Request
The request to add to the scheduler.
Returns bool
True if the request was successfully enqueued, False otherwise.
from_crawler
Create the scheduler, reading the async-thread timeout from the Scrapy settings.
The
APIFY_ASYNC_THREAD_TIMEOUT_SECSsetting (in seconds) caps how long each coroutine run on the background event loop may take before timing out; it defaults to 60 seconds.Parameters
crawler: Crawler
Returns ApifyScheduler
has_pending_requests
Check if the scheduler has any pending requests.
Resolves the requests Scrapy has finished with first, as their outcome is what decides the answer.
Returns bool
True if the scheduler has any pending requests, False otherwise.
next_request
Fetch the next request from the scheduler.
Returns Request | None
The next request, or None if there are no more requests.
open
Open the scheduler.
Parameters
spider: Spider
The spider that the scheduler is associated with.
Returns Deferred[None] | None
A Scrapy scheduler that uses the Apify
RequestQueueto manage requests.A request stays unresolved in the RQ until Scrapy is done with it, so an interrupted run leaves it pending for the next one. When the platform is about to migrate the Actor run, the scheduler stops handing out requests; then, as when the run is being aborted, it marks the ones Scrapy finishes as handled, so the next run does not repeat them.
This scheduler requires the asyncio Twisted reactor to be installed.