POST /api/corriere/articles
Price: 5 credits
Get one Corriere della Sera article by its URL: headline, standfirst, full body text, first publication and last update times, byline, section breadcrumb, lead image with caption, in-body images and, on live pages, the timeline of live updates with their own timestamps. Works for the national site and for the local editions on the regional subdomains.
Fetch one Corriere della Sera article by its full page URL: an article URL ending in .shtml, or a video.corriere.it URL, which carries no extension. The content id alone does not resolve, so pass the whole URL. Returns article_title, alternative_headline (the editorial sub-headline, when the page carries one different from the title), description, text, published_at and updated_at (unix), section and sections (the breadcrumb path, falling back to the section the page declares in its own metadata), byline, image with image_caption, images, edition (www for the national site, or the regional subdomain such as roma or milano), video_duration on video pages and content_type (article, video or live). Live pages also return live_updates, each with its own id, update_title, text and published_at. is_accessible_for_free is the paywall flag Corriere publishes in its own page metadata, reported as-is; it does not predict whether the body is served, and the body text comes back in full to an anonymous client across the whole archive, 2010 included. Photo galleries and the 'chiedi all'esperto' Q&A pages are the templates that carry no prose. Feed web_url values from corriere/articles/search or corriere/sections/articles into this endpoint.
access-token string requiredtimeout integer — Max scrapping execution timeout (in seconds) (default: 300; min: 20; max: 1500)article string required — Corriere della Sera article URL (examples: "https://www.corriere.it/esteri/26_agosto_14/trump-riporta-le-portaerei-al-vapore-e-intanto-l-arsenale-usa-si-consuma-8046f3a9-6e8f-4a19-9ed1-d70429339xlk.shtml", "https://roma.corriere.it/notizie/politica/26_agosto_13/caso-ranucci-rai-danni-immagine-b792bc9f-4213-4fcc-a9f0-f75efabfdxlk.shtml"; minLength: 1)@type string (default: "CorriereArticle")id string requiredarticle_title string requiredalternative_headline string nullabledescription string nullabletext string nullableweb_url string requiredpublished_at integer nullableupdated_at integer nullablesection string nullablesections array (default: [])content_type string requirededition string requiredbyline string nullableimage string nullableimage_caption string nullableimages array (default: [])@type string (default: "CorriereImage")image string requiredwidth integer nullableheight integer nullablecredit string nullablepreset string nullablecaption string nullablealt string nullablevideo_duration string nullableis_accessible_for_free boolean nullablelive_updates array (default: [])@type string (default: "CorriereLiveUpdate")id string requiredupdate_title string nullabletext string nullablepublished_at integer nullable422 — Not a Corriere della Sera article URL Check the fields against this schema. A URN with the wrong prefix is the most common cause.408 — The request ran past its time limit Raise `timeout` in the request body, up to the maximum this endpoint documents. Lowering `count` or turning off the `with_*` flags also helps, because less work finishes sooner.412 — No Corriere article at this URL, or the URL redirected to a different article Retrying will not help: either the entity does not exist, or the input points at a different one.429 — Too many requests: a rate limit or a usage window is exhausted When the response carries an X-Retry-After header, wait that many seconds and retry: the same number is in the body as `detail.retry_after`, and the limit clears once that window passes. The message in the body names the limit that was hit.500 — Something broke on our side Retrying will not help. If it keeps happening, send us the X-Request-ID from the response headers.529 — Rate limit reached, or the endpoint is overloaded Wait at least 30 seconds, then retry.X-Error — Error message text (present only on error)X-Request-ID — Unique request identifierX-Execution-Time — Execution time in secondsX-Result-Count — How many records the body carries. 0 means an empty result, which is a normal answer and not by itself an error. A non-zero count can come back together with X-Error when the failure happened partway through — read this header and X-Error independently.X-Total-Available-Results — How many records exist for this query, when the endpoint can say. On a `dry_run` request this is the answer and the body is empty. It saturates: the endpoint's documented maximum means 'at least that many', any smaller number is exact.X-Warning — Present when the request body carried keys this endpoint does not document. They were ignored, so any filter you meant to apply through them did not apply. Check the spelling against this schema and retry.X-Retry-After — Seconds to wait before retrying. Present only on 429.